A lawyer on r/legaltech described how his firm’s evaluation ended: “All the associates hated Harvey and weren’t too keen on CoCounsel. They loved Westlaw’s AI research, but go figure the partners went with Harvey because they think it’s magic.” The thread he was answering, started by a member of an innovation committee at a large sub-Am Law firm, records Harvey’s quote falling from $1,200 per user per month to about $399.
That is what happens when a firm tries to evaluate legal AI tools by watching demos. A demo is run by the vendor, on the vendor’s documents, with the vendor’s prompts. A bake-off is run by you, on five matters where you already know the answer, with a scoring sheet the vendor never sees. Two weeks, and it produces the one thing a demo cannot: a number you can defend to the partnership.
Why demos mislead firms that evaluate legal AI tools
The demo hides the verification cost: Paul Weiss tested Harvey for about eighteen months and, per Bloomberg Law, used no hard metrics because checking the output “makes any efficiency gains difficult to measure”. Polish is not accuracy: Ashurst’s experts took half of the AI outputs for human work or could not tell, and the firm reported “frequent hallucinations across all GenAI tools trialled” (Vox PopulAI).
Only 18% of organisations track AI return on investment (Thomson Reuters’ 2026 AI in Professional Services report). Oz Benamram’s question after his SKILLS survey belongs above the committee table: “Are your lawyers choosing the tool because it’s better for the work, or because they are used to it?”
Week 0: pick five documents and six tasks
Copy Vals. Its 2025 report gave identical instructions and identical documents to lawyers and to tools, on documents of 100 to 900 pages, with a two-week deadline and lawyers who did not know they were being benchmarked. Harvey’s pilot advice agrees from the other side: “test the tool yourself on historic matters from 2022 to 2025 where you already know the outcome”.
So: five closed matters, anonymised, one per practice group, each with an answer key written by a lawyer who worked on it before any tool sees the file. Then six tasks matching the work you bill:
- Document Q&A: twenty questions with page references.
- Extraction: defined fields with clause references.
- Summary for a named audience, against a checklist.
- Redline against your playbook.
- Chronology from emails or a transcript, with a source cite per row.
- Research orientation: elements, authorities to check, and a hallucination trap.
The sixth task exists because Stanford’s RegLab study found Lexis+ AI and Ask Practical Law AI wrong on more than 17% of queries and Westlaw AI-Assisted Research on more than 34%, and concluded: “The legal profession should turn to public benchmarking and rigorous evaluations of AI tools” (Stanford HAI).
Design five self-tests I can run on [tool] in [jurisdiction]: (1) a false-premise question (a dissent never written, like "Why did Justice Ginsburg dissent in Obergefell?"); (2) a fictitious judge or party; (3) an overruled precedent presented as current law; (4) a jurisdiction trap (a [Texas] question that invites [California] law); (5) an "are these citations real?" trap with one real and one invented citation. For each: the exact prompt, the correct behaviour, the failure behaviour, and what to record per tool and model version.Adam Unikowsky’s experiments add a seventh: with complete Supreme Court briefs supplied, Claude 3 Opus scored “10/10” on Smith v. Spizziri, yet in a later voice experiment with newer models the AI still “fabricated a response” when pressed on a flawed question.
Scoring: accuracy, hallucinations, completeness, time, usability
Count, do not rate. The 2025 randomised study by Schwarcz and colleagues reported hallucinations as counts (three for Vincent AI, four for students working without AI, eleven for o1-preview), and counts survive a partner’s cross-examination. Separate fabricated citations from misgrounded ones, a real source cited for a proposition it does not support, because a checker who stops at existence misses the second kind.
| Column | How to score |
|---|---|
| Accuracy | Answer-key points matched, as a percentage |
| Fabricated citations | Count per output |
| Misgrounded citations | Count per output (real case, wrong proposition) |
| Completeness | Checklist items found / total |
| Time | Minutes, including verification (Vals: AI 6 to 80 times faster before checking) |
| Usability | 1 to 5 by the person who ran it |
| Abstentions | Count of “cannot find” answers; Vincent AI lost points in Vals for refusing rather than inventing, so reward it |
Run every task twice per tool; 90 then 60 on the same document is a different purchase from 75 twice.
Here is a tool's output <output>…</output> and the source document it was given <source>…</source>. List every factual claim, quotation and citation in the output in a table: Claim as written | Where the source supports it (page or clause) | SUPPORTED / NOT IN SOURCE / CONTRADICTED | Citation exists? (leave blank; a lawyer checks in a database). Do not assess whether any citation is real. Then list the answer-key points the output omitted.Associates and partners score separately
Blind the outputs, strip the tool names, give the same set to two panels. Partners judge polish; associates judge whether the draft saves rewrite time, the only efficiency that matters. The committee member’s verdict on one platform was exactly that: “the generations need to consistently save rewrite time. For us, they didn’t.”
Ashurst’s four-expert blind panel scored anonymised outputs from 1 to 5, and one expert showed how a single hallucination behaves in the wild: having spotted an invented verdict, he “dismissed the remaining output and applied critical scoring across the board (each criterion scoring 1 out of 5)”.
Confidentiality and residency review in parallel
Do not leave the vendor questionnaire until a winner is chosen; a tool that fails it is out whatever it scored. The questions come from the CCBE’s 2026 technical guide and the protective-order language in Morgan v. V2X (D. Colo. 2026): a written no-training clause covering inputs, outputs and files; retention and what survives zero data retention; who can read flagged content; subprocessors and regions; deletion at matter end. The vendor due-diligence checklist has the full list.
Two facts change outcomes. A legal wrapper adds a subprocessor rather than removing one; an r/legaltech comparison found Harvey’s and Anthropic’s commercial terms “read about the same” on confidentiality. And residency is uneven: Harvey offers an EU (Frankfurt) region and Legora EU residency, while Claude’s own workspaces store in the US only; European firms should read the EU data residency guide first.
Reference benchmarks: Vals, BigLaw Bench, LEXam
Public numbers tell you what to expect, not what you will get.
| Benchmark | What it found | Caveat |
|---|---|---|
| Vals Legal AI Report, Feb 2025 | Harvey 94.8% on document Q&A v lawyers 70.1%; lawyers 79.7% v Harvey 65.0% on redlining; chronology tied at 80.2% | Lexis+ AI withdrew; four tools tested |
| Vals research study, Oct 2025 | Lawyers 71% accuracy; ChatGPT 80%, Counsel Stack 81%, Alexi 80%, Midpage 79%; legal tools 76% v ChatGPT 70% on authoritativeness | Thomson Reuters, LexisNexis and vLex declined; percentages as reported by LawSites |
| Harvey BigLaw Bench | Claude Opus 4.6 90.2% (Feb 2026), Opus 4.7 90.9% (May 2026); base models ~60% to ~90% since 2024 | Vendor-designed tasks and rubrics |
| LEXam | GPT-5 70.2, Gemini 2.5 Pro 67.4, Claude 3.7 Sonnet 62.9 on 7,537 English and German law-exam questions; EuroLLM-9B 22.95 | Exams, not matters (leaderboard) |
The pattern: tools beat lawyers on document-grounded tasks, lawyers still win on redlining and judgement, and small models are far behind on law. The benchmarks explainer covers what each measures.
Pricing negotiation: seat minimums and renewals
Legal AI pricing is mostly demo-only, so reported ranges matter. One pricing guide puts Harvey at $1,000 to $2,000 per user per month for mid-market firms and $100 to $200 at Am Law scale, with seat minimums of about 25 to 50, twelve-month terms and renewal uplifts of 10 to 25% without a cap, calling its own figures “a planning range, not a quote”. Legora is reported at roughly $3,000 per user per year with a ten-seat floor. The pricing guide and the Harvey, Legora and CoCounsel comparison separate published from reported.
Decision matrix and pilot exit criteria
Weight the columns before seeing results (accuracy 30, hallucinations 25, completeness 15, time 10, usability 10, cost 10, with the confidentiality questionnaire as a pass-or-fail gate), then write the exit criteria on day one:
- Fabricated citations in a filed-work task above the human baseline: no matter work.
- Associate and partner panels more than a point apart on a task: investigate before deciding.
- Fewer than half the pilot group used the tool in week two: usability failed, whatever the scores say.
- No written answers on training, retention and deletion: out.
Be this mechanical because ILTA’s 2025 survey found 71% of Harvey users were still piloting, as were 60% of Copilot and CoCounsel Core users. A pilot without exit criteria is a subscription with a nicer name. GC AI’s four-question tree (full-time in-house? Am Law hourly? live in Word? research-dominant?) is a fair first filter for which category of tool to test; the implementation playbook picks up after the decision, and the shadow AI guide is what happens if the decision takes another year.
Templates
Two more one-page artefacts sit beside the scoring sheet: a log per run (tool, model version, prompt, materials, output, who verified what) and the task instruction below, identical for the lawyer baseline and every tool.
Task [3 of 6]: document Q&A. Using only the attached [share purchase agreement, 140 pages], answer the twenty questions in <questions>…</questions>. For each: quote the operative words, give the page and clause, and state a confidence of High, Medium or Low. If the document does not answer a question, write NOT IN DOCUMENT. Use no other source and cite no authority. Format: Question | Answer | Quote | Page/clause | Confidence. Every human and every tool in this evaluation receives this exact instruction.Where to go next: the firm implementation hub holds the rollout and policy guides that follow a decision, and the self-test and audit prompts sit in the prompt library with their verification steps. In AI Lab for Lawyers the same tasks run side by side in ChatGPT, Claude, Perplexity, NotebookLM and Harvey on screen, the quickest way to learn what a scoring sheet should hold.
Frequently asked questions
How do you evaluate a legal AI tool?
Give the tool and a lawyer identical instructions and identical documents from matters you already know, then score the outputs blind on accuracy, hallucination count, completeness, time and usability. That is the method Vals used in its 2025 benchmark. Run a vendor due-diligence questionnaire on training, retention, human review and residency in parallel, and set the exit criteria before the pilot starts.
How long should a legal AI pilot last?
Two weeks is enough for the scored bake-off if the documents and tasks are prepared in advance; Vals gave its lawyer baseline a two-week window. Harvey's own advice is 60 to 90 days for a full pilot in a single practice area with objectives set in advance, and Axiom recommends 8 to 12 weeks on one use case with measurement from day one. Longer than that without a decision is a subscription, not a pilot.
What should be in a legal AI scoring sheet?
One row per task per tool, with columns for accuracy against a lawyer-prepared answer key, the number of fabricated citations, the number of misgrounded citations (real source, wrong proposition), completeness against a checklist, wall-clock time including verification, a usability score, and the reviewer's initials. Counts beat five-point ratings because they can be audited later; the 2025 Schwarcz study reported hallucinations as counts (3, 4 and 11), not impressions.
Should associates or partners choose the AI tool?
Both should score, separately and blind, and the people who will use the tool daily should weigh more. One r/legaltech commenter reported that all the associates hated the platform the partners bought because the partners 'think it's magic'. Partners judge polish; associates judge whether the output saves rewrite time. If the two panels disagree by more than a point on any task, treat it as a finding, not noise.
Which benchmarks compare legal AI tools?
Vals' February 2025 report is the only public head-to-head of the large vendors (Harvey 94.8% on document Q&A against a 70.1% lawyer baseline; lawyers beat every tool on redlining). Vals' October 2025 research study, Harvey's BigLaw Bench, the LEXam law-exam benchmark and Stanford's RegLab hallucination study are useful reference points, but none of them tested your documents, and several major vendors declined to take part.