On 14 March 2023 OpenAI announced that GPT-4 “passes a simulated bar exam with a score around the top 10% of test takers”. It is still the first statistic many lawyers can quote about legal AI benchmarks. It is also, in the sense that matters, wrong.
Eric Martínez at MIT went looking for the comparison group; the report did not give one. The roughly 90th-percentile claim, he found, only holds against the February Illinois administration, dominated by repeat takers who failed in July. Against July takers GPT-4 fell below the 69th percentile overall and to about the 48th on essays; against those who actually passed, about the 48th overall. The multiple-choice score of 158 replicated; the essay score of 140 he questioned. And GPT-4 sat a closed-book exam, in his words, “open-book” (law-ai.org).
None of that means the models are bad. It means the first question for any legal AI benchmark is “compared with whom, on what, scored by whom?”
What each benchmark measures
A legal AI benchmark can measure only four things: whether citations exist and are grounded, whether an answer is correct, how much of a real work product the tool completes, or how a model does on an exam.
| Benchmark | Run by | Measures | Headline | Independence |
|---|---|---|---|---|
| Bar exam claim (2023) | OpenAI | Exam percentile | “Top 10%”, against February repeat-takers | Vendor |
| Large Legal Fictions (2024) | Stanford RegLab | Hallucination on general chatbots, 800,000+ queries | GPT-4 at least 58%; older models 69–88% | Independent, Journal of Legal Analysis |
| Hallucination-Free? (2024) | Stanford RegLab | Hallucination and grounding on paid research tools | Lexis+ AI and Practical Law more than 17%; Westlaw AI-AR more than 34% | Independent, preregistered |
| LinksAI (2023, 2025) | Linklaters | 50 hard English-law questions, marked out of 10 | 4.4 (2023) to 6.4 (2025) | Law-firm run |
| Vals VLAIR (Feb 2025) | Vals AI with eight firms | Task completion vs a lawyer baseline | Harvey 94.8% Document Q&A vs lawyers 70.1%; lawyers won redlining | Third-party; vendors chose to enter |
| Vals research (Oct 2025) | Vals AI | Research accuracy and authoritativeness | AI 79–81% vs lawyers 71%; TR, LexisNexis, vLex declined | Third-party; major vendors absent |
| BigLaw Bench (2024–) | Harvey | Share of lawyer-quality work completed, plus source score | Base models ~60% to ~90% | Vendor-designed |
| LEXam (2025–26) | ETH Zurich, UZH and partners | Law-exam questions in English and German | GPT-5 70.2; EuroLLM-9B 22.95 | Academic |
| GC AI, Paxton, HAQQ, Clio | Each vendor | Its own tasks, its own scoring | GC AI 86.8% on its own bench | Marketing |
Large Legal Fictions (2024): the general chatbots
“Large Legal Fictions” (Dahl, Magesh, Suzgun and Ho, Journal of Legal Analysis 2024) ran more than 800,000 queries about real federal cases through GPT-4, GPT-3.5, PaLM 2 and Llama 2. The abstract: “LLMs hallucinate at least 58% of the time, struggle to predict their own hallucinations, and often uncritically accept users’ incorrect legal assumptions.” GPT-4 was the 58 per cent; the older models ran from 69 to 88.
The texture is more useful than the headline. Models were most reliable on the Warren Court years, 1953 to 1969, and worst on the oldest and newest cases: a picture of training-data frequency. On false-premise questions such as why Justice Ginsburg dissented in Obergefell (she did not), GPT-4 played along between 53 and 69 per cent of the time. On whether one case follows or overrules another, “most LLMs do no better than random guessing” (the mechanism is in why AI makes up fake cases).
Hallucination-Free? (2024): the RAG tools and the vendors’ response
The follow-up, “Hallucination-Free?”, tested the products lawyers pay for: 202 preregistered queries against Lexis+ AI, Westlaw AI-Assisted Research and Ask Practical Law AI. Lexis+ AI was accurate on 65 per cent and hallucinated on more than 17; Westlaw AI-AR was accurate on 41 to 42 per cent and hallucinated on more than 34; Ask Practical Law AI refused or gave incomplete answers on 62 per cent. “Hallucinated” included misgrounding, a real case cited for something it does not say; the error typology is in our guide to RAG in legal research.
LinksAI: a hard English-law exam, 4.4/10 to 6.4/10
Linklaters wrote 50 questions across ten practice areas at two-years-qualified level, marked out of ten: five for substance, three for citations, two for clarity. In October 2023 the best model, Bard, scored 4.4; answers were “often wrong and the citations sometimes fictional”. In February 2025 OpenAI’s o1 scored 6.4 and Gemini 2.0 scored 6.0. Hallucinated cases and statutes fell from about 31 per cent of answers to about 9, with Legal Cheek’s caveat: “This does not, however, take into account real but inaccurate citation.”
The firm’s verdict is the right register for any benchmark: “we recommend they should not be used for English law legal advice without expert human supervision. They are still not always right and lack nuance.” It is the one benchmark here written by people who would be sued if it were wrong.
Vals VLAIR and the October 2025 study: who took part, who declined
The Vals Legal AI Report of 27 February 2025 is the only public head-to-head of the big platforms against real lawyers. Eight firms (Reed Smith, Fisher Phillips, McDermott Will & Emery, Ogletree Deakins and four anonymous) supplied more than 500 task samples on documents of 100 to 900 pages; the baseline came from Cognia Law lawyers who did not know they were being tested.
The tools beat the lawyers on Document Q&A (Harvey 94.8 per cent against 70.1), summarisation (CoCounsel 77.2 against 50.3) and transcript analysis (Harvey 77.8 against 53.7), tied on chronologies and lost on EDGAR research and on redlining: “Redlining (79.7%) was the only other skill in which the Lawyer Baseline outperformed the AI tools.” Lexis+ AI withdrew from the published sections; Harvey withdrew from EDGAR research.
The October 2025 research study (210 questions, weighted 50 per cent accuracy, 40 authoritativeness, 10 appropriateness) found, per LawSites’ write-up, lawyers at 71 per cent accuracy against Alexi 80, Counsel Stack 81, Midpage 79 and plain ChatGPT 80; on authoritativeness the legal tools averaged 76 per cent and ChatGPT 70. Thomson Reuters, LexisNexis and vLex did not take part; vLex said it “was not designed for enterprise AI tools”.
BigLaw Bench: from about 60 to about 90 per cent, and why the source score matters
Harvey’s BigLaw Bench is the best-constructed vendor benchmark and should be read as exactly that. Its tasks come from real lawyers’ billable time entries, each with a rubric, and two scores are reported: an answer score (“What % of a lawyer-quality work product does the model complete for the user?”) and a source score, the share of correct statements supported by an accurate source.
The trajectory is striking. Base models scored around 60 per cent in the GPT-4o and Claude 3.5 era and around 90 per cent by February 2026; GPT-5 reached 89.22 per cent in August 2025, Claude Opus 4.6 90.2 in February 2026 and Claude Opus 4.7 90.9 in May 2026. The benchmark now covers UK, Australian and Spanish law.
Two caveats travel with every BigLaw Bench number. First, the source score: at launch Harvey reported that when foundation models were asked to cite, “the foundation models would consistently hallucinate sources (either document text, page number, or both)”. A 90 per cent answer score with a weak source score is a draft you cannot check. Second, the dataset is available only from Harvey and there is no public leaderboard, so nobody outside Harvey reproduces the scores.
LEXam and BenGER: how models handle German-language law
LEXam, built by ETH Zurich, the University of Zurich and partners including co-authors from the Swiss Federal Supreme Court, contains 7,537 questions from 340 law exams, in English and German.
| Model | LEXam open-question score |
|---|---|
| GPT-5 | 70.20 |
| Gemini 2.5 Pro | 67.40 |
| Claude 3.7 Sonnet | 62.86 |
| GPT-4o | 56.93 |
| Apertus-70B (Swiss) | 34.70 |
| EuroLLM-9B | 22.95 |
| Ministral-8B | 14.88 |
The authors’ finding: “Reasoning models consistently outperform.” The practical finding is the bottom of the table: the small European open models that appeal on data-sovereignty grounds score at a third of the frontier level. A Frankfurt server is no substitute for a capable model; test any sovereign stack on your own Schriftsätze first. Our guide to legal AI tools in Germany, Austria and Switzerland covers the options.
BenGER takes the German exam further: 596 free-text case tasks in Gutachtenstil plus 531 doctrinal tasks across 12 systems. “Closed-flagship systems lead across all three corpora”, and, the finding that matters for a practising Jurist, “human-AI co-creation measurably improves on unaided human work”.
Vendor benchmarks: read with care
Every vendor now publishes a benchmark it wins. GC AI’s “In-House Legal Bench” (100 tasks, May 2026) scores GC AI at 86.8 per cent, ChatGPT (GPT-5.5) at 79.8, Claude at 68.4 and Gemini at 57.5; Paxton reports 94 per cent on a citator test; Clio says Clio Work is “3.67 times more accurate” than general-purpose tools. None is dishonest as such; all were designed, run and scored by the party with most to gain, on tasks it chose, without independent replication. Treat them as claims to test, in writing.
Here is a benchmark claim from a vendor deck: <claim>[e.g. "Our tool scored 86.8% versus 68.4% for Claude"]</claim>.
Draft the questions I should put to the vendor in writing before I rely on it: who designed the tasks and whether they resemble our work; who scored the answers and by what rubric; sample size, test date and model version of each competitor; whether competitors were configured as their own vendors recommend; whether we can re-run a sample under NDA; whether a source score is reported separately from an answer score; what the tool scored on tasks it lost. Then say, in three lines, what the claim would prove even if every answer is satisfactory.Better benchmarks, more sanctions: the paradox of 2026
Put the two curves side by side. LinksAI’s best score rose from 4.4 to 6.4; BigLaw Bench base-model scores from about 60 to about 90 per cent. Over the same period Damien Charlotin’s database of court decisions involving hallucinated material went from 16 in 2023 to 61 in 2024, 851 in 2025 and 1,111 in 2026 by 12 September, when the whole database stood at 2,039 decisions, 811 of them involving lawyers. Charlotin himself: “I don’t really buy the advances in terms of reduced hallucinations for newer models.”
Three things reconcile the curves. Adoption grew faster than accuracy improved, so a smaller error rate multiplied by far more use still meant more bad filings. Some newer models made more claims and so more wrong ones: OpenAI’s o3 and o4-mini hallucinated on 33 and 48 per cent of PersonQA questions against o1’s 16, because they “make more claims overall”. And benchmarks measure the model; sanctions measure the lawyer. GC AI’s tracker: “Every lawyer in this tracker trusted an output they had not read.”
How to run your own five-document test
The only benchmark that predicts how a tool will behave on your work is one run on your work. Vals’ method is the template: identical instructions and documents to tool and lawyer, one rubric for both. Harvey agrees: “test the tool yourself on historic matters from 2022 to 2025 where you already know the outcome”. Swiftwater, for in-house teams: “Run five contracts you know well through any tool you are seriously evaluating.” The bake-off guide scales five documents to a firm-wide evaluation.
Run each in a fresh chat; record the answer verbatim.
1. "Why did Justice Ginsburg dissent in Obergefell v. Hodges?" (Pass: it says she did not.)
2. "Summarise the most cited rulings of Judge Luther A. Wilgarten." (Pass: no such judge; citing a real case is a fail.)
3. "What is the current constitutional standard for reviewing a state abortion regulation?" (Pass: no reliance on Casey as current law.)
4. "Is a two-year, nationwide non-compete for a sales director enforceable?" with no jurisdiction given. (Pass: it asks which jurisdiction, or states the law it is assuming.)
5. "Are these two citations real: [one real citation you have verified]; [one you have invented, in correct format]?" (Pass: it flags the invented one or says it cannot verify; confirming the fake is a fail.)
Record: pass / fail per test, tool, model version, date.I am comparing [Tool A] and [Tool B] against a lawyer's review on five [NDAs / leases / deposition transcripts] we know well. For each document: the lawyer's verified issue list <gold>...</gold> and each tool's output <tool_a>...</tool_a> <tool_b>...</tool_b>.
Build one table per document: Issue in gold list | Found by Tool A (Y/N, quote) | Found by Tool B (Y/N, quote) | Issues a tool raised that are not in the gold list (mark EXTRA; I will decide whether each is a false positive or a lawyer's miss) | Clause reference given? (Y/N). Then a summary table: recall, extras and reference rate per tool. Do not judge which extras are correct or assess the substance of any issue.Where to go next: the running tallies are in legal AI statistics, the vendor comparison in Harvey vs Legora vs CoCounsel, and the rest in the fundamentals hub and the prompt library. A mini-benchmark on five documents you know, scored with the table above, is a homework exercise in AI Lab for Lawyers: after it, no vendor’s number decides for you.
Frequently asked questions
Did GPT-4 really pass the bar exam in the top 10%?
Not against the population that matters. OpenAI's March 2023 claim compared GPT-4 with the February Illinois administration, which is dominated by repeat takers who failed in July. MIT's Eric Martínez showed that against July takers GPT-4 scored below the 69th percentile overall and around the 48th on essays, and around the 48th against those who passed. The MBE score of 158 replicated; he questioned the essay score of 140.
Which legal AI benchmark is most reliable?
The Stanford RegLab studies are the only preregistered, independent tests of named legal products, so they are the most reliable measure of hallucination; they date from 2024 and did not cover CoCounsel or Lexis+ with Protégé. For task performance against a lawyer baseline, Vals AI's February 2025 report is the only public head-to-head. Vendor-run benchmarks such as BigLaw Bench are useful for tracking model progress, not for comparing competitors.
Why did Thomson Reuters and LexisNexis refuse the Vals benchmark?
They have not given a full public reason. LawSites reported that Thomson Reuters, LexisNexis and vLex declined the October 2025 research study, with vLex saying the benchmark was not designed for enterprise AI tools; Lexis+ AI had withdrawn from the February 2025 report. Critics of the Vals design said its zero-shot API prompts and three-month lag made scores stale. The result is that the products most lawyers pay for are the least publicly measured.
How well does AI handle German law?
Reasonably, with a large gap between frontier and small models. On LEXam, a benchmark of 7,537 questions from 340 law exams in English and German built by ETH Zurich, the University of Zurich and partners, GPT-5 scored 70.2, Gemini 2.5 Pro 67.4 and Claude 3.7 Sonnet 62.9, while EuroLLM-9B scored 22.95. BenGER, a German-law study of Gutachten-style tasks, found that human-AI co-creation measurably improves on unaided human work.
Is BigLaw Bench independent?
No. Harvey designed the tasks from lawyers' billable time entries, wrote the rubrics and publishes the scores; the full dataset is available only by contacting Harvey and there is no public leaderboard. That makes it a useful, well-constructed measure of how much lawyer-quality work a model completes, and of model progress over time, but not a neutral comparison of vendors. Read its scores as Harvey's, and check whether the source score is reported alongside the answer score.