Steven Schwartz did check. Before he filed the brief that became Mata v. Avianca, he asked ChatGPT whether Varghese v. China Southern Airlines was a real case, and the model told him it could “be found in reputable legal databases such as LexisNexis and Westlaw”. Three years later, an Illinois lawyer with a “premier corporate subscription of ChatGPT” filed ten false citations and quotations and was fined $1,500 for each. Between those two filings sit roughly 2,000 court decisions, and the recurring mistake is not failing to verify AI legal citations but verifying them the wrong way: confirming that something exists and stopping there.
To verify AI legal citations you need six separate checks, in a fixed order. This is the protocol, with the minutes each step takes, the sanction each error type has already produced, and the certificate you can sign at the end.
Why “the case exists” is not verification
Damien Charlotin’s database of court decisions involving hallucinated material stood at 2,039 on 12 September 2026: 811 involving lawyers, 1,173 self-represented litigants, 32 judges. Look at the nature of the errors rather than the headline. Charlotin tags 1,689 entries as fabricated, but 853 as misrepresented and 549 as false quotes, with overlap. A large share of the profession’s failures involve cases that exist.
Stanford’s pre-registered benchmark of the paid research tools made the same point with numbers. Lexis+ AI and Ask Practical Law AI gave incorrect or misgrounded answers on more than 17% of queries, Westlaw AI-Assisted Research on more than 34%. The paper’s lasting contribution is the word misgrounded: a genuine source cited for a proposition it does not support. LeanLaw’s verification checklist calls that more dangerous than fabricating outright, because a checker who stops at existence waves it through.
A poster on r/LawSchool put it more bluntly than any judge: “Westlaw 100% hallucinates case holdings. It just may not hallucinate case names.”
The Ninth Circuit turned the point into a standard in Lnu v. Blanche:
a competent and diligent attorney must do more than prompt generative AI, check that the citations provided by the AI are real and the subject matter roughly on point, and call it a day. … A competent and diligent attorney must also read and reason. — Ninth Circuit, Lnu v. Blanche, No. 24-4790 (3 June 2026)
Two details show what “real” can hide. In Ayinde v Haringey one of the five invented English cases carried the neutral citation [2020] EWHC 2435 (Admin); the citation exists, but it belongs to an unrelated business-rates matter. In Concord Music v. Anthropic a Latham & Watkins associate asked Claude to format a citation to a real article; the link came back right, the author and title wrong, and the firm’s “manual citation check did not catch that error”. Both would have passed a link check.
The six error types, each with a real sanction
Sort every AI citation error into one of six bins, because each bin needs a different check and a different tool.
| Error type | What it looks like | A court that saw it | Consequence |
|---|---|---|---|
| 1. Fabricated | The case does not exist | Mata v. Avianca (S.D.N.Y. 2023): six invented cases | $5,000 jointly and severally on two lawyers and the firm |
| 2. Misquoted | Real case, invented words in quotation marks | Noland v. Land of the Free (Cal. Ct. App. 2025): 21 of 23 quotations fabricated | $10,000, State Bar referral |
| 3. Misdescribed | Real case, wrong holding or posture | Scott v. Illinois Human Rights Commission (2026): five real cases that do not say what was claimed | $1,500 per false citation or quotation, ARDC referral |
| 4. Superseded | Real case, overruled or limited, presented as current | Stanford’s test: Lexis+ AI applied Casey after Dobbs | 33 “outdated advice” entries in Charlotin; the citator is the cure |
| 5. Wrong citation details | Right case, wrong reporter, file number, author or pin cite | Concord v. Anthropic (2025); KG Berlin 17 WF 144/25 (2025): real FamRZ page range, wrong BGH file number | Declaration portion struck; formal Ermahnung |
| 6. Unsupported proposition | Real case, real words, does not prove your point | Lnu v. Blanche (9th Cir. 2026): “roughly on point” is not enough | $2,500 each, six-month suspension, two years of sworn disclosures |
The bins matter because tools cover them unevenly. A link check catches bin 1. A citator catches bin 4. Nothing but a lawyer reading the case catches bins 3 and 6, and that is where the Stanford numbers live: 61% of Westlaw’s errors were reasoning errors made on the correct documents.
The six-step protocol to verify AI legal citations
The order is deliberate: existence first, because there is no point reading a case that does not exist; documentation last, because the record is what a court or insurer will ask for. Time estimates are LeanLaw’s, split across my six steps (LeanLaw folds the pin cite into the holding check), and they match my experience.
Step 1: existence in a primary database (30 seconds to 2 minutes)
Open the case where it lives: Westlaw, Lexis, CourtListener, BAILII, the National Archives, RIS, juris or beck-online. The Divisional Court in Ayinde listed the acceptable sources for English law: the government’s legislation database, the National Archives case-law database, the official Law Reports and “the databases of reputable legal publishers”. A blog summary does not count. If you cannot open the decision itself, the citation does not go in. Judge Wingate of the Southern District of Mississippi, after his own chambers docketed an order naming non-existent parties, adopted a rule that every cited case is printed from Westlaw and attached to the draft. Crude, and it works.
Start by getting the list out of the document without letting the model do anything else.
List every case, statute, rule, regulation and secondary source cited in the document below in a table: Citation exactly as written | Type (case / statute / rule / secondary) | Proposition it is cited for (quote the sentence) | Pinpoint given? (yes/no) | Direct quotation? (yes/no).
Do not verify, correct or comment on whether any citation exists. I will check each row in a primary database myself. List repeated authorities once per occurrence.
<document>
[paste]
</document>Red flags at this stage, from LeanLaw’s and BriefCatch’s lists: a citation that matches your argument too perfectly; an overly descriptive case name; a reporter volume outside the range that reporter reached; a page number that does not exist in the volume; a docket number in a format the court never used; a court that did not exist on the date given; a judge you have never heard of. Ask Lexis+ AI about the fictitious “Judge Luther A. Wilgarten” and, in Stanford’s test, it returned a real case.
Step 2: quotation against the reporter (2 to 5 minutes)
Every string inside quotation marks is checked character for character against the pinpointed page. The Illinois Appellate Court in Scott held that cross-referencing against a platform is insufficient if the attorney does not confirm that the cited text actually appears, and fined four false statutory quotations alongside the case-law errors. In LG Frankfurt’s 2-13 S 56/24 (25 September 2025) three verbatim “quotes” from BGH decisions were, in the court’s words, “komplette Fälschungen”, and the file numbers alone should have raised suspicion because the BGH does not hear the type of appeal they described. A quotation that reads beautifully is a reason to check, not a reason to relax.
Step 3: holding and proposition (3 to 5 minutes)
Read the case: all of it if short, the facts, the question and the operative paragraphs if long. Then ask one question: does this authority hold what my sentence says, in the posture I need?
This is the step the paid platforms fail. Stanford found the tools “struggle with elementary legal comprehension: describing the holding of a case, distinguishing between legal actors, and respecting the hierarchy of legal authority”; Westlaw described one holding as the “opposite of” the actual opinion. In Johnson v. Dunn one of the five bad Butler Snow citations, Wilson v. Jackson, was a real case cited for the wrong proposition; the three lawyers were disqualified anyway.
A model can allocate your reading time, as long as it is not the checker.
Review the draft section below. For every sentence asserting a legal proposition, label it SUPPORTED (name the authority in the draft and quote the words that do the work), OVERSTATED (say how the draft goes further than the authority), UNSUPPORTED (no authority cited) or FACTUAL CLAIM NEEDING RECORD CITE.
Do not add any authority of your own and do not tell me whether a citation exists; that is my job. End with the three propositions you would read the source most carefully for, and why.
<draft>
[paste]
</draft>Treat OVERSTATED and UNSUPPORTED as your reading order. Treat SUPPORTED as unverified until you have read the passage yourself.
Step 4: subsequent history and status (1 to 3 minutes)
Run the citator: KeyCite, Shepard’s, BCite, or the equivalent in juris or RIS. A citation that passes steps 1 to 3 can still be dead law. Stanford’s clearest example was Lexis+ AI applying the undue-burden standard from Casey after Dobbs overruled it; models report what they were trained on as current. In Prososki v. Regan the Nebraska Supreme Court found “real cases with fake quotations, real cases with mischaracterized holdings” alongside invented ones, struck the brief and referred counsel to the Counsel for Discipline.
Step 5: pin cite and page accuracy (done during the step 3 read)
The pin cite has to land on the passage that does the work. The KG Berlin’s reconstruction in 17 WF 144/25 shows how a plausible German citation is assembled: the fake “BGH, Beschl. v. 14.11.2007 – XII ZB 183/07, FamRZ 2008, 137” combined a genuine journal page range (FamRZ 2008, 134–136 holds a real BGH maintenance decision) with a plausible Senate designation and a file number that belongs to no decision at all. The court called its checking “aufwändige Prüfung” and issued a formal Ermahnung.
Step 6: jurisdiction and precedential weight (3 to 5 minutes)
Binding or persuasive? Right court, right circuit, right country, right era? Thirty-eight per cent of Lexis+ AI’s errors in the Stanford study were “inapplicable authority”: wrong jurisdiction, wrong period, overruled. Westlaw once claimed a Nebraska Supreme Court decision had reversed the United States Supreme Court on federal law. In the earlier “Large Legal Fictions” study GPT-4 misidentified the court in 83.1% of district-court cases. An English tell in Ayinde was Americanised spelling (“emphasized”) in a domestic filing; models import the jurisdiction with the most training data.
Then document. LeanLaw’s sixth step is the record: who checked which citation, in which database, on what date, and which citations were AI-generated in the first place. An Austrian practice guide written after the OGH’s 14 Os 95/25i decision recommends a “Kurzvermerk zur Prüfung und zum RIS Gegencheck der Zitate” in the file; Legal AI Governance’s renewal checklist puts a “pre-filing verification log” among the artefacts to have ready for your insurer.
Summarise this session as a verification log entry: Date and time; Tool and model used; Matter reference (anonymised); Document checked; every authority the document cites, each marked "verified at source by [initials] on [date] in [database]" or "NOT YET VERIFIED"; authorities removed and why; quotations corrected and how; open items. Plain text, no commentary, ready to paste into the file.Time budget: how long this takes and how to bill it
Add the steps up and a citation you actually rely on costs seven to fifteen minutes of a lawyer’s attention; a supporting “see also” costs less because steps 2 and 5 fall away. A twenty-citation brief is two to five hours of verification even when a checker handles existence, which is why the honest Reddit complaint is that “double checking Claude or Chat GPT’s output to make sure it’s true takes almost as long as me opening Westlaw and doing it myself”. Sometimes it does. That is the price of using the tool for authority at all, and the reason to use AI for issue-spotting and a grounded database for cases, as set out in the research-verify-cite workflow.
| Step | LeanLaw estimate | Who | Tool |
|---|---|---|---|
| 1. Existence | 30 s – 2 min | Paralegal or checker; lawyer confirms | Database, citation checker |
| 2. Quotation, 3. Holding and 5. Pin cite | 2 – 5 min | Lawyer | Reporter text |
| 4. Status | 1 – 3 min | Paralegal or lawyer | KeyCite / Shepard’s / BCite |
| 6. Jurisdiction and log | 3 – 5 min | Lawyer | File note |
Billing answers itself once verification is seen as work rather than overhead. ABA Formal Opinion 512 says a lawyer may charge for inputting information and “for the time necessary to review the resulting draft”, but “a fee charged for which little or no work was performed is an unreasonable fee”. California’s 2026 guidance lists “reviewing and editing generative AI outputs” as chargeable actual time; North Carolina’s worked example cuts the other way, so if AI turned three hours of drafting into one, you bill one. A time entry reading “verified 14 authorities in Westlaw, read pinpoints, corrected two quotations” is defensible to a client, to a court asking about your process and to an insurer at renewal.
Tool-assisted verification: what the Verify buttons check and what they miss
Shepard’s Verify Trust Markers, added to Lexis+ with Protégé in May 2026, flag citations that cannot be verified as existing in Lexis content. That is step 1. Thomson Reuters’ Deep Research Verify, released in June 2026 and carried into the next-generation CoCounsel Legal of 20 August 2026, goes further by the vendor’s description: it highlights supporting passages and flags misattributions, checking that cited authority “actually supports each legal assertion”. That is a vendor’s claim about steps 3 and 6 for a feature a few months old; treat it as a reading aid until an independent benchmark says otherwise.
The standalone tools in Nicole Black’s ABA Journal roundup of June 2026 are mostly step-1 and step-5 machines: Clearbrief’s Cite Check Report and GroundTruth as Word add-ins, Benchly’s rules-based check inside ezBriefs, CaseRead (free up to 50,000 characters), LawDroid’s CiteCheck AI (free for up to five documents, cross-referenced against CourtListener), CiteSentinel at $19.99 per document. BriefCatch RealityCheck pairs a deterministic database check with an AI pass for misquotes and mischaracterised holdings, an attempt at steps 2 and 3. Tool by tool, they are compared in AI citation checkers compared.
| Layer | What confirms it | What it cannot tell you |
|---|---|---|
| Existence and format | Citator, Shepard’s Verify, any checker | Whether the words or holding are right |
| Quotation | Reporter text; RealityCheck as a first pass | Whether the words support your point |
| Holding and proposition | A lawyer reading the case; Deep Research Verify as a pointer (vendor claim) | Nothing replaces the read |
| Status | KeyCite, Shepard’s, BCite, juris, RIS | Whether the case was ever on point |
| Jurisdiction and weight | The lawyer | — |
The self-test: five citations, three fakes
Before you trust any tool for research, find out how it fails. The test takes fifteen minutes; repeat it whenever the model changes.
Take two citations you know are real and on point, from a brief you wrote yourself. Build three fakes by hand, one for each bin a link check misses: a real case name with the wrong reporter volume and year; a plausible case name that does not exist, in the court’s format; and a real case cited for a proposition it does not support, ideally one that has been overruled. Ask the tool to confirm that each of the five exists and supports the stated point. A tool that only checks existence will bless the third fake. A tool that flatters you will bless all three. Record the result by tool, model and date.
Add Stanford’s three probes, which have already caught commercial products. Ask why Justice Ginsburg dissented in Obergefell (she did not; Ask Practical Law AI agreed with the premise). Ask about the rulings of “Judge Luther A. Wilgarten” (a fiction; Lexis+ AI returned a real case). Ask for the current test for abortion restrictions and see whether Casey is presented as good law.
Design a self-test I can run on [tool] to probe how it handles legal citations. I will supply two real citations I have verified <real>[...]</real> and three I have deliberately constructed <constructed>[one real case name with a wrong reporter volume and year; one invented case name in the correct format for this court; one real case cited for a proposition it does not support]</constructed>.
Write: (1) the exact prompt asking the tool to confirm that each of the five exists and supports its stated proposition; (2) the correct behaviour for each citation; (3) the failure behaviour to watch for; (4) a results table with tool, model version and date. Do not evaluate the citations yourself and do not add citations of your own.Two cautions. A tool can pass a test and still fail under pressure: Adam Unikowsky found Claude 3 Opus produced a “10/10” disposition of a Supreme Court case from the briefs alone, yet in a later voice experiment recreating one of his own arguments the AI still “fabricated a response” when pressed on a flawed question. Charlotin’s explanation: “The harder your legal argument is to make, the more the model will tend to hallucinate, because they will try to please you.” And the people who most need the test are the least likely to run it. Habeas.ai reports research in which users with greater AI literacy showed more overconfidence about verification, not less. If you are reading a 3,500-word page on citation verification, that is you.
A certification template for standing orders
Courts have converged, from different directions, on the same three questions. Judge Brantley Starr’s standing order in the Northern District of Texas, entered on 30 May 2023 while Mata was still pending, requires a certificate that no portion of a filing was drafted by generative AI or that any AI-drafted language “was checked for accuracy, using print reporters or traditional legal databases, by a human being”. The Upper Tribunal’s judicial review claim form, since UK and R (Munir) v SSHD [2026] UKUT 81 (IAC), requires a statement of truth that every cited authority “(a) exists; (b) may be located using the citation provided; and © supports the proposition of law for which it is cited”. The NSW Supreme Court’s Practice Note SC Gen 23 asks authors of AI-assisted submissions to verify that citations “(a) exist, (b) are accurate, and © are relevant”.
Those are steps 1, 2, 3 and 6, written as a promise. The Ninth Circuit in Lnu v. Blanche went further for one firm: every lawyer there must, for two years, include in every filing a statement under penalty of perjury disclosing whether generative AI was used and certifying personal review of every citation. The Fifth Circuit, which declined a circuit-wide rule in June 2024, explained why it did not need one: “‘I used AI’ will not be an excuse for an otherwise sanctionable offense.” The Legal AI Governance tracker counted 113 active court orders binding attorney filings in spring 2026, 82 requiring verification or disclosure; check your judge’s standing order the day you file.
A certificate that satisfies all of these at once looks like this. Fill it in from the step-6 log, not from memory.
CERTIFICATION OF CITATION VERIFICATION
I, [name], counsel for [party], certify that:
1. [No portion of this filing was drafted by generative artificial intelligence.] / [Generative artificial intelligence ([tool and version]) was used to assist with [e.g. summarising the record, drafting headings, editing for clarity]. No authority in this filing was supplied by an AI tool without independent verification.]
2. Every case, statute, rule and secondary source cited in this filing (a) exists; (b) can be located using the citation provided; and (c) supports the proposition for which it is cited.
3. Each citation, quotation and pinpoint was checked by a human being against [print reporters / Westlaw / Lexis / CourtListener / BAILII / the National Archives / RIS / juris] on [date(s)], and the subsequent history of each authority was checked in [KeyCite / Shepard's / BCite] on [date].
4. A verification log recording each authority, the database used, the reviewer and the date is retained in the file and available to the Court on request.
Signed: ____________________ Date: __________
If you cannot sign paragraph 2 for every citation, you have not finished verifying; the certificate tests the protocol rather than decorating it. And if you discover after filing that you could not have signed it, the sequence in what to do when you find a fake citation matters more than the error: the Fifth Circuit said in Fletcher v. Experian that had counsel “accepted responsibility and been more forthcoming, it is likely that the court would have imposed lesser sanctions”.
The cases behind each error type are collected in the sanctions timeline and, for readers outside the United States, in the European decisions; why the errors happen at all is explained in why AI makes up fake cases, and the same six checks apply when the fakes arrive from the other side, as in responding to AI-generated pro se filings. More verification prompts are in the prompt library, and the rest of this cluster is at /verification/. The sanctioned lawyers in Charlotin’s database share one failure, trusting an output they had not read; in AI Lab for Lawyers you build the verification habit on your own documents, tool by tool, until the six steps are reflex rather than checklist.
Frequently asked questions
How do I check whether an AI citation is real?
Open it in a primary database: Westlaw, Lexis, CourtListener, BAILII, the National Archives, RIS, juris or beck-online. If the citation does not resolve to a document with matching party names, court, year and reporter, it does not go in the filing. Then keep going: existence is only step one of six. Never ask the AI that produced the citation whether it is real; in Mata v. Avianca ChatGPT assured the lawyer the cases could be found in LexisNexis and Westlaw.
What does Shepard's Verify actually check?
Shepard's Verify Trust Markers, added to Lexis+ with Protégé in May 2026, flag citations that cannot be verified as existing in Lexis content. They confirm existence, not that the authority supports the proposition you cite it for, and not that the quotation appears at the pinpoint. You still need to read the passage, run the citator for subsequent history and check jurisdiction and posture yourself.
How long does citation verification take?
LeanLaw's checklist budgets 30 seconds to 2 minutes for existence, 2 to 5 minutes to read the holding and confirm the quote, 1 to 3 minutes for KeyCite or Shepard's, and 3 to 5 minutes for applicability, so roughly seven to fifteen minutes per citation you rely on. A brief with twenty citations is two to five hours of work when a checker handles existence and you do the reading.
Can I rely on Westlaw AI citations without checking?
No. Stanford's pre-registered benchmark found Westlaw AI-Assisted Research producing incorrect or misgrounded answers on more than 34% of queries and Lexis+ AI on more than 17%, and the Ninth Circuit relied on that study in its June 2026 sanctions order in Lnu v. Blanche. The Illinois Appellate Court added in July 2026 that paying for 'premier' or 'corporate' versions of AI products 'does not negate an attorney's obligation to verify all citations of authority'.
What should a verification certificate say?
Model it on the three tests the Upper Tribunal now requires in judicial review claim forms: that every cited authority exists, can be located using the citation given, and supports the proposition for which it is cited. Add Judge Starr's element that the check was done by a human in print reporters or a traditional database, name the person and the date, and keep the underlying log so you can prove it if asked.
What are the six types of AI citation error?
Fabricated (the case does not exist), misquoted (real case, invented words), misdescribed (real case, wrong holding), superseded (real case, overruled or limited), wrong citation details (real case, wrong reporter, file number, author or pin cite) and unsupported proposition (real case, real words, but they do not prove your point). Stanford calls the last one 'misgrounded'; LeanLaw's checklist warns it is more dangerous than fabrication because a checker stops at existence.