Steven Schwartz’s firm did not have Westlaw or Lexis. Levidow, Levidow & Oberman had a limited Fastcase plan, so when Schwartz needed authority against Avianca he turned to ChatGPT for AI legal research, and when Judge Castel demanded the six opinions he asked ChatGPT for those too. He even asked whether Varghese was a real case and was told it “can be found in reputable legal databases such as LexisNexis and Westlaw”. That was June 2023.

By 12 September 2026, Damien Charlotin’s database of court decisions involving hallucinated material stood at 2,039 entries, 811 involving lawyers. The models got dramatically better in those three years; the failure mode did not change. Someone treated a language model as a search engine that returns answers instead of documents, and filed the answers.

The fix is to notice that “research” is three different jobs, that AI is only safe for two, and that the third, verification, must become a habit that survives a deadline.

Three kinds of research, and why AI is only safe for two

“Research this” means one of three things: orientation (what are the elements, where do I look), authority (which cases decide it) or certainty (is this citation real, current, and does it say what my sentence says).

Mode Safe tools What goes wrong Time
1. Issue-spotting Any frontier model in a no-training tier Wrong jurisdiction imported; a false premise in your question accepted Minutes
2. Finding authority Grounded platforms (Lexis+ with Protégé, CoCounsel Legal, vLex Vincent); web-cited tools pointed at official databases Fabricated citations and, worse, “misgrounded” ones Hours, plus verification
3. Verification You, a primary database and a citator Delegating it to the model that produced the citation Ten minutes per cite

Stanford’s pre-registered benchmark found Lexis+ AI and Ask Practical Law AI producing incorrect or misgrounded answers on more than 17% of queries, and Westlaw AI-Assisted Research on more than 34%. Misgrounded is the paper’s lasting contribution: a genuine case cited for a proposition it does not support, which is more dangerous than an invented case because a checker who stops at “does it exist” waves it through.

In the Vals legal research benchmark of October 2025, as reported by LawSites, lawyers scored 71% while Counsel Stack scored 81%, Alexi 80%, Midpage 79% and plain ChatGPT 80%. But ChatGPT scored 70% on authoritativeness against 76% for the legal tools, and Vals concluded that “access to proprietary databases, even if composed mainly of publicly available data, does result in differentiated products”. Accuracy is what you want in a memo. Authoritativeness is what you need in a filing.

Mode 1: issue-spotting with any frontier model

This is the safest use of AI legal research, because the output contains no citations: elements, sorted facts, defences, search queries.

Elements-first research skeleton (no citations)
You are assisting a [jurisdiction]-qualified litigator. Do not cite any case. Cite statutes only where you are confident, tagged [VERIFY].
Build the analytical skeleton for a memo on whether [client, anonymised] can [establish / defend] a claim for [cause of action] under [jurisdiction] law on these facts: [anonymised facts].
Output: (1) the elements, numbered, with the standard of proof; (2) for each element, the facts that support it, cut against it, and are still unknown; (3) the three most likely defences; (4) eight search queries for [Westlaw / Lexis / CourtListener / BAILII], phrased as I would type them.

Two rules keep this safe. Name the jurisdiction every time; as Clio puts it, “you’ve been thinking about Texas employment law all morning. The AI hasn’t.” And never ask for “the leading case on X”, which LexisNexis Canada’s Timon Sisic calls “the most common mistake people make” with legal AI: it is the request that produces a confident invention.

Mode 2: finding authority with grounded tools

Now move to a tool that retrieves documents and links to them: a legal platform, or a web-cited tool such as Perplexity Enterprise or ChatGPT with browsing. Retrieval-augmented research still fails because legal relevance depends on jurisdiction, date and posture, not text similarity: Stanford found 47% of Lexis+ AI’s errors were retrieval mistakes and 61% of Westlaw’s were reasoning errors made on the right documents.

Grounded authority search with a source fence
Identify the controlling authority in [jurisdiction] on [precise question, e.g. whether a liquidated damages clause in a commercial lease fixed at 150% of rent is an unenforceable penalty].
(a) Binding authority first, then persuasive, each labelled. (b) For every case: citation, court, year, pinpoint, and one sentence on what it actually holds on this point. (c) Note splits and later doubt.
(d) NEGATIVE CONSTRAINT: if you cannot find binding [jurisdiction] authority for a point, write "NO VERIFIABLE LOCAL AUTHORITY FOUND" and stop. Do not extrapolate from other jurisdictions or guess.
(e) End with a "Sources cited" index listing every authority once, in the order I should check them.

When nothing is squarely on point, ask for the analogy rather than the answer. LexisNexis Canada’s own Lexis+ AI prompt asks for cases where an insurance broker’s failure to verify vehicle ownership led to liability and adds: “If you are unable to find caselaw that specifically finds a broker liable for this type of failure, provide arguments for how analogous caselaw can be used to argue that such a standard of care exists.” Fenced to authorities you already hold, the same request is safe on any frontier model.

Perplexity has 80% of Gunderson Dettmer’s lawyers active, and one of them describes it in the right register: “I now use it for all my internet searches.” And AI agents, as Harvey’s own benchmark team found when testing them on diligence, have a bias “to search efficiently, not completely”; a research agent stops when it finds something, which is not the same as finding everything.

Mode 3: verification, which nobody gets to delegate

The Illinois Appellate Court set the standard in Scott v. Illinois Human Rights Commission: “The only acceptable standard is zero false citations.” The lawyer had used a “premier corporate subscription of ChatGPT”; the court held that “no matter how much one pays for ‘premier’ or ‘corporate’ versions of AI products, it does not negate an attorney’s obligation to verify all citations of authority”. The Ninth Circuit put it more briefly in Lnu v. Blanche: “A competent and diligent attorney must also read and reason.”

Six layers:

  1. Existence. Open it in Westlaw, Lexis, CourtListener, BAILII, the National Archives, RIS, juris or beck-online. If it cannot be opened, it does not go in.
  2. Identity. Names, court, year and reporter match; in Ayinde a fake case carried a genuine [2020] EWHC number.
  3. Status. Run the citator. Vendor trust markers are not your signature.
  4. Holding and pinpoint. The Illinois court was explicit: cross-referencing against a platform is insufficient unless the cited text actually appears in the source.
  5. Quotations. Character for character. False quotes from real cases account for 549 entries in Charlotin’s database.
  6. Jurisdiction and currency. Stanford caught Lexis+ AI applying the Casey standard after Dobbs had overruled it.

Judge Brantley Starr, author of the first US AI standing order, drew the same line in May 2023: these platforms have “many uses in the law: form divorces, discovery requests, suggested errors in documents, anticipated questions at oral argument. But legal briefing is not one of them.”

The prompt pattern: jurisdiction, negative constraint, sources-cited index

Every AI legal research prompt that survives contact with a judge has three load-bearing parts.

Jurisdiction, stated affirmatively and dated. “Cite only English authorities” beats “do not cite US cases”; models follow affirmative instructions better than prohibitions. “As at 13 September 2026” forces the model to surface currency limits instead of silently applying stale law.

A negative constraint with an exit. Justia’s “NO VERIFIABLE LOCAL AUTHORITY FOUND” gives the model a permitted way out other than invention. OpenAI’s own paper on hallucination explains why that matters: models are trained on evaluations that “penalize uncertainty”, so they learn to guess. Your prompt has to reward the blank.

A sources-cited index. Every authority once, in checking order, so verification becomes working down a table rather than re-reading a memo.

Fourth, check your own question: asked why Justice Ginsburg dissented in Obergefell, Ask Practical Law AI agreed with the premise, so add “list every premise in my question and tell me which you cannot confirm”. And on a reasoning model, drop “think step by step” and “double-check your answer”; the rules for prompting reasoning models are different, and a reasoning model is not a checking model.

Where the paid tools stand in September 2026

Tool Independent evidence Verification aids
Lexis+ with Protégé Stanford 2024: more than 17% hallucination, 65% accurate; entered neither 2025 Vals test Shepard’s Verify Trust Markers flag citations that cannot be verified against Lexis content: existence, not support
CoCounsel Legal / Westlaw (August 2026) Stanford 2024: Westlaw AI-Assisted Research more than 34%; Thomson Reuters declined the October 2025 Vals test Westlaw Brief Builder and Deep Research Verify, which the vendor says checks that cited authority “actually supports each legal assertion”; untested independently
vLex Vincent Vals February 2025: it “refused to answer, rather than hallucinate”; a 2025 randomised trial counted 3 hallucinations for Vincent, 4 for students without AI, 11 for o1-preview Grounded corpus; refusal is a feature
ChatGPT (Business or Enterprise) Vals October 2025: 80% accuracy, 70% authoritativeness; Stanford: at least 58% hallucination for GPT-4 Web citations only; never the source of a filed citation

Stanford tested in April 2024; as of March 2026 neither major vendor had published an independent re-test. The Harvey, Legora and CoCounsel comparison covers the work platforms; for research, vendors that submit to public benchmarks (Alexi, Counsel Stack and Midpage did) earn more trust than vendors that decline.

What lawyers say about research tools in practice

The Reddit consensus on AI legal research is sharper than most vendor decks. An r/biglaw associate on the general models: “it is really good for very basic questions, and then completely falls apart when you are looking for any edge case”. On r/LawSchool: “Westlaw 100% hallucinates case holdings. It just may not hallucinate case names.”

The enthusiasts are specific. A lawyer on r/Lawyertalk: “It does typically direct me towards a case that contains the head note I need in about 1/3 of the time of doing it myself on westlaw.” Pierce, a lawyer at a small business and real estate firm in Missouri, told Clio’s 2025 Legal Trends Report: “Now it takes literally five minutes and my first two hours of research are done.”

In the State Bar of Texas 2026 survey, “the single most requested improvement was AI that produces reliable, hallucination-free legal research with accurate citations”.

A solo’s stack without Westlaw or Lexis

The honest answer to “what is the best consumer AI for legal research” is that no consumer tool is a research database, and those asking are most at risk: Stanford’s analysis of US lawyer sanctions found solos accounted for 50.4% of the firms involved and firms of 2-25 lawyers another 39.5%. The gate Judge Castel described in Mata, “existing rules impose a gatekeeping role on attorneys”, is a database you can open, not a subscription tier. A defensible stack without the big two:

  1. A frontier model in a no-training tier (ChatGPT Business, Claude Team, Copilot with enterprise data protection) for Mode 1 of AI legal research. The prompt library has the standard header to paste in front of client-adjacent work.
  2. A web-cited tool pointed at official sources for Mode 2. The Divisional Court in Ayinde listed the English ones: legislation.gov.uk, the National Archives, the official Law Reports and reputable publishers.
  3. A closed-universe tool. Gemini Notebook answers only from the cases you upload, with clickable citations.
  4. A free citation checker for layer one. CaseRead, LawDroid CiteCheck AI and GroundTruth catch non-existent cases; none catches misgrounding. The citation checker comparison shows what each layer costs.
  5. A log. Tool, prompt, what was verified, by whom, when.

The ten-minute cite-check before anything leaves your desk

Budget it. LeanLaw’s checklist puts existence at 30 seconds to two minutes per citation, holding and quote at two to five, citator at one to three and applicability at three to five. Ten minutes a cite, twenty citations, three hours: the price of Mode 2.

Build the verification table (the model lists, you check)
List every case, statute, rule and secondary source cited in <document> in a table: Citation as written | Proposition it supports (quote my sentence) | Pinpoint given? | Quotation? | Red flags (reporter, volume or year mismatch; suspiciously on-point case name; too-perfect quotation).
Do not tell me whether any citation exists or is good law; I will check each row in a primary database. Add a blank column "Verified by / database / date".

Check the other side’s brief too: in Noland v. Land of the Free the winning party lost its fee award because it “did not alert the court to the fabricated citations”.

This is the exercise we run in AI Lab for Lawyers: one research question, three tools, every citation opened live, so the habit forms in class rather than in front of a judge.

Where to go next: the six-layer verification protocol expands Mode 3 into a printable checklist; the ChatGPT prompts for lawyers collection carries the research prompts with verification steps attached; and the overview of how lawyers use AI shows where research sits beside drafting and review in the use-cases cluster. If you would rather build the workflow with someone watching, that is what the four live sessions of AI Lab for Lawyers are for.

Frequently asked questions

Can I use ChatGPT for legal research?

Yes for issue-spotting, structuring an argument and drafting search queries; no for producing citations you will file. In Stanford's benchmarks general chatbots hallucinated on at least 58% of legal queries, and even the Vals test where ChatGPT scored 80% on accuracy put it below the legal platforms on authoritativeness. Use ChatGPT in a Business or Enterprise tier, tag every authority it names as unverified, and open each one in a real database.

What is the best AI for legal research?

There is no single winner. For finding authority, a grounded platform that links to source documents (Lexis+ with Protégé, CoCounsel Legal, vLex Vincent) is safer than a chatbot because you can click through. For orientation, any frontier model works. The best tool is the one you verify against a primary database, since Stanford found even the paid platforms hallucinating on more than 17% of queries (Lexis+ AI) and more than 34% (Westlaw AI-Assisted Research).

Is Perplexity better than ChatGPT for legal research?

Perplexity's advantage is that it cites the web pages it read, so you can check the source in one click; Gunderson Dettmer reports 80% of its lawyers active on it, and one of them says it now handles all his internet searches. It is not a legal database, so it can cite a blog summary of a case rather than the case. Point it at official sources (CourtListener, BAILII, legislation.gov.uk) and treat it as a finding tool, not an authority.

Do I still need Westlaw or Lexis if I have AI?

For anything you file, you need a primary database of some kind, but it does not have to be Westlaw or Lexis. Mata v. Avianca happened at a firm with only a limited Fastcase plan. CourtListener, BAILII, the National Archives and RIS are free primary sources; the point is that the case must be opened and read where it actually lives. AI tools find candidates; a database confirms them.

How do I verify AI research before citing it?

Six layers, in order: the case exists in a database; the party names, court, year and reporter match; a citator shows it is still good law; the pinpointed passage says what you say; every quotation matches character for character; and the jurisdiction and posture fit. Budget roughly ten minutes per citation, log who checked what, and never ask the AI whether its own citations are real.

Written by

Dr. Niklas Schmidt, Partner at Wolf Theiss

Partner at Wolf Theiss Attorneys-at-Law, where he heads the firm-wide tax team; lawyer, author, TEDx speaker and technologist. He has spent well over 1,000 hours testing practical AI applications for legal work, runs a toolkit of roughly 80 AI tools in daily practice, founded the WT Crypto Academy (1,000+ participating lawyers) and has given around 450 talks over 20 years. He teaches the live course AI Lab for Lawyers on Maven.