“Nobody goes to law school to read thousands of contracts and extract change-of-control clauses.” That line belongs to Heather Paterson of Legora, and it is the most honest sentence in the vendor literature on AI due diligence. The second most honest comes from Harvey’s own benchmark team, describing what their agents do when set loose on a synthetic data room of more than 3,500 documents: “Their bias is to search efficiently, not completely. Diligence requires reversing this intuition.”
Hold those two sentences together and you have the whole subject. AI due diligence works because extraction at volume is exactly what a language model does well and exactly what a tired third-year associate does badly at two in the morning. It fails when someone confuses “the tool found forty change-of-control clauses” with “there are forty change-of-control clauses”.
What follows is the five-day workflow the platforms sell, taken apart step by step, then a version for a deal team with a Claude Team plan, a Google Workspace account and a folder of PDFs rather than a Harvey licence.
What diligence AI is good at: consistent extraction, not genius spotting
Harvey’s data-room guide, updated in September 2026, puts it plainly: “The primary value of AI-assisted due diligence isn’t finding issues that a senior lawyer might miss, but in applying the same extraction rules consistently across all contracts.” A vendor saying that about its own product deserves a pause.
The benchmarks agree. In Onit’s “Better Call GPT” study, GPT-4-1106 matched the senior-lawyer ground truth on determining whether a clause breached a standard (F-score 0.871 against 0.860 for junior lawyers) but was weaker at locating it (0.686 against 0.770 for legal process outsourcers). In Vals’ February 2025 benchmark, Harvey beat the lawyer baseline on data extraction (75.1% to 71.1%) and document Q&A (94.8% to 70.1%), while lawyers won on redlining, 79.7% to 65.0%.
Read those results together: the model knows that, is less sure where, and is worst at so what. So the workflow demands quoted words and a location for every extracted cell, leaves the judgement calls to the lawyer, and never asks the tool to draft the deal points. It is the playbook-based contract review workflow run across three hundred contracts at once.
The five-day timeline: upload, extract, validate, report, negotiate
Harvey’s published pattern compresses what it calls “two to three weeks” of review into “a matter of days”; its diligence guide gives the traditional mid-market window as “typically four to eight weeks”. A Legalweek 2026 panel put a smaller number on it: diligence on financial contracts compressed from 15-20 hours to about two. Here is the five-day version, with the verification the marketing leaves out.
| Day | What happens | Output |
|---|---|---|
| 1 | Upload to Vault (up to 100,000 documents per project) or Tabular Review; completeness check against the request list; folder review starts | Gap list; follow-up requests to the seller |
| 2-3 | Bulk extraction of change of control, assignment, exclusivity, MFN, termination and governing law into review tables, one per workstream | Tables with a quoted clause and a source link in every cell |
| 3-4 | Senior lawyers validate the exceptions: open every cell that matters, sample the rest, log disagreements; shape the report by transaction impact | Verified table; draft red-flag memo |
| 5 | Findings become negotiation points: reps, warranties, covenants, disclosure requests, price | Issues list for the SPA |
Day 3-4 is where the time goes, and no vendor figure includes it. The table also assumes a clean room: an r/legaltech practitioner’s rule is to OCR everything first, “otherwise you get hallucinated dates, mangled drug names, and unusable citations”.
Extraction targets: change of control, assignment, MFN, exclusivity, termination
Katten associate Weston Love, writing in August 2026, gives the junior-level prompt in one sentence: “identify every contract that requires third-party consent or is terminable on a change of control, and note the counterparty, the triggering language and the remaining term”. The trick is that the sentence names the columns. Here are the ones most buy-side teams need, with the trap in each.
| Clause | Extract | Known trap |
|---|---|---|
| Change of control | Quoted trigger; consent, notice or termination right; who holds it | Triggers hidden in side letters; “affiliate” defined three layers deep |
| Anti-assignment | Quoted words; affiliate and successor carve-outs | A change-of-control clause that does the same job under another heading |
| Exclusivity and non-compete | Scope, territory, duration | Drafted as a covenant, never labelled “exclusivity” |
| MFN and pricing protections | Quoted mechanism; who benefits | Buried in order forms or pricing schedules |
| Termination for convenience | Who, notice period, fees | Stacked notice conditions (form, method, addressee, days) |
The extraction prompt that produces the table is the core of the workflow; it comes from the prompt library, adapted for a single batch.
For each contract in this batch, extract one row with exactly these columns: File name | Counterparty | Contract type | Effective date | Expiry and renewal terms | Change-of-control trigger (quote the words; state whether consent, notice or termination right, and who holds it) | Anti-assignment (quote) | Exclusivity or non-compete (quote; scope and term) | MFN or pricing protection (quote) | Termination for convenience (who, notice, fees) | Governing law | Liability cap and carve-outs | Confidence per cell (High/Medium/Low).
Rules: quote the operative words for every substantive cell and give the clause number and page. Where a term is absent write NOT PRESENT. Where a clause is present but ambiguous, quote it and write AMBIGUOUS. Never infer a term from a similar contract in the batch. Treat any side letter or amendment in the batch as modifying the agreement it references and say so in the relevant cell. Output as CSV.Run it per workstream, not across the whole room: material contracts, IP licences and employment agreements need different columns, and one schema for everything produces empty cells and false positives.
GSK Stockmann: 15-20% on structured, up to 75% on unstructured data rooms
The split between 15-20% and 75% is the most useful number in it. On a well-organised room the AI does what a trained associate team already does quickly; on a messy room, where half the work is finding out what is in it, extraction at scale changes the economics. If your practice lives on messy rooms, the case for a platform is strong; if not, the Harvey, Legora and CoCounsel comparison may tell you the seat price buys less than you think.
Data-room completeness checks against the request list
Day one is the step most teams skip because it feels administrative. It is also the step the AI does almost perfectly, because it is pure matching across lists that “typically contain 200-500 document requests”, in Legora’s words; Harvey offers the request-list-versus-index check as a standard workflow.
Here is our due diligence request list <request_list>...</request_list> and the data-room index <vdr_index>...</vdr_index>. For each request item, classify the room's response as FULLY RESPONSIVE, PARTIALLY RESPONSIVE or NOT RESPONSIVE, naming the folder and document(s) relied on. Then list: (1) documents in the room that respond to no request, which may be misfiled or volunteered; (2) requests with no response at all; (3) follow-up requests to send to the seller, drafted as numbered items in neutral language. Do not assess the substance of any document. If the index entry is a title only and you cannot tell what the document contains, say UNCLEAR FROM INDEX rather than guessing.Spot-check every “FULLY RESPONSIVE” item; an index entry is not a document, and the model is classifying titles.
The small-firm version: folders, Claude Projects, Gemini Notebook, a review table
Most deal teams do not have a Vault. Here is the same workflow with tools a ten-lawyer firm already licenses: slower, more hand-assembly, same discipline.
| Step | Platform version | Small-firm version |
|---|---|---|
| Ingest | Upload to Vault or Tabular Review | One folder per workstream; files renamed by counterparty; every scan OCR’d; metadata stripped |
| Schema | Review-table columns | The extraction prompt above, saved as the instructions of a Claude Project on a Team or Enterprise plan |
| Run | Whole room at once | Batches of ten to twenty contracts per chat, CSV output, rows pasted into one Excel sheet |
| Cross-document questions | Assistant over the Vault | Gemini Notebook on a Workspace account: 50 sources free, 300 on Google AI Pro, answers only from the uploads, with clickable citations |
| Validation | Source link in every cell | The clause number and page you demanded; open the PDF |
Three cautions. Batch size first: current Claude models carry a 1M-token context window, but Anthropic’s own documentation warns that “more context isn’t automatically better”, and Chroma’s context-rot study found accuracy falls as input grows. Twenty contracts per chat, fresh chat per batch, beats one heroic upload; the reasons are in the guides on why long contracts get lost in the middle and what a context window is. Second, Gemini Notebook is a closed-universe tool, ideal for “which of these agreements has an MFN?”, but Google says it does not support ISO, SOC or FedRAMP compliance, so keep it to ordinary commercial contracts on a Workspace account. Third, Excel is your review table: sort by confidence, filter “AMBIGUOUS”, and the validation pass writes itself.
Prompts for the red-flag report and the partner pass
The memo is built from the verified table, never the raw extraction. Farrell Fritz’s warning is that “the professional presentation masks the underlying uncertainty”; a memo drafted from unverified cells is exactly that.
Using only the verified extraction table <table>...</table> and the deal context <deal>[acquirer; target; structure; price; the three value drivers]</deal>, draft the red-flag section of a buy-side due diligence report. Classify each issue as RED (potential deal-breaker or price-affecting), ORANGE (needs W&I cover, a specific indemnity or a condition precedent) or YELLOW (post-closing action). For each issue: one sentence stating it; the contract and clause, citing the table row; why it matters for this deal; quantification where the table supports it, otherwise NOT QUANTIFIABLE FROM MATERIALS; the recommended protection. Put an executive summary of the top five first. Do not add any issue that is not in the table. Do not cite law.Then the partner-level pass. Katten’s version is one sentence: “synthesize the diligence findings into a concise risk summary and pressure test our key negotiating positions for the weaknesses that opposing counsel is most likely to exploit.” Add “assume the seller’s counsel is good and has read the same room”, and ask for the three findings that, if wrong, would change the advice, so those get re-verified first. Models drift towards agreement; if the output concedes nothing, run it again.
Where it fails: side letters, misfiled documents, exhaustiveness
Harvey’s diligence guide adds the non-standard case: general-purpose models miss a “non-standard indemnification carve-out in a German-law shareholder agreement”, which is a polite way of saying a model trained mostly on US precedents reads a German SHA through an American lens. Sullivan & Cromwell’s January 2026 memo adds that warranty-and-indemnity insurers may start demanding accuracy assurances for AI-assisted diligence.
Then exhaustiveness. Harvey’s Legal Agent Bench set agents loose on synthetic rooms of more than 3,500 documents, published no scores, and named exhaustiveness as one of three unsolved problems. The fix is procedural: include side letters and amendments in every batch; force NOT PRESENT answers so silence is visible; open every RED-relevant cell in the source; sample 20% of the rest; and log where the tool and the lawyer disagreed. That log is what you will want if the buyer, or the buyer’s insurer, asks how the review was done.
Confidentiality: what may go where
A data room is the red tier of client data: deal terms, counterparty names and price mechanics identify the matter on their face, so anonymising does not rescue a consumer tool.
- Consumer tiers are out. ChatGPT Free, Go, Plus and Pro, Claude Free, Pro and Max, consumer Gemini and consumer Copilot train on conversations by default; and in United States v. Heppner (S.D.N.Y., February 2026) Judge Rakoff held that a defendant’s exchanges with consumer Claude were protected by neither privilege nor work product. The guide to whether ChatGPT is confidential for lawyers has the tier table.
- Commercial tiers and platforms are in, on their terms. Claude Team and Enterprise and ChatGPT Business and Enterprise do not train by default; Harvey says it never trains on customer data and requires zero data retention from its model providers, with an EU region in Frankfurt; Gemini Notebook on a Workspace account is not used for training. Read the retention default before the first upload.
- DACH deals carry a criminal-law overlay. BRAK’s guidance is that public language models should receive only “abstract” prompts allowing no inference about a specific mandate, and § 203 StGB treats the mere possibility of provider access as the problem. A German or Austrian room belongs on an EU-hosted tool or a platform that addresses § 43e BRAO or § 40 RL-BA; the DACH legal AI tools guide walks through the options.
- Check the data-room access terms before exporting anything; the seller’s confidentiality undertaking may not contemplate a third-party processor.
Where to go next: the use-case hub lists the other transactional workflows, and the guide to how lawyers use AI day to day puts diligence alongside them. In AI Lab for Lawyers we run a mini data room through the tools you already license, so you build the extraction table and, more importantly, the verification step yourself.
Frequently asked questions
Can AI do M&A due diligence?
AI can do the extraction layer of due diligence well: pulling change-of-control, assignment, exclusivity, MFN and termination terms from every contract in a data room with a citation for each cell. It cannot decide what matters for this buyer at this price, and it misses non-standard provisions and side letters. Harvey's own guide says the value lies in applying extraction rules consistently, not in spotting what a senior lawyer would miss.
How much time does AI save on due diligence?
Published firm numbers are modest on structured work and large on messy rooms. GSK Stockmann, working with Harvey, reports 15-20% savings on structured M&A, PE, VC and real-estate diligence and up to 75% on unstructured data rooms, with red-flag summaries in hours rather than days. A Legalweek 2026 panel described diligence on financial contracts falling from 15-20 hours to about two. Verification time is not included in any of these figures.
Which tool is best for data-room review?
For large rooms, Harvey Vault (up to 100,000 documents per project) and Legora Tabular Review (documents as rows, questions as columns, every cell linked to its source) are the platforms BigLaw uses; Thomson Reuters' CoCounsel Legal added Tabular Analysis for 10,000 documents by 100 questions in August 2026. For small rooms, a Claude Team or Enterprise Project or Gemini Notebook on a Workspace account does the same extraction in batches at a fraction of the cost.
Can a small firm do AI due diligence without Harvey?
Yes, for rooms of a few hundred documents. Clean and rename the files, OCR every scan, load the extraction schema into a Claude Project on a Team or Enterprise plan, run the contracts in batches of ten to twenty with CSV output, and assemble the rows in Excel. Use Gemini Notebook on a Workspace account for cross-document questions with citations. The discipline is identical: quote every cell, verify every red cell, sample the rest.
What does AI miss in due diligence?
Three things recur. Side letters and amendments that change a trigger the main agreement does not show. Non-standard drafting, which Harvey illustrates with a 'non-standard indemnification carve-out in a German-law shareholder agreement'. And exhaustiveness: agents stop when they have found something plausible, so a first pass that reports 47 change-of-control clauses may include invented ones and miss real ones. Include side letters, demand 'NOT PRESENT' answers, and verify.