Ask a frontier model to structure an exchange under IRC § 351 and it will give you something that satisfies every literal requirement of the section. Bloomberg Tax has a name for the result: a “technically perfect sham”, an arrangement that ticks each statutory box and fails the economic-substance doctrine because nothing about it has a business purpose beyond the tax result. The model was not being careless. It was matching words, and tax law is the one field whose anti-abuse doctrines exist to defeat word-matching.
That is the structural problem with AI for tax lawyers, and it shapes everything here. I head a firm-wide tax practice and use these tools daily; the question is not whether but where. The answer is a line: statute parsing, preliminary memos and indirect-tax questions on one side; substance, structuring and arithmetic on the other; a memo protocol in between.
Tax is structurally hostile to LLMs: literal compliance is not enough
Bloomberg Tax’s Pramod Kumar Siva sets out the failure modes by section. The vulnerable provisions he names are §§ 351, 199A, 6662, 6700 and 7701(o), and the common thread is that each turns on characterisation rather than text: whether a transaction has substance, whether a position has substantial authority. His conclusion: “Hallucinated authority isn’t merely a drafting defect, but a substantive legal failure.”
A general model’s strengths are the opposite of what these doctrines demand: it excels at producing text that resembles compliant text, and has no view on whether the arrangement survives an examiner who asks why the steps were taken in that order.
AI for tax lawyers, the acceptable half: parse, draft, answer the rate question
Bloomberg Tax’s dividing line is the one I use. Acceptable under supervision: “parsing statutes and drafting preliminary memoranda”. Not acceptable: independent research without verification, anti-abuse doctrine analysis, statutory interpretation after Loper Bright, and transaction structuring.
| Task | Verdict | Why |
|---|---|---|
| Parsing a statute or regulation into elements, definitions and cross-references | Acceptable under supervision | The text is supplied; the output carries section references you can check |
| Preliminary memorandum from authorities you have verified | Acceptable under supervision | Structure and drafting, not research; every citation is already yours |
| Indirect-tax Q&A: rate, exemption, certificate on file | Acceptable at volume, from your own tables | Lookup and classification against a source you control |
| Economic substance, step transaction, § 7701(o) analysis | Not acceptable | Characterisation, not text; the “technically perfect sham” problem |
| Statutory interpretation after Loper Bright | Not acceptable | The interpretive framework moved; training data did not |
| Transaction structuring | Not acceptable | Produces what matches the words, which is what the doctrines defeat |
The TEI roundtable adds the corporate-department uses: tax-notice responses, file retrieval, product classification, and Senay Redda’s orientation example, “If you spun off a C corp, how does it become an S corp?”, which yields “a step-by-step answer”. The parsing prompt is the one I run most.
From the statutory text and regulations pasted below only, produce for IRC § [351] as at [date]: (1) each requirement as a numbered element with the operative words quoted; (2) every defined term used, with the section that defines it; (3) every cross-reference to another section, and what it does; (4) the exceptions and anti-abuse hooks, quoted; (5) a list of what the text does not answer and would need a ruling, regulation or case to resolve. Cite only the pasted material. Where you would draw on outside knowledge, tag it [VERIFY] and keep it out of the elements list.
<statute>[paste]</statute>
<regulations>[paste]</regulations>The fence is the same logic that keeps AI legal research honest: the model works on what you gave it, and every line is checkable.
What I do not delegate: substance, structuring and the answer it wants to give
Redda’s line explains the second half of the table: “Generative AI is designed with a bent to give an answer”. Ask whether a structure works and you get one that works, in the sense that the paragraph reads as if it does. Damien Charlotin, who runs the hallucination database, put the related risk bluntly: “The harder your legal argument is to make, the more the model will tend to hallucinate, because they will try to please you” (CalMatters). Aggressive positions are hard arguments by definition.
The research on false premises is worse. Dahl and colleagues at Stanford found that models “often uncritically accept users’ incorrect legal assumptions” and that GPT-4 hallucinated on 53 to 69 percent of false-premise questions (Large Legal Fictions). A tax question almost always contains a premise: that the entity is a partnership, that the safe harbour applies, that the ruling is still good. The sycophancy guide covers the mechanism; the response is a prompt that attacks the premise first.
The arithmetic problem
Will Matthews of Bloomberg Tax says models “still tend to trip over things that require that kind of calculation”, and Thomson Reuters lists “math, counting, and sorting” among their weaknesses. Anthropic’s interpretability work explains why: asked how it added 36 and 59, Claude “describes the standard algorithm involving carrying the 1” although that is not what happened internally (Anthropic). A model’s account of its own arithmetic is a story; asking “are you sure?” is not verification.
My rule: the model may write the formula and explain the mechanism; the spreadsheet computes.
Explain the computation of [the § 199A deduction / the § 6662 substantial-understatement threshold / the earn-out adjustment] under [jurisdiction, year] as a step-by-step formula I can enter in a spreadsheet: each input with its source, each intermediate value with the operation, each threshold or cap with the section that imposes it [VERIFY]. Do not compute a result. Where the treatment depends on an election or a fact I have not given you, stop and list it as [INPUT NEEDED].Indirect-tax questions at volume
The high-volume use case the TEI roundtable likes is Michael Bernard’s at Vertex: “Can I get the rate? Is it exempt? Is it exempt for a particular purpose? Do we have a certificate online?” Those are lookup and classification questions against data the department owns: the shape a model handles well when the data is supplied and badly when it is not.
Using only the rate table, exemption matrix and certificate register pasted below, answer for each line item in <request>: the applicable rate and its table row; whether an exemption applies, quoting the matrix entry; whether a valid certificate is on file, with reference and expiry; and your confidence. Where the tables do not cover the jurisdiction or product, write NOT IN TABLES and route to [name]. Use no rate or rule from outside these materials.
<tables>[paste]</tables>
<request>[paste]</request>Research tools: Ask Blue J and the fourfold reasoning gain
Grounding in a curated corpus changes the failure mode rather than removing it. Ask Blue J curates federal tax case law with primary sources and positions itself as predicting how a court or tax authority would rule; OpenAI’s reasoning guide reports that Blue J saw a “4x improvement in end-to-end performance” moving from GPT-4o to the o1 reasoning model (OpenAI).
Two caveats. OpenAI’s o3 and o4-mini hallucinated more than o1 on its PersonQA test (33% and 48% against 16%) because they “make more claims overall” (TechCrunch); a reasoning model is not a checking model, and the reasoning-model guide explains how to prompt one. And a premium subscription changes nothing: the Illinois Appellate Court held in Scott v. Illinois Human Rights Commission (July 2026) that “no matter how much one pays for ‘premier’ or ‘corporate’ versions of AI products, it does not negate an attorney’s obligation to verify all citations of authority”.
A poster on r/LawSchool called a firm’s tax model “scary good”, completing “all junior associate and law clerk research assignments with ease”, then added the operative clause: “I just have to spend a couple hours double checking the work” (/u/igtr). Those hours are the job.
Thomas v. Commissioner and the struck memorandum
Will Matthews’ remark at the TEI roundtable is the client’s view of the same episode: “Nobody’s going to be comfortable in the foreseeable future—telling a stakeholder, ‘Hey here’s the answer; I got it from a chatbot.’” The fake cases explainer covers why a section number that sounds right is the dangerous kind.
Jurisdictional flattening: federal thick, state and international thin
Siva’s second structural point is “jurisdictional flattening”: federal law is over-represented in training data, state and international law under-represented, so a model defaults to the federal answer and, in the Tax Court, risks a Golsen problem by applying the wrong circuit’s law. For a European practice the flattening is worse: a model asked a cross-border question reaches for US concepts unless told otherwise, and the confidentiality rules are stricter. German and Austrian bar guidance limits public tools to abstract prompts that allow no inference about a specific mandate, and in Germany § 203 StGB makes the breach a criminal one; the DACH tools guide has the options.
The verification protocol: authority levels and annotated databases
The TEI roundtable’s recommendation is a memo standard rather than a tool rule: specify technical merit, the authority level of the conclusion (“more likely than not”, “should”, “will”), accounting implications and disclosure, and check every authority in “reputable annotated databases”. Matthews adds that “The quality of the response you’re getting is going to depend entirely on the quality of the underlying material.” A verification appendix makes the checking visible.
Draft a tax memorandum on [question] under [jurisdiction, year]. Sections: Issue; Short answer with an explicit confidence level (more likely than not / should / will) and the biggest reason for uncertainty; Facts from <facts> only; Law, with every statute, regulation, ruling and case tagged [VERIFY] with the exact section or citation; Analysis, distinguishing the taxpayer's position from the authority's likely characterisation; Risks, penalties and disclosure; and an Appendix listing every authority with the proposition it supports, the words relied on, and an empty column "Verified by / date / database". Where the law changed after [date], state the effective date and what applied before. Do not compute figures.A human fills in the appendix, in Bloomberg Tax or another reputable annotated database, before the memo goes out. The citation verification guide has the six-layer check; in tax, currency is where the damage happens.
Where to go next: the practice-area hub and the other practice-area guides draw the same supervision line for other specialisms, the real estate guide covers extraction with clause references, and the prompt library has the tax memo recipe. In AI Lab for Lawyers I show the tax workflows I trust and the ones I refuse to delegate, tools open on screen.
Frequently asked questions
Can ChatGPT do tax research?
It can orient you: which sections are in play, how a step is sequenced, what a term of art means. Thomson Reuters' Senay Redda cites 'If you spun off a C corp, how does it become an S corp?' as a question that gets a good step-by-step answer. What it cannot do is supply authority you can rely on; the Tax Court has struck a memorandum built on fabricated cases, and every section number needs checking in an annotated database.
Why is AI risky for tax structuring?
Because tax law punishes literal compliance without substance, and a language model optimises for literal compliance. Bloomberg Tax's example is an IRC § 351 exchange that satisfies every statutory element and is still a 'technically perfect sham' under the economic-substance doctrine. The model does not weigh business purpose, step-transaction risk or how the IRS would characterise the arrangement; it produces the structure that matches the words, and the anti-abuse doctrines exist to defeat exactly that.
Is AI good at tax calculations?
No. Bloomberg Tax's Will Matthews says models 'still tend to trip over things that require that kind of calculation', and Thomson Reuters lists 'math, counting, and sorting' among documented weaknesses. Anthropic's own research found that Claude's explanation of how it adds two numbers does not match what it actually does internally. Use the model to write the formula and explain the mechanism, then compute in a spreadsheet you can audit.
What tax tasks can AI safely do?
Under supervision: parsing statutes and regulations into elements and cross-references, drafting preliminary memoranda from authorities you supply, answering high-volume indirect-tax questions such as rates, exemptions and certificate status from your own tables, classifying products, and drafting responses to tax notices from the file. In each case the model works on material you provided and the output carries references you can check, which is the line that separates safe uses from research it did on its own.
Which AI tools do tax lawyers use?
Corporate tax departments use Vertex-type indirect-tax systems for rate and exemption questions, Bloomberg Tax and other reputable annotated databases for research, and Ask Blue J, which curates federal tax case law with primary sources and positions itself as predicting how a court or tax authority would rule. Frontier models such as ChatGPT and Claude are used for structure and drafting; OpenAI reports that Blue J saw a fourfold performance gain moving from GPT-4o to the o1 reasoning model.