In May 2023, a New York lawyer named Steven Schwartz filed a brief in Mata v. Avianca that cited six cases. The prose was clean, with reporter numbers, parallel cites, and the full apparatus a judge expects to see.
There was just one problem: the cases did not exist.
ChatGPT had written all of it. When Schwartz asked the chatbot whether the cases were real, it cheerfully confirmed they were.
That story is now legal-tech folklore, and the easy lesson everyone took is "don't trust ChatGPT." But the easy lesson is the wrong one.
The brief was not bad writing. It was excellent writing attached to nothing real.
Which brings us to the question this whole genre of blog post keeps asking and answering badly.
The question is a trap
When someone searches for the best AI chatbot for legal writing, they usually picture a leaderboard. ChatGPT versus Claude versus Gemini, ranked by who drafts the cleanest motion. That framing assumes the hard part of legal writing is the writing.
It is not. Every sentence a lawyer files tells a court that something is true. A chatbot has no idea whether anything it says is true.
So the honest 2026 answer is simple. The best AI chatbot for legal writing depends on which job you are handing it.
For first-draft scaffolding, structure, and tone, the generic chatbots are genuinely excellent and roughly interchangeable. For anything that asserts a legal authority, a holding, a citation, a statutory subsection, none of them are safe.
"Which one is best" is the wrong contest. The real axis is not writing quality, it is verifiable grounding.
TL;DR
- The "best chatbot for legal writing" framing is a trap. The gap between tools is not prose quality, it is whether a cited proposition ties back to a real, checkable source.
- Generic chatbots (ChatGPT, Claude, Gemini) are excellent drafting assistants and unreliable legal authorities. Use them for scaffolding and tone, never for citations.
- "Legal" branding does not mean hallucination-free. A Stanford study found purpose-built legal research tools still hallucinated on roughly 1 in 6 queries or worse.
- ABA Formal Opinion 512 already names the duty: understand a tool's propensity to hallucinate and independently verify its output, with more scrutiny for research and drafting.
- The fix is not a smarter model. It is grounding architecture (retrieval over real opinions and statutes) plus human verification.
- For raw drafting, ChatGPT and Claude (both $20/mo consumer tiers) are interchangeable. For anything you file, use a grounded tool whose citations link to a source.
- If your work is specifically contract review, the calculus shifts. See our contract-review guide.
What is the real axis for picking a legal-writing AI, per this post?
Part of our document tools, redline, and matrix guide series.
For related document-tools coverage, see Which Legal AI Is Best in 2026? A Capability Comparison of Agentic Suites and Legal AI Workflows: How Law Firms Chain Multi-Step AI Tasks in 2026.
How we judged these tools
We did not score chatbots on prose quality. The leaders are all good writers, and that is not where the risk lives. We judged each tool on one axis: can a lawyer verify what it told them, cheaply, before it hits a filing.
That means asking one thing. Does a cited proposition tie back to a real, checkable source, or is it a confident string with nothing behind it?
Every factual claim here is grounded in real behavior and public sources, linked where they are verifiable. Prices are vendor-published consumer numbers where we are confident, and omitted where we are not.
Where a chatbot is not a grounded legal tool, we say so plainly instead of pretending the gap does not exist.
The tools at a glance
The split that matters is the last column. "Grounded" means the tool retrieves real source documents and links its citations back to them, so you can check a cite in a click. Every price below is either the vendor's own published number or a named, dated third-party source.
| Tool | Price | Access | Grounded in real law? | Best for |
|---|---|---|---|---|
| ChatGPT | $20/mo Plus (free tier) | Self-serve | No | Fast first drafts, tone work |
| Vaquill AI | Published | Self-serve | Yes (US opinions and statutes) | Solo GCs, in-house teams |
| Claude | $20/mo Pro (free tier) | Self-serve | No | Long-document drafting, careful prose |
| Gemini | Free tier; paid consumer tier | Self-serve | No | Drafting inside Google Workspace |
| Spellbook | ~$500/seat/mo (quote-based) | Sales | Partial (clause data; no case law) | Contract drafting in Word |
| CoCounsel | ~$225 to $400+/user/mo (TR) | Sales | Yes (Westlaw corpus) | Litigation research and drafting |
Sources and the honest tradeoffs for each are in the per-tool sections below. ChatGPT, Claude, and Gemini prices are vendor-published consumer tiers. Spellbook does not publish a price; the ~$500/seat figure is a market estimate we cannot link to a public sticker. CoCounsel's range is a market estimate as well, since Thomson Reuters quotes per firm.
Where the generic chatbots actually shine
I want to be fair to the chatbots, because the "AI is dangerous" crowd tends to throw out something genuinely useful.
ChatGPT, Claude, and Gemini are, in 2026, extraordinary first-draft engines. Hand any of them a fact pattern and a request for a demand letter, a deposition outline, or a client update email. You will get back something coherent in seconds.
They are good at structure: IRAC, headings, transitions, the rhythm of a persuasive paragraph.
They are good at tone-shifting too. They can soften a blunt internal memo for a client, or sharpen a polite draft into something that lands harder.
Hand three senior litigators the same blind drafting prompt, with no citations involved. Most could not reliably tell you which chatbot wrote which output. The prose ceiling is high and flat across the leaders.
This is exactly why the "which one writes best" comparison feels so unsatisfying when you run it. They are all good. That is not where the risk lives.
The risk lives the moment the draft needs to say "as the court held in" anything.
The trap is the confidence, not the typos
Here is the part most people get backward. They worry the AI will produce sloppy, obviously-wrong text they can catch. The opposite is true.
The danger is fluent, confidently formatted output that looks exactly like competent legal work.
A fabricated citation from a good model does not look fabricated. It has a plausible case name, a real-sounding court, and a reporter cite in the correct format. The pull quote reads like something a judge would actually write.
Mata v. Avianca was not exposed because the writing was bad. It was exposed because opposing counsel tried to pull the cases and could not find them.
And the problem is not receding as models improve. According to reporting from Cronkite News, documented AI legal-hallucination incidents passed 300 since mid-2023. Roughly 200 of those were logged in 2025 alone.
The cadence went from a couple of incidents a week to two or three a day. Better models have not closed the gap. If anything, more lawyers trusting better-sounding output has widened it.
The sanctions have teeth. In 2025 a California appellate court fined a lawyer after 21 of the 23 quotes in a brief turned out to be fabricated. A federal judge in New Orleans fined a lawyer $1,000 over roughly eleven citations that were fabricated or misused.
These are not edge cases from 2023 anymore. This is a steady drip of professional embarrassment, two years after everyone supposedly learned the lesson. We track the pattern and the workaround in legal AI that avoids hallucinating cases.
The nuance that actually matters: grounded is not the same as infallible
If the answer were simply "use a legal-specific tool instead of a generic chatbot," this post would be short. It is not that simple, and pretending otherwise is how vendors sell.
The most important piece of research here is the Stanford RegLab and HAI study, "Hallucination-Free?", published in 2024 and later peer-reviewed in the Journal of Empirical Legal Studies. The researchers tested the purpose-built legal research tools, the ones marketed as "hallucination-free."
They found that Lexis+ AI and Ask Practical Law hallucinated on more than 17% of queries, while Westlaw's AI-Assisted Research hallucinated on more than 34%. Roughly one in six queries, or worse, despite the marketing.
Read that twice, because it kills the lazy conclusion. A "legal" brand on the box does not mean the tool cannot make things up.
The difference between a generic chatbot and a well-built legal tool is not that one hallucinates and one does not. The difference is degree, and more importantly, it is auditability.
A grounded tool retrieves real documents and answers on top of them. It hallucinates less because it has something real to anchor to. The deeper value is traceability.
When it cites Carpenter v. United States, 585 U.S. 296 (2018), you can click through to the actual opinion and read the holding.
With a raw chatbot, the citation is a guess dressed as a fact. The only way to check it is to go find the source yourself, which most people under deadline do not do.
That is the entire game. Generic chatbots optimize for fluent text, grounded legal tools optimize for traceable claims. Both can be wrong, but only one lets you catch it efficiently.
Why grounding is an architecture problem, not a model problem
The instinct, when you learn a chatbot hallucinates, is to wait for a smarter model. GPT-6 will fix it, the next Claude will fix it. This is a misunderstanding of what is happening.
A language model is a prediction engine. Asked for a citation supporting a proposition, it predicts the most statistically plausible citation-shaped string. Sometimes that string corresponds to a real case, sometimes it does not.
The model has no internal database of what is real. A smarter model produces a more plausible-sounding string, which, if anything, makes a wrong answer harder to catch, not easier.
The fix is to change the architecture, not the model. Retrieval-augmented generation (RAG) does this. Before the model writes anything, the system pulls real opinions or statutes from an actual corpus, then makes the model answer from those documents with citations pointing back.
The model is still a model. But now it reasons over real text instead of free-associating from training data.
We walk through the mechanics in how AI legal research works with RAG. The cost and reliability tradeoffs of doing this properly, versus calling a raw model, are in API call vs RAG pipeline.
This is also why the data layer matters more than the chat layer. A grounded tool is only as good as the corpus it retrieves from. In the US, that foundation is a corpus of millions of federal and state opinions, the actual text a tool must retrieve and cite.
A tool that retrieves over that corpus can show you the actual Loper Bright v. Raimondo, 603 U.S. 369 (2024) opinion that overruled Chevron. A chatbot can only tell you that Loper Bright overruled Chevron and hope it remembered correctly.
The duty is already written down
If you are a US lawyer, you do not get to treat this as an open question of taste. ABA Formal Opinion 512, issued in July 2024, already settled the professional obligation. It requires lawyers to understand the tools well enough to grasp how they hallucinate, and to independently verify AI output before relying on it.
The opinion scales that duty to the stakes: drafting and legal research, the things that go to a court or a client, demand more scrutiny than brainstorming does.
That single principle resolves most of the chatbot-comparison anxiety. The question is not "which chatbot can I trust," because the answer to that is none, fully. The real question is "which workflow lets me satisfy my verification duty without it eating my entire afternoon."
A tool whose every citation links to its source makes verification a click. A chatbot whose citations are unsourced strings makes verification a research project you run from scratch.
We unpack the opinion in plain English in our ABA 512 guide. The data-handling side of competence, where your client confidences actually travel, is in where your legal AI data actually goes.
So which tool, for which job?
Here is the honest division of labor I would give a colleague.
Use a generic chatbot (ChatGPT, Claude, Gemini) for:
- First drafts where you supply the facts and the law yourself
- Restructuring, tightening, and tone adjustments on text you already wrote
- Outlining arguments, brainstorming counterarguments, drafting client-facing emails
- Plain-language explanations of a concept you will verify elsewhere
Do not use a generic chatbot to:
- Find supporting authority for a proposition
- Quote or paraphrase a holding it surfaced on its own
- Confirm a statute's text or a case's procedural posture
- Generate anything you will file without independently sourcing every cited claim
Use a grounded legal tool (the kind built on retrieval over real opinions and statutes) for:
- Anything that asserts what the law is and needs a citation a court can check
- Comparing how a doctrine evolved across multiple opinions
- Pulling and verifying statutory text against the actual code
The AI-native drafting suites (Harvey, Legora, CoCounsel, and the leaner US-focused stacks) live in that last category. They are not "better chatbots." They are chat surfaces sitting on top of retrieval over real corpora, which is a different product even when the text box looks identical.
The honest version of the pitch is that grounding makes verification cheap, not that it makes verification unnecessary. Anyone telling you their legal AI eliminates the need to check is selling you the Mata v. Avianca sequel.
If you want specifics on the named tools, here is the compact version. None of the general chatbots are grounded legal tools, so treat them as drafting assistants, not authorities.
ChatGPT
At a glance: ChatGPT Plus $20/mo (free tier available) · web and apps · best for: fast first drafts and tone work.
It is the default for most people, and for good reason. The writing is clean and the turnaround is instant.
What's good:
- Excellent at structure, IRAC, and reworking text you already wrote
- Huge ecosystem of integrations and a free tier to try first
- Strong at tone-shifting a draft for a different audience
Where it falls short:
- It invented the fake cases in Mata v. Avianca, and it still fabricates citations
- No traceable source behind a legal claim, so every cite is a research project to check
What users say: Lawyers debating tool budgets on r/legaltech keep landing on the same point: a $20 ChatGPT seat does most of the drafting a far pricier "legal" wrapper does, so the question is whether you also need grounding (r/legaltech pricing thread, 2026).
Bottom line: Great for scaffolding you supply the law for. Never let it source authority you plan to file.
Claude
At a glance: Claude Pro $20/mo (free tier available) · web and apps · best for: long-document drafting and careful prose.
Many lawyers prefer its tone and its handling of long context. It tends to hedge rather than bluff, which helps a little.
What's good:
- Strong long-context handling for big drafts and memos
- Often more measured in tone than other general chatbots
- Free tier to test before you pay
Where it falls short:
- Still no grounding, so a confident citation can be invented
- Cautious phrasing can read as authority it does not have
What users say: In a 2026 r/legaltech budget thread, several lawyers said they prefer Claude's tone for memos over the pricier suites, while noting it is still a raw model with no source behind its cites (r/legaltech pricing thread, 2026).
Bottom line: A solid drafting partner for text where you own the law. It is not a legal research tool, full stop.
Gemini
At a glance: Free tier available; paid consumer tier exists (vendor pricing varies) · web, apps, Google Workspace · best for: drafting inside Google Docs and Gmail.
Its pull is integration. If your work lives in Google Workspace, it is already there.
What's good:
- Drafts directly inside Docs and Gmail where lawyers already work
- Free tier and broad availability
- Competent at outlines and client-facing email drafts
Where it falls short:
- Same grounding gap as the others; citations are guesses
- Quality can swing more across legal-specific prompts
Bottom line: Convenient if you live in Workspace. Same rule applies: drafting yes, authority no.
Spellbook
At a glance: ~$500/seat/mo, market estimate (Spellbook does not publish a price; 7-day trial) · sales · best for: transactional lawyers drafting contracts in Word.
This is the tool most often named when the search is specifically "AI for legal drafting." It lives inside Microsoft Word, suggests clause language, and benchmarks terms against its contract corpus, which is a real edge for transactional work.
What's good:
- Native Word add-in, so drafting happens where transactional lawyers already work
- Clause suggestions and benchmarking trained on a large contract corpus
- Strong fit for redlining and first-pass contract markup
Where it falls short:
- Built for transactional drafting, weak for litigation and brief work
- No public price, so budgeting means a sales call
- It can still insert wrong citations, so the verification duty does not disappear
What users say: Transactional lawyers rate it among the best-liked drafting add-ins (Lawyerist scores it around 4.1/5), with the recurring caveat that it sometimes inserts incorrect citations and is narrow outside contracts (Lawyerist Spellbook review).
Bottom line: A genuinely good pick if your day is contracts in Word. Look elsewhere for litigation drafting or open research.
CoCounsel
At a glance: ~$225 to $400+/user/mo, market estimate (Thomson Reuters quotes per firm) · sales · best for: firms that want drafting plus research on the Westlaw corpus.
CoCounsel sits on Thomson Reuters' research data, so unlike a raw chatbot it can ground answers in a real legal corpus. Thomson Reuters bought Casetext for a reported $650 million in 2023 to own exactly this grounding layer.
What's good:
- Grounded in the Westlaw corpus, so cites can trace to real authority
- Handles research, document review, and drafting in one assistant
- Familiar ChatGPT-style interface backed by a vetted database
Where it falls short:
- Priced and contracted for firms, not solo or lean in-house budgets
- Users warn it still fabricates citations and runs thin on appellate material
- Long contract commitments are common, so it is a heavier buy
What users say: A University of Michigan Law legal-tech evaluation found CoCounsel useful for speed and its Parallel Search, but warned it still fabricates citations and is thin on appellate coverage, with the blunt advice to verify everything (Michigan Law legal-tech series).
Bottom line: Worth it for a firm that needs grounded research and drafting together. Overkill, and overpriced, for a solo who mostly drafts.
Vaquill AI
Vaquill AI answers sit on retrieval over real US opinions and statutes, so each cite links back to a source you can open.
At a glance: Self-serve, published pricing (sign up to see the latest) · best for: solo GCs and in-house teams that need traceable answers.
This is the grounded side of the line, not a general chatbot. It retrieves over real US opinions and statutes so a cited claim links back to a source you can read.
What's good:
- Answers sit on retrieval over real opinions and statutes, so cites are traceable
- Document matrix handles extraction across many files, which a chat window cannot
- Self-serve pricing, far below the enterprise suites
Where it falls short:
- Grounding lowers hallucination risk but does not remove your duty to verify
- Built for in-house and solo workflows, not a giant litigation shop's stack
What users say: We are a newer entrant, so there is no large independent forum thread to quote yet. Rather than borrow someone else's praise, the honest pitch is the price and the traceable cites; check them against the alternatives above.
Bottom line: Pick it when you need answers a court can check, beyond clean prose. Verification still falls on you, but here it is a click, not a project.
Editing tools are a separate tier
If your search is really "polish my brief" rather than "draft it," a different class of tool fits: editing assistants like Clearbrief and BriefCatch that score and tighten prose, check citations against your own attached sources, and run inside Word or Outlook. They do not draft from scratch and they are not research engines. They sharpen text you wrote, which sidesteps the grounding problem entirely because the law in the document is already yours.
The single-document blind spot
One more thing the chatbot-comparison framing misses entirely. Most chat tools, generic or legal, are built around one conversation about one document. That is fine for "redline this NDA" or "summarize this opinion."
It falls apart on a task like this: across these forty leases, which ones have an assignment clause that triggers on a change of control.
That is not a writing problem. It is an extraction-across-a-set problem, and a chat window is the wrong shape for it. The answer is a tabular grid that pulls the same fields out of dozens of documents at once, what we call a document matrix.
Pair that with document comparison for redlines. Together they cover the part of legal work no single-thread chatbot was ever designed to do, however well it writes.
If your use case is contract review, that workflow matters more than prose quality. We go deep on it in the contract-review guide and a roundup of contract-review tools.
The market is converging on this
You can see the industry quietly admitting all of the above. Thomson Reuters did not pay a reported $650 million for Casetext in 2023 because Casetext wrote pretty paragraphs. It paid for grounded research infrastructure.
The AI-native peers, Harvey, Legora, CoCounsel, all market on grounding and citations, not on prose. (Their pricing reality is its own story; we break it down in Harvey vs Legora vs CoCounsel pricing.) Nobody serious is competing on "our chatbot writes the most elegant sentence." They are competing on "you can trust where our citations come from," because everyone learned, the hard and public way, that the sentence was never the problem.
So the best AI chatbot for legal writing, as a standalone question, has no good answer, because it measures the wrong thing. The chatbots are all good writers.
The question that separates a tool you can build a practice on from one that gets you sanctioned is simpler and older: can you check what it told you?
FAQ
Is any AI chatbot safe to cite in a legal brief?
No general chatbot is safe to cite from on its own. Even grounded legal tools hallucinate on a meaningful share of queries, so you verify every cite. ABA Formal Opinion 512 makes that verification a professional duty, not a preference.
ChatGPT or Claude for legal drafting?
Both write well at the same $20/mo consumer tier, so pick by feel. Claude handles long documents and tends to hedge; ChatGPT has the bigger ecosystem. Neither is grounded, so use either for scaffolding and supply the law yourself.
What makes a grounded legal tool different from a chatbot?
A grounded tool retrieves real opinions and statutes first, then answers on top of them with citations that link back. A raw chatbot predicts a citation-shaped string with no source behind it. The difference is whether you can check the answer in a click or have to rebuild the research from scratch.
What is the best AI for legal drafting?
For raw contract drafting in Word, Spellbook is the most-named specialist. For drafting tied to real authority you can cite, a grounded tool like CoCounsel or Vaquill AI fits better. For free first drafts you supply the law for, ChatGPT and Claude are fine. The right answer depends on whether the draft has to cite the law or just read well.
Can AI write legal documents and briefs?
Yes, AI can produce a clean first draft of a brief, motion, demand letter, or contract in seconds. It cannot reliably find or verify the authority that draft cites. You stay responsible for every cited proposition, so treat the output as scaffolding and source the law yourself.
What is the best free AI for legal writing?
ChatGPT, Claude, and Gemini all have free tiers that draft and edit well. None is grounded, so a free chatbot is safe for structure and tone but not for sourcing authority. If you need citations a court can check, that capability is not free; it requires retrieval over a real corpus.
What is the best AI for legal writing for a solo attorney?
A solo usually wants a self-serve price and traceable cites without a sales call. The enterprise suites (Harvey, CoCounsel) are priced for firms. A $20/mo chatbot covers drafting, and a self-serve grounded tool covers the cites you file. We go deeper in our solo-practitioner tools guide.
Do these tools keep client data confidential?
Consumer chatbot tiers are not built for legal confidentiality by default, and free tiers may train on your inputs unless you opt out. Legal-specific tools usually offer data-handling terms (no training on your data, a DPA). Confirm the terms before you paste a client matter; we cover where the data goes in where your legal AI data actually goes.
A closing takeaway
The chatbot you pick for first-draft scaffolding matters less than the workflow that catches a fabricated citation before it hits a filing.
For more on grounded legal research workflows, see /features/legal-research.
New legal AI guides, weekly.
Further Reading
Top Legal AI Agentic Suites Compared (2026)
Read postTop Contract Management Vendors for Fortune 500 Legal Teams
Read postWhat Is NDA Triage? How Lawyers Sort Inbound NDAs in Minutes With AI
Read postNDA Triage AI: How AI NDA Review Works and How to Evaluate a Tool (2026)
Read postIntelAgree vs DocuSign CLM (and Ironclad, ContractWorks): AI Contract Management Compared
Read postTop 16 AI Contract Review Tools (2026)
Read post
Product & Content
Legal AI suite for US working lawyers: research, drafting, document comparison, document matrix, matters, and citation-verified answers, in one tool.