Damien Charlotin's Hallucination Tracker, Read Like a Risk Manager

AI-hallucination sanctions escalating from $5K to $109.7K

The penalty tracks the response after a cite is flagged, not the fabrication itself.

The AI hallucination tracker is a public database that Damien Charlotin maintains of court cases where someone filed AI-fabricated content, usually invented citations. It lists 1,624 cases as of 18 June 2026 (damiencharlotin.com/hallucinations) and grows almost daily. Read as news it is a parade of embarrassing footnotes. Read as a risk-manager's logbook it tells you which tools fail, which filings fail first, who gets caught, and what turns a warning into a suspension.

TL;DR

Zachariah Crabill lost two years of his Colorado law license over a motion citing cases ChatGPT had invented. The fabrications were the smaller half of his problem. What turned the matter into a suspension was a single decision on the morning of the hearing: blame an unnamed intern instead of conceding the citations.

The AI hallucination tracker for legal cases that Damien Charlotin maintains (1,624 cases as of 18 June 2026) is full of arcs shaped like Crabill's, and the through-line a managing partner reading only the headlines will miss is exactly this one. Sanction severity in these cases is set far more by the seventy-two hours after a court flags a citation than by the original use of AI.

Treated as news, the tracker is a parade of embarrassing footnotes. Treated as a risk-manager's logbook, it explains which models fail, which practice areas fail first, which filer profile dominates the sanctions docket, and what tips a judge from a warning into a discipline referral.

Part of our legal AI verification and hallucination guide series.

For related verification / hallucination / vendor-trust coverage, see AI Hallucinations in Legal Research: How to Avoid Sanctions and How to Verify AI Legal Citations Before You File (ABA 512 Checklist).

Quick check

On the tracker's data, what sets sanction severity most?

The tracker is a public, hand-maintained database of court cases in which a lawyer, pro se litigant, or expert filed a document containing AI-generated hallucinations. Damien Charlotin, a researcher and lecturer at Sciences Po in Paris, started compiling it in early 2024 after Mata v. Avianca made the rounds.

Entries come from court filings (often via PACER), media coverage, and reader submissions. Each row carries nine fields: case caption, court and jurisdiction, date, the party using AI, the AI tool where known, the nature of the hallucination, the outcome or sanction, the monetary penalty, and links to the underlying report. The live version sits at damiencharlotin.com/hallucinations and listed 1,624 cases as of its 18 June 2026 update (Charlotin's own running count on the page), with new rows added nearly every day.

That count has climbed fast. A widely cited write-up logged 120 cases in May 2025, the page itself read 1,174 in April 2026, and it crossed 1,600 by mid-June 2026. The dataset now spans roughly a dozen countries, with United States filings the clear majority.

Two things make it useful. Charlotin keeps it narrow, counting only incidents tied to a specific court filing or sworn submission, which strips out the noise of "lawyer says dumb thing about AI on LinkedIn." And the dataset crosses jurisdictions and firm sizes, letting a reader compare a federal magistrate's reaction in Wyoming to a state appellate panel in California without retyping anything.

How to read the tracker as a risk manager

The table is searchable and sortable. Three passes give you most of the value:

  • Sort by Date, descending. The newest rows show what judges are sanctioning right now. The 16 and 17 June 2026 rows alone already include a bar referral over a nonexistent criminal appellate opinion.
  • Filter by the AI Tool column. This tells you whether the failures cluster in consumer chatbots or in named legal tools, which is the single fact a firm AI policy turns on.
  • Read the Outcome and Monetary Penalty columns together. The gap between a "no sanction, strong disappointment" row and a five-figure penalty is almost never the size of the fabrication. It is the lawyer's response, which is the pattern below.

One caution Charlotin himself flags: the tracker is a best-effort harvest, not a census, and the authoritative version of any decision is the order itself. Use it to spot patterns, then pull the order before you rely on a specific case.

Methodology

The dataset was sliced three ways. By court level: federal trial, federal appellate, state trial, state appellate, state bar discipline, tribal, immigration. By filer profile: pro se, solo, small firm under twenty lawyers, midsize firm, AmLaw 200, government, expert witness.

By sanction severity, scored on a four-step ladder: warning or denied motion, monetary sanction under $5,000, monetary sanction at or above $5,000 or a discipline referral, and suspension, disbarment, or pro hac vice revocation.

Grouping started from Charlotin's summary field; for any row above the lowest severity rung, the underlying order was pulled and read. Where the entry lacked a firm size, the firm's website headcount or NLJ 500 listing was used; where the filer was solo or small-firm, that was confirmed against state bar records.

Honest limits up front: not every row identifies the model, some pro se entries lack enough detail to tell why the litigant was self-represented, the "firm size" column is occasionally a judgment call, and the tracker is a best-effort harvest, not a census. Patterns below are directional.

With that said, roughly two-thirds of the identified-lawyer entries sit in federal court, the state-trial bucket is the next largest, immigration courts are a smaller but visibly growing category, and solo plus small-firm filers together account for the majority of identifiable lawyer rows.

Pattern one: the dominant tool is still general-purpose ChatGPT

Across the entries where a model is identified, OpenAI's ChatGPT (free or ChatGPT Plus) accounts for the clear majority. Google Gemini, Microsoft Copilot, and Anthropic's Claude show up, in smaller numbers. Purpose-built legal tools account for a meaningful minority, but a minority.

The dominant failure mode is not lawyers buying bad legal AI. It is lawyers using a general chatbot they already had on their phone for a task it was never built to do.

This matters for policy. A firm-wide rule that says "no ChatGPT for case research" is more defensible than one that tries to whitelist tools by name, because whitelisting becomes immediately outdated and gives associates a reason to argue at the edges. The rule that maps cleanly to the data is: research using consumer chatbots is prohibited; legal research happens in a tool whose vendor will sign a contract describing its retrieval corpus.

Pattern two: litigation motions are the failure surface

The tracker is dominated by motions. Motions to dismiss, motions in limine, oppositions to summary judgment, post-trial briefs, motions for reconsideration. Contracts, transactional opinions, demand letters, and policy memos barely appear.

The reason is mechanical: motions live or die on case law, and case law is what large language models are most willing to fabricate in confident, citation-shaped prose.

An "AI is fine for low-stakes work" policy fails on inspection. Low-stakes by client weight is not low-stakes by sanction exposure. The motion in a slip-and-fall case is filed in the same court, under the same Rule 11 standard, as the motion in a securities matter. Charlotin's data has plenty of small-stakes underlying disputes attached to large sanctions.

Pattern three: solo and small-firm dominance, with a long pro se tail

By filer profile, solo practitioners and firms under twenty lawyers are the largest group in the tracker, with pro se litigants close behind. Midsize and AmLaw 200 entries exist (Morgan and Morgan in Wadsworth v. Walmart is the headline) but are a minority.

Reading this as "BigLaw is safe" misses what the tracker measures. BigLaw has more internal review layers, mandatory partner sign-off on filings, and conflicts-and-research desks that catch errors before the brief leaves the building. Structural defense, not virtue.

The tracker captures incidents that reached a public order, and larger firms catch more of their own before they get there. The honest read for a solo or small firm is that the workflow is operating without those structural defenses, so a single missed check has a far higher chance of becoming a sanction.

Pattern four: sanctions scale with the response, not the original error

Reading the higher-severity rows closely reveals a pattern that the headlines miss. Judges who learn that a brief contains a fabricated citation are usually willing to treat the first incident as a competence failure and impose a modest sanction. What turns a modest sanction into a career-altering one is the response when the fabrication is raised.

The clearest articulation of this is still in Mata v. Avianca, Inc., 678 F. Supp. 3d 443 (S.D.N.Y. 2023). After Schwartz was given a chance to confirm the citations, he went back to ChatGPT for confirmation rather than to a database. Judge P. Kevin Castel wrote:

In researching and drafting court submissions, good lawyers appropriately obtain assistance from junior lawyers, law students, contract lawyers, legal encyclopedias and databases such as Westlaw and LexisNexis. Technological advances are commonplace and there is nothing inherently improper about using a reliable artificial intelligence tool for assistance. But existing rules impose a gatekeeping role on attorneys to ensure the accuracy of their filings.

The bad-faith finding came from what happened next, not from the initial use of ChatGPT. The same arc shows up in Wadsworth v. Walmart, Inc., 348 F.R.D. 489 (D. Wyo. 2025), where Judge Kelly H. Rankin tied the size of the sanction to the absence of a meaningful internal review at the firm rather than to the initial mistake by a single lawyer.

The California appellate order from October 2025 that produced a $10,000 fine emphasized that twenty-one of twenty-three quoted holdings were fabricated and that counsel did not concede the issue until the panel had walked through the citations on the record. The penalties keep climbing toward the top of the ladder: see the $110K Oregon AI hallucination sanction for the same arc at a far higher number.

The Crabill arc makes the cleanest single-docket version of the point. The underlying error was a motion to set aside a summary judgment supported by ChatGPT-fabricated citations, the kind of slip lower-severity tracker rows resolve with a warning or a small monetary sanction.

What turned it into a two-year Colorado suspension was the morning-of-hearing decision, captured in the disciplinary record, to blame an unnamed legal intern when the judge asked about the citations on the record. A reader who only saw the headline would assume the suspension came from the fabrication. The discipline order makes clear it came from the response.

A firm's sanctions exposure, on this data, is set less by whether someone will eventually file a bad brief (the tracker suggests this is close to inevitable at scale) and more by how the firm reacts in the seventy-two hours after a court raises the issue.

Pattern five: the trend line is rising, not falling

Charlotin's monthly intake has been climbing through 2024 and 2025 and into 2026. Single-digit new entries per month in early 2024 became double-digit months by mid-2025, with no quarter on record where new incidents fell meaningfully against the prior quarter.

Three plausible drivers are stacking. More lawyers are using AI at all. More judges are actively looking for fabricated citations as part of their pre-hearing review. And a small but growing share of opposing counsel are filing motions for sanctions specifically targeting AI use, which both increases the discovery rate and creates a record that goes into the tracker.

This rules out the "we are past the worst" reading that surfaced briefly after Mata. Risk models that assumed incidents would self-correct as awareness grew have been overtaken by the data.

The diagnostic three

Most of the useful exposure work for a risk manager is in three questions.

Workflow controls. What tools are approved for legal research, what tools are explicitly prohibited, and is there a verification checkpoint embedded in the filing workflow rather than left to lawyer discretion? Charlotin's data is overwhelming on the point that ad-hoc verification, treated as a professionalism norm rather than a workflow step, fails at scale.

ABA Formal Opinion 512, issued July 29, 2024 by the Standing Committee on Ethics and Professional Responsibility, reinforces this by treating supervisory responsibility under Model Rules 5.1 and 5.3 as the partner-level obligation that AI use specifically triggers.

Training and disclosure. Does every lawyer and paralegal who touches AI output know the local rule in every court where the firm files? AI disclosure orders are now standing rules in pockets of California federal court, Pennsylvania state court, the Western District of North Carolina's Charlotte Division, the Northern District of Texas, and a growing list of state appellate panels.

The tracker has several entries where the sanction was driven less by the fabrication itself than by the failure to comply with a local disclosure order the lawyer did not know existed.

Sanctions response plan. If a judge or opposing counsel raises a citation as fabricated, who at the firm gets the call within the hour, what does the holding response say, and what is the firm's standing position on candor versus litigation defense? This is the highest-leverage question of the three. The tracker's higher-severity rows show that the sanction's slope is almost always set here, not at the moment of the original mistake.

The verification gap underneath all of this

One mechanical fact sits underneath the patterns. Language models generate text by predicting the next token, with no internal index of cases and no native concept of verification. Retrieval-augmented systems improve on this by searching a real corpus first, but retrieval can miss, generation can mischaracterize what was retrieved, and fabricated and verified output appear in the same confident register.

A concrete failure pattern: a retrieval system pulls a real opinion that uses a phrase ("substantially similar") in a procedural context, and the generation layer drops that phrase into a substantive-law sentence and cites the opinion. Both the citation and the phrase are real. The proposition they support together is not.

That kind of fabrication does not look like a fake case number and is invisible to a "did this citation exist" check, which is why Stanford HAI's 2024 legal-RAG study found non-trivial hallucination rates even in purpose-built tools. The tracker's growth curve will not bend until the verification gap closes at the tool layer; firm-level policy has to assume it has not.

What to do this quarter

A short list, drawn straight from the patterns above. Ban general-purpose chatbots for legal research in writing, with a named partner-level approver for exceptions. Replace them with a legal-AI tool whose vendor will describe its retrieval corpus, show source text, and surface per-claim confidence.

Make verification a workflow step inside the brief-prep checklist, not a professionalism norm. Track local AI disclosure rules in every court where the firm files, and update the tracking quarterly.

Write the sanctions response plan before there is a sanctions event: who responds, the tone of the first filing, the firm's standing position on candor. Tabletop it the same way a firm tabletops a data breach.

This is regulatory hygiene other operationally serious professions absorbed for their own AI failures: medical-imaging AI after the first wave of false-negative incidents, algorithmic trading after the 2010 flash crash, autonomous driving after Uber's Tempe case.

Each profession absorbed it once an incident dataset reached critical mass. Charlotin's tracker, read structurally, is the legal industry's version of that dataset.

FAQ

What is the Charlotin hallucination database? It is a public, hand-maintained database of court cases worldwide where a lawyer, pro se litigant, or expert filed a document containing AI-generated hallucinations, usually fabricated citations. Damien Charlotin, a researcher and lecturer at Sciences Po, started it in early 2024 after Mata v. Avianca. It lives at damiencharlotin.com/hallucinations.

How many AI hallucination cases are in the tracker? The page listed 1,624 cases as of its 18 June 2026 update, per Charlotin's own running count. That figure is up from about 1,174 in April 2026 and grows almost daily, so check the live page for the current number before you cite it.

Is the AI hallucination tracker only about US cases? No. It is global and spans roughly a dozen countries, including the UK, Australia, Israel, and Brazil. United States filings are the clear majority, which is why a US risk manager can treat most of the dataset as relevant to their own courts.

What does the tracker count as a hallucination case? Only incidents tied to a specific court filing or sworn submission where AI produced hallucinated content that a court or tribunal addressed in more than a passing reference. It does not track every fake citation or every mention of AI in a filing, which keeps the dataset narrow and usable.

How do you use the legal AI hallucination cases tracker as a risk manager? Sort by date to see what judges are sanctioning now, filter the AI Tool column to see whether failures cluster in consumer chatbots or named legal tools, and read the Outcome and Monetary Penalty columns together to learn what escalates a warning into a suspension. Then pull the underlying order before relying on any single case.

What is the biggest pattern the tracker reveals? Sanction severity is driven by the response after a court flags a citation, not by the original use of AI. Judges often treat a first fabrication as a competence slip and impose a modest penalty; what produces suspensions is doubling down, blaming staff, or failing to concede once the court raises the issue.

Does the duty to verify still apply if a legal AI tool gave you the citation? Yes. Under Rule 11 and ABA Formal Opinion 512, the lawyer signing the filing owns the accuracy of every citation regardless of which tool produced it. A grounded tool lowers the error rate, but it does not move the duty to verify off the attorney.

Are purpose-built legal AI tools in the tracker too? Yes, in a meaningful minority of the rows where a tool is identified. General-purpose ChatGPT still accounts for the clear majority, but the presence of named legal tools is the reason verification has to stay a workflow step rather than a trust assumption.

Closing note

Charlotin's tracker, read the way this post reads it, is the best free risk-management artifact the legal-AI conversation has produced. For the rule-by-rule lens on the same patterns, the ABA Formal Opinion 512 explainer walks through Model Rules 1.1, 1.6, 3.3, 5.1, and 5.3 against this exact incident shape, and is the natural next read.

Closing the verification gap at the tool layer is what Vaquill AI's citation-verified research is built for: it grounds answers in a real corpus, shows the source text, and surfaces per-claim confidence so the check is a workflow step rather than a professionalism norm. You can see how grounded research works.

For more on grounded research that makes verification a workflow step, see /features/legal-research or /legal-api.

Legal AI that reads your documents and knows the law.
Ask a legal question, review a contract, or search thousands of your files. Every answer shows where it came from. 7-day free trial, no card.
Updated June 20, 202617 min read

New legal AI guides, weekly.

Arshita Anand

Arshita Anand

Co-Founder & CEO · Attorney

Arshita leads product and strategy at Vaquill, building the legal AI suite that solo, small-firm, and in-house US lawyers use to run a matter end to end.