To evaluate a legal AI vendor, score it across seven dimensions (accuracy and grounding, data handling and training, security and DPA, pricing and total cost, integration and workflow fit, support and onboarding, references and stability), then run a paid or free pilot on your own messiest documents. The demo is theater. The pilot is the only test that predicts production. The scorecard is further down this post.

TL;DR
Part of our in-house counsel guide series.
- A vendor demo runs on documents the vendor picked. It tells you the tool works on easy inputs, which you already assumed. It predicts nothing about your contract repository.
- The only evaluation that maps to production reality is a pilot on your own files, scored on one question: does the tool show a checkable source for every claim it makes?
- Score every vendor across seven dimensions and weight them to your risk. Accuracy and grounding plus data handling carry the most weight for in-house teams.
- Stanford researchers found leading legal research tools still hallucinated on a large share of queries. Source-checking is not optional. Bake it into the pilot.
- The red flags that should stop a deal: the vendor will not let you pilot, the tool cannot show its sources, the answers go vague on training, or the only option is an all-in annual lock-in.
Which dimension carries the most weight on the scorecard?
Why the demo lies and the pilot does not
A sales demo is a controlled environment. The vendor chose the documents, tuned the prompts, and rehearsed the flow. You are watching the tool's best day on its easiest inputs.
Your real work is not the demo. It is a 90-page master services agreement with two conflicting amendments. It is a scanned PDF from 2014 with a coffee stain on the indemnity clause.
The gap between those two worlds is where legal AI fails quietly. A tool that summarizes a clean template flawlessly can miss a carve-out buried in a messy redline. You will not see that in a demo. You will see it in month two, in front of a client.
The framework below has two halves. The seven dimensions tell you what to ask. The pilot tells you what is true. Run both, and weight the pilot heaviest.
The seven dimensions
Weight these to your own risk. A team handling PHI weights data and security higher. A team drowning in contract volume weights workflow fit higher. The dimensions stay the same. The weights move.
1. Accuracy and grounding
This is the dimension that decides whether the tool is usable at all. Two things matter, and they are different.
Accuracy is whether the answer is correct. Grounding is whether the tool shows you the source so you can check. A tool can be accurate and ungrounded, which still forces you to verify everything by hand.
The risk here is not theoretical. A 2024 Stanford RegLab and HAI study by Magesh, Surani, Dahl, and co-authors tested major legal research tools and found hallucination rates above 17 percent for some products and above 34 percent for another (Stanford HAI, 2024). They flagged two failure modes: answers that state the law wrong, and answers that cite sources which do not support the claim.
The second mode is the dangerous one. A confident answer with a real-looking citation that does not say what the tool claims. That is the case you sign off on and regret. Grounding is the defense, because it lets you catch the misgrounded cite before it ships.
ABA Formal Opinion 512 (July 2024) on generative AI reminds lawyers that the duty of competence does not transfer to the software. You stay responsible for verifying outputs. A tool that makes verification fast is doing real work. A tool that hides its sources is shifting risk onto you.
Ask: Does every output link to a checkable source? Can I see the exact passage it relied on? What is your own measured accuracy, and how did you measure it?
2. Data handling and training
The question is simple. What happens to the documents you upload, and does the vendor train models on them?
You want a contractual no, not a marketing no. "We do not train on your data" on a webpage is a press release. The same commitment in the DPA is a control. We cover that gap in our guide on verifying the no-training claim.
Legal AI tools route through third-party model providers. You need the named list and a zero-retention agreement with each one. For how the major vendors word this, see our Harvey subprocessors guide.
Ask: Do you train on customer data, in writing? Which LLM providers are in the chain? Do you hold zero-retention agreements with each?
3. Security and DPA
Security is where legal and IT meet. You want SOC 2 Type II as a floor and a signed DPA. Add US data residency if you need it, and a BAA if the deployment touches PHI.
This dimension has its own deep playbook. Rather than restate it, send the vendor our 47-question vendor security questionnaire and review the answers against the data residency and hosting commitments in our signed-DPA and US/EU hosting roundup.
Ask: Current SOC 2 Type II under NDA? Signed DPA with the subprocessor list attached? Where does the data physically live?
4. Pricing and total cost of ownership
The sticker price is the smallest number in the deal. Total cost includes implementation, training time, the seats you pay for but never activate, and the overage when usage spikes.
Per-seat pricing punishes teams with occasional users. Usage-based pricing punishes heavy ones. Neither is wrong. They fit different shapes. We break down the real numbers in our legal AI pricing benchmark.
Ask: What is the total annual cost at our seat count and usage? What is the overage rate? What happens to price at renewal?
5. Integration and workflow fit
A tool that lives in a separate tab gets used for a week, then forgotten. Fit means the tool meets your team where work already happens, in the document system, the email client, the matter folder.
Ask how it handles your actual file types. Scanned PDFs, marked-up Word docs, email threads with attachments. The tool that handles your mess beats the one with the prettiest interface.
Ask: Does it work inside our existing systems? How does it handle scanned and messy documents? What is the setup lift on day one?
6. Support and onboarding
A great tool with no onboarding becomes shelfware. You want a named onboarding contact, a clear ramp plan, and a support channel that answers in hours, not days.
Ask what onboarding actually looks like. A self-serve product should get a small team productive in a day. An enterprise product should assign a human. Either is fine. A vague answer is not.
Ask: Who owns our onboarding? What is the support SLA? How fast does a typical team reach steady use?
7. References and stability
You are buying a relationship. A vendor that folds in 18 months leaves you migrating data under pressure. References tell you whether the tool survives contact with real teams.
Ask for two reference customers who match your size and practice area. Then ask them the question vendors hate: what broke, and how fast did support fix it?
Ask: Two references like us? How long have you been operating, and who funds you? What is the data export process if we leave?
The evaluation flow: score the dimensions, then let the pilot decide.
How to run a real pilot
The pilot is the part most teams skip, and it is the part that matters. Here is how to run one that produces a decision, not a vibe.
Use your own documents, including the ugly ones. Pull 15 to 25 real files from your actual repository. Mix clean templates with messy redlines, scanned PDFs, and the deal folder everyone avoids. If you only test clean inputs, you only learn the tool handles clean inputs.
Build a grounded accuracy test. For each document, write down 3 to 5 questions you already know the answer to. The liability cap. The governing law. The auto-renewal date. The carve-outs to the indemnity. You know the truth, so you can grade the tool against it.
Grade on grounding, then on the answer. For every response, score two things. Was the answer correct? Did the tool show a source you could click and verify? An answer with no checkable source scores low even when it is right, because you cannot trust it at scale.
Score it the same way for every vendor. Use one scorecard, one set of documents, one set of questions, across every tool in the running. A pilot you score differently per vendor is not a comparison. It is a preference dressed up as data.
Set a time box. Two weeks is enough for most tools. Give each vendor the same window and the same files. If a vendor needs a month to look good, that is a finding.
The scorecard
Copy this into a sheet. One row per dimension, one column per vendor. Weight the dimensions to your risk, score each one to 10, multiply by the weight, and total it. The example weights below suit a typical in-house team. Move them to fit yours.
| Dimension | Weight | What a 9-10 looks like | What a 1-3 looks like |
|---|---|---|---|
| Accuracy and grounding | 25% | Checkable source on every answer, measured accuracy disclosed | No sources shown, accuracy "trust us" |
| Data handling and training | 20% | Contractual no-training, named LLMs, zero-retention agreements | Marketing-only no-training, no LLM list |
| Security and DPA | 15% | Current SOC 2 Type II, signed DPA, US residency, BAA if needed | SOC 2 "in progress," no signed DPA |
| Pricing and TCO | 10% | Clear total at your usage, predictable renewal | Opaque pricing, surprise overages, big renewal jump |
| Integration and workflow fit | 15% | Works in your systems, handles messy files | Separate tab, chokes on scans and redlines |
| Support and onboarding | 10% | Named contact, fast SLA, quick ramp | No owner, slow support, vague ramp |
| References and stability | 5% | Two matching references, stable funding, clean export | No references, unclear runway, locked-in data |
A worked example. Vendor A scores 9 on accuracy (weight 25), so 2.25 points. It scores 4 on pricing (weight 10), so 0.40 points. Sum the weighted scores across all seven rows for a total out of 10, then run the same math for every vendor. The highest total wins, unless it scored low on a dimension you cannot compromise on.
That last clause matters. A tool that aces six dimensions and fails data handling is the wrong pick for a team handling regulated data. Treat accuracy, grounding, and data handling as gates rather than weights. A failing grade on a gate ends the deal regardless of the total.
Red flags that should end the deal
Some answers are not negotiating positions. They are exits.
The vendor will not let you pilot. Covered above, and worth repeating. A product that works survives a test on your documents. One that needs a guided demo to look good will not survive month two.
The tool cannot show its sources. Ask "where did this come from." If the answer is a shrug or a regenerated summary, the tool is guessing with confidence. You will spend more time checking it than doing the work yourself.
The answers go vague on training. "We take privacy seriously" is not "we do not train on your data, see clause 7 of the DPA." Vagueness where the question was specific is the finding. Hold the deal until the commitment is in the contract.
The only option is an all-in annual lock-in. A confident vendor offers a pilot or a monthly start. A vendor that demands a year upfront, no trial, with a steep early-termination fee, is pricing in one risk: that you leave once you see the product.
The subprocessor list is hidden or stops at the first layer. You need the full chain, including the LLM providers. A vendor that will not disclose past its own infrastructure does not know where your data goes, or does not want you to.
The verdict
Run the seven dimensions to build your shortlist. Run the pilot to break the tie. Weight the pilot heaviest, because a demo measures the vendor's preparation and a pilot measures your reality.
One question separates a tool you can trust from one you cannot: does it show a checkable source for every claim? Everything else is negotiable. That is not. A team that buys on demos buys confidence. A team that buys on a grounded pilot buys a tool it can defend to a client, a partner, or a bar.
If you want a starting point for the pilot, Vaquill AI shows the source clause or statute behind every answer, so the grounded-accuracy test runs on day one. Score it the same way you score everyone else, and test the clause library against your own templates.
FAQ
How do you evaluate a legal AI vendor?
Score the vendor across seven dimensions: accuracy and grounding, data handling and training, security and DPA, pricing and total cost, integration and workflow fit, support and onboarding, and references and stability. Then run a pilot on your own documents. Weight the pilot heaviest, because it measures production reality and the demo does not.
Why is a pilot better than a demo?
A demo runs on documents the vendor picked and prompts the vendor tuned. It shows the tool's best day. A pilot runs on your real files, including the messy ones, and shows what happens when the tool meets your actual work. The gap between the two is where legal AI quietly fails.
What documents should I use in a legal AI pilot?
Use 15 to 25 real files from your own repository. Mix clean templates with messy redlines, scanned PDFs, and conflicting amendments. If you test only clean inputs, you only learn the tool handles clean inputs, which you already assumed.
How do I test legal AI accuracy?
For each pilot document, write 3 to 5 questions you already know the answer to, such as the liability cap or governing law. Grade the tool against the truth you know. Score two things: whether the answer was correct, and whether the tool showed a source you could click and verify.
What are the biggest red flags in a legal AI vendor?
The vendor will not let you pilot, the tool cannot show its sources, the answers go vague on whether it trains on your data, and the only option is an all-in annual lock-in with no trial. A hidden or first-layer-only subprocessor list is another. Each one should stop the deal.
Do legal AI tools still hallucinate?
Yes. A 2024 Stanford RegLab and HAI study found leading legal research tools hallucinated on a meaningful share of queries, including answers that cited sources which did not support the claim. That is why grounding matters: a checkable source for every answer lets you catch the bad cite before it ships.
How should I weight the evaluation dimensions?
Weight them to your risk. A team handling PHI or regulated data weights data handling and security higher. A team buried in contract volume weights workflow fit higher. Treat accuracy, grounding, and data handling as gates: a failing grade on any of them ends the deal regardless of the total score.
Should I sign an annual contract before testing a legal AI tool?
No. Start with a pilot or a monthly plan. A vendor confident in its product offers a hands-on trial. A vendor that demands a year upfront with no trial and a steep early-termination fee is pricing in the risk that you leave once you see the product.
Last updated: June 2026.
New legal AI guides, weekly.
Further Reading
How to Verify a Legal AI Accuracy Claim (Before You Trust the Number)
Read postWhat Lawyers Want From Legal AI in 2026: A 534-Lawyer Study
Read postAI Medical Record Review for Lawyers: How It Works, What It Costs, and Where It Breaks
Read postIs Legal AI a Bubble? What the 2026 Market Data Says
Read postWhat Lawyers Really Think of Legal AI in 2026 (Reddit + Reviews)
Read postThe Vaquill AI Legal Benchmark: 93% vs 33% on US Law That Changed This Year
Read post
Co-Founder & CEO · Attorney
Arshita leads product and strategy at Vaquill, building the legal AI suite that solo, small-firm, and in-house US lawyers use to run a matter end to end.