How to Run a Legal AI Pilot in 30 Days (2026)

To run a legal AI pilot in 30 days: pick one high-volume workflow (start with NDA triage), lock your data posture and baseline metrics in week zero, route real work through one tool for two to three weeks, then make a go or no-go call against a written scorecard in week four. The rest of this guide is the week-by-week version, with the metrics, the scorecard template, and the failure modes that sink most pilots.

The fastest way to lose a quarter is to start a legal AI pilot the way most teams do: spin up trials of four tools, let everyone "play with it," and reconvene in ninety days to compare gut feelings. You end up with anecdotes, not evidence, and a GC who still cannot tell the CFO whether the spend paid off.

A legal AI pilot run that way is not a test. It is a vibe check with a budget line.

There is a better version, and it fits in thirty days. The constraint is the point: thirty days forces you to pick one workflow, measure one baseline, and make one decision.

This is the playbook a in-house team can run between everything else on the queue, built to produce a number you can defend in a budget meeting.

TL;DR

  • Pilot one high-volume workflow (NDA triage or first-pass contract review), not the whole suite. Breadth is how pilots die.
  • Write down your success metrics before day one: turnaround time, escape rate (errors that slip past the AI), and hours reclaimed. No baseline means no verdict, just impressions.
  • Set data posture and access controls up front, in week zero, so the pilot itself is not the thing that leaks a privileged document.
  • Run a tight week-by-week plan: week 1 setup and baseline, weeks 2 to 3 real work through the tool, week 4 measure and decide.
  • The common failure modes are boringly consistent: no baseline, too many tools at once, and no named owner. MIT found 95 percent of enterprise GenAI pilots produce no measurable business impact. Most of that is process, not the model.
  • End on a go or no-go call against a written scorecard, not a meeting where the loudest opinion wins.
4-question check
Question 1 of 4

What share of enterprise GenAI pilots did MIT NANDA find delivered no measurable P&L return?

How we built this playbook

We build a legal AI workbench, so treat this as interested but checkable: the method below is the one we walk in-house teams through, and every external number is named, dated, and linked so you can verify it yourself.

The structure pulls from three places: the way enterprise pilots actually fail (the MIT NANDA data below), the baseline-first ROI method that finance teams will actually accept (Clio and Axiom both publish versions of it), and the workflow-first sequencing that vendor pilot guides converge on. The composite pilot used as a worked example is illustrative, drawn from how these pilots tend to run on a two-to-three-lawyer team, not a single named client.

The headline number is worth sitting with. MIT's NANDA initiative, in its widely-cited mid-2025 report "The GenAI Divide: State of AI in Business," reported that roughly 95 percent of enterprise generative AI custom-tool pilots delivered no measurable P&L return, against the report's estimate of 30 to 40 billion dollars in enterprise GenAI spend.

That figure is cross-industry, not legal-specific, so read it as a base rate for how enterprise pilots behave, not a legal benchmark. Its diagnosis is not that the models are bad.

It is that organizations never closed the learning gap between a tool that demos well and a tool wired into how the work actually moves, and that generic tools stall precisely because they never adapt to a team's workflow.

For a legal department, that gap has three usual shapes, all visible before you open the trial.

No baseline. The cardinal sin. If you do not know how long an NDA takes you today, "the AI feels fast" is not a finding, just a pile of warm impressions no finance partner signs off on.

Pilots rarely fail because the tool was bad. They fail because nobody measured the before, so there is no after to compare it to.

Too many tools at once. Running four trials in parallel feels rigorous and is the opposite. You split thin attention four ways, none of the tools gets enough real volume to show its true behavior, and you cannot isolate which tool produced which result. A pilot is an experiment. Change one variable.

No owner. When the pilot belongs to "the team," it belongs to no one. Someone has to own the baseline, route the real work through the tool, log the misses, and own the decision. Without a name attached, the pilot evaporates the first week the queue spikes, which is every week.

A fourth, subtler killer: misaligned incentives. If the pilot lawyers bear all the visible risk of a bad AI call but get no credit for time saved, the rational move is quiet non-compliance.

They run the pilot on easy contracts and do the hard ones the old way, so your pilot measures the AI on a sample it will never see in production.

There is a category error sitting underneath most of these failures. In-house teams routinely line up a point contract-review AI, the AI module bolted onto a CLM like Ironclad or DocuSign, Microsoft 365 Copilot, and a legal-specific workbench in the same evaluation, then grade them on one rubric, as if they were the same product.

They are not. Copilot is a horizontal assistant that happens to open a Word document; the CLM add-on optimizes for getting signed paper into a repository, not for catching a weakened indemnity; a point contract-review tool does one job deeply; a workbench tries to chain research, drafting, and review.

Each is built for a different job, so a single pilot scored against a single workflow will flatter whichever tool happens to match that workflow and bury the rest. The fix is not to test fewer tools.

It is to be honest that you are testing one job, name the job, and accept that the winner of an NDA-triage pilot has earned the right to do NDA triage and nothing else yet.

The four categories, and the one job each is actually built for:

CategoryExample toolsThe job it is built forWhere it breaks in a pilot
Horizontal assistantMicrosoft 365 Copilot, ChatGPTGeneral drafting and summarizing inside OfficeNo legal playbook, no citations, no escape-rate discipline
CLM add-onIronclad AI, DocuSign IQMoving signed paper into a repositoryOptimized for throughput, not for catching a weakened indemnity
Point contract-review toolStandalone redline appsOne workflow (e.g. NDA triage) done deeplyStops at the edge of that one workflow
Legal workbenchHarvey, CoCounsel, LegoraChaining research, drafting, and review on your positionsBroadest to pilot; you must still name a single workflow to score

Score each against the one job it claims, not against each other. A CLM add-on that loses an NDA-triage bake-off has not failed; it was never trying to win that job.

Week zero: pick the workflow and lock the data posture

Before the clock starts, do two things. They take an afternoon and decide whether the rest of the month means anything.

Pick one workflow. For most in-house teams the right pick is the thing clogging the queue: inbound NDA triage or first-pass review of a recurring contract type (vendor MSAs, order forms, DPAs). Choose by volume and repetition.

You want a workflow that runs often enough to generate real data inside three weeks and is standardized enough that "right" and "wrong" are knowable. NDAs are the canonical starter: you see dozens, the failure modes are well understood, and a miss is recoverable.

Do not pilot on your most bespoke, highest-stakes M&A paper. You will not get enough reps, and the variance will drown the signal.

Lock the data posture. The step in-house teams skip and regret. The pilot moves privileged, PII-laden documents into a vendor's system, and privilege does not protect itself once the file leaves your control.

This is not just a procurement formality; it sits inside your ethics duties. ABA Model Rule 1.6 confidentiality obligations, and the Rule 1.1 duty of competence (whose technology-competence comment most US states have now adopted in some form), put the burden on counsel to make reasonable efforts to vet how a vendor stores, uses, and exposes client information before the data goes in, not after a breach.

Settle the posture before a single contract goes in, and get specifics in writing, not a "we take security seriously":

  • Are prompts and outputs used to train the vendor's or model provider's models? You want a flat no in the contract, not the marketing page.
  • Is there a zero-data-retention agreement with the model provider, and what is the vendor's own retention and deletion SLA?
  • Is data US-resident and US-processed? Is a BAA available if any contracts touch PHI?
  • SOC 2 Type II (the audited-over-time one, not a Type I snapshot), current sub-processor list, and notice before it changes.
  • SSO, SCIM, role-based access, and audit logs, so you can prove later who saw which pilot matter.

Set access controls so only named pilot participants see pilot matters, segregated from your live matter store. Our vendor security questionnaire for in-house counsel is built for this gate; run the data, access, AI, and sub-processor sections before the trial, not at renewal.

Be precise about the legal exposure rather than reaching for "waiver" as a scare word: feeding privileged or work-product material to a vendor whose terms permit broad disclosure, training use, or open-ended retention can create confidentiality, attorney-client privilege, work-product, and professional-responsibility risk, and exactly how much depends on the controls and contract terms you put in place first.

Define success before you start: the three metrics that count

Pick your numbers now, while you have no result to rationalize. Three metrics carry almost all the weight for an in-house pilot, and they fit on one card.

Turnaround time. Median time from "contract lands in queue" to "reviewed and routed." Measure it as wall-clock, not active-work minutes, because the business feels calendar days, not your billing increments.

Published benchmarks give a rough sanity range, but be precise about what they actually measured: vendor and consultancy figures on AI-assisted contract review and processing tend to cluster around a 25 to 60 percent reduction in review time, and most of those numbers are cross-functional or vendor-reported rather than audited legal-department results.

So treat the range as a sanity ceiling other teams have claimed, not a legal benchmark and not a promise. Your number, measured on your paper, is the only one that counts.

Escape rate. The metric teams forget and the one that protects you: the share of AI-reviewed contracts where a material issue slipped past, caught only on human spot-check. Define "material" before you start so the count is not argued after the fact.

On an NDA pilot, material usually means a confidentiality obligation that came back non-mutual, an over-broad residuals clause, a weakened carve-out, an auto-renewal you would have struck, or governing law or venue that moved. On a vendor MSA, add liability cap erosion, indemnity expansion, and a missing DPA reference.

A tool that is fast and wrong is worse than slow and right, because it manufactures false confidence.

Have a senior reviewer re-check a fixed sample (every fifth contract is a sensible cadence) at full depth and log every miss against that list. If a vendor flinches at the escape-rate question, that is your answer.

Hours reclaimed. Lawyer hours returned per week on the piloted workflow. This is the number the CFO buys. Be conservative and net: time saved minus time spent supervising the AI and fixing its misses.

Vendor surveys love to quote double-digit hours a week, and it may be real for some teams, but treat every external figure as marketing until your own log confirms it on your own paper. GC AI's December 2025 ROI study of its own customers, for one, claims an average of roughly 14 hours saved per week and about $252,000 in annual savings, built on a 14 percent cut in outside-counsel spend against a $1.8M median (GC AI, Dec 2025). Read that the way you would read any self-reported customer survey: a ceiling some teams hit under that vendor's own measurement, not a number to bank before your log earns it. The pilot exists to produce your number, not import someone else's.

A note on honesty: industry surveys in 2025 found fewer than one in five organizations measure AI return on investment at all, and most that do count usage and cost, not outcome. Measuring outcomes at all puts your pilot ahead of the field.

A 30-day legal AI pilot: scope, baseline, run, decide

Scope and success metrics are set in week zero; the go/no-go decision is made on evidence in week four.

The week-by-week plan

Week 1: setup and baseline

Spend the first week proving the present, not testing the future. Pull your last thirty to sixty days of the chosen workflow and compute the baseline: median turnaround, current escape rate (yes, your human process has one too, you have just never measured it), and hours the workflow consumes per week.

Non-negotiable. The baseline is the only thing that turns week 4 into evidence instead of opinion.

In parallel, stand up the tool with the week-zero data posture in place. Load your actual templates and standards so the AI reviews against your positions, not generic ones (the MIT finding bears repeating: generic tools stall because they never learn your workflow).

Encode your fallbacks and playbook now so the pilot tests the tool you would actually deploy. Pick the owner and the one or two lawyers who will run real work through it, and keep the circle small.

Weeks 2 to 3: real work, real volume

Route live contracts through the tool. Real inbound, not a curated test set. The lawyers still own the decision; the AI does the first pass. Every reviewer logs three things per contract: AI turnaround, whether they accepted, edited, or discarded the AI's read, and any miss the spot-check caught.

Two patterns recur often enough to plan around. First, reviewers under-log edits. They quietly fix the AI's missed carve-out and move on, so the log shows "accepted" when the truth is "accepted after I rewrote it," and your hours-reclaimed number inflates. The fix is a dropdown, not a free-text box: accepted clean, accepted with edits, discarded, with edits forced into their own category.

Second, the cherry-pick creeps in sideways. In a composite drawn from how these pilots tend to go (think a roughly 40-person SaaS company, a two-lawyer legal team, an NDA-and-order-form pilot at maybe 60 inbound a month), the tell was not lawyers refusing hard contracts outright; it was the messy reseller paper and the heavily-marked-up customer redlines quietly arriving "after" the pilot window, batched and run the old way, while only the clean counter-signatures hit the tool.

The intake log gave it away: pilot volume sagged on exactly the weeks the queue spiked, the inverse of what real adoption looks like. Catch it by pulling intake from the queue yourself rather than letting reviewers self-select what they feed the AI.

The single most useful thing a GC can ask in the week-four readout is not "is it good," it is "show me the three worst contracts it touched and what it missed on them," because the average case was never the question.

Resist adding a second tool "just to compare." You committed to one variable. If you genuinely need a head-to-head, run them sequentially on comparable volume, never tangled on the same week's queue. Two weeks of honest volume on one workflow tells you more than four tools dabbled with for a month.

Mind the sample size. Ten contracts is not a pilot, it is a coincidence; one good or bad run swings the whole number. This is the real reason to pick a high-volume workflow: an NDA stream that throws thirty to fifty contracts in three weeks gives you signal, while your quarterly enterprise MSA gives you three data points and a shrug.

The curated test set is the opposite trap: cherry-picked clean contracts make every tool look brilliant and tell you nothing about the messy paper in the real queue.

Week 4: measure and decide

Stop new intake mid-week and close the books. Compute the same three metrics you baselined and set them side by side. Turnaround: baseline versus pilot. Escape rate: human-only versus AI-assisted. Hours reclaimed: net, after supervision.

Then score it. Do not let the readout become a meeting where the most enthusiastic associate or the most skeptical partner sets the verdict by volume.

Make the math explicit, and be honest about what the sample can and cannot tell you. A worked example, round numbers a three-lawyer team might see: you baseline 42 NDAs over the prior month at a median 2.5 days turnaround, roughly 21 lawyer-hours a week, with a human escape rate near 7 percent on full re-review.

In weeks 2 and 3 you route 38 live NDAs through the tool, median turnaround drops to 0.75 days, and your every-fifth-contract spot-check (8 of the 38, re-reviewed at full depth) surfaces one material miss. That is one miss in a sample of 8, a 12.5 percent rate in the sample, which is your best estimate of the true escape rate across all 38, not a measured fact about them.

With a sample that small the confidence band is wide: one more miss in those 8 would read as 25 percent, zero would read as 0, so do not treat the point estimate as precise. The honest read is "estimated escape rate roughly 12 percent, single-digit sample, low confidence." Net hours reclaimed land around 9 a week.

Now you have a real decision: turnaround crushed it, hours are material, but the estimated escape rate sits above your human baseline on thin evidence. Not an automatic yes, and not a clean no either.

It is a reason to fix the playbook and re-run the spot-check at a heavier sampling cadence before you scale, because the cheapest thing in a pilot is one more week of re-review, and the most expensive is scaling a 12 percent miss rate you never confirmed.

The pilot scorecard

Score each line 0 to 2 (0 fails, 1 marginal, 2 clears). This is the artifact that survives the meeting. Copy the table straight into the pilot memo.

LineClears (2) whenScore
Turnaround vs baselineMedian cut by 25 percent or more (higher bar on NDAs)
Escape rate vs human baselineAI-assisted rate at or below your human number
Net hours reclaimedBeats the seat cost several times over at a loaded rate
Data posture in writingZDR, no training on your data, residency, BAA where requiredAutomatic no-go if zero
Unprompted adoptionLawyers keep using it without nagging
Self-serve priceNo six-figure committee to buy a productivity tool

The prose behind each line:

  • Turnaround beat baseline by a margin you would pay for. Faster at all is not the bar. A useful default is a 25 percent or greater median reduction (the low end of published benchmarks); on a high-repetition workflow like NDAs, set it higher. (0-2)
  • Escape rate at or below your human baseline. Score 2 only if the AI-assisted rate meets or beats your human number. More material issues slipping through means speed is a liability. (0-2)
  • Net hours reclaimed are real and material. Time saved minus supervision and rework. Convert to dollars at a loaded rate and require it to clear a threshold you set in advance (a simple one: the annual seat cost, several times over). (0-2)
  • Data posture cleared in writing. ZDR, no training on your data, residency, BAA where required. (0-2)
  • Lawyers would keep using it unprompted. Adoption that needs nagging does not survive a busy week. (0-2)
  • Price fits self-serve reality. In-house teams should not need a six-figure committee to buy a productivity tool. (0-2)

Add it up out of 12, and read the total against thresholds you wrote down before week 4, not after. A workable default scale: below 7 is a no-go, full stop, the tool did not earn the next workflow.

A 7 or 8 is a retune, not a verdict: you run exactly one fixed tuning cycle (tighten the playbook, reload templates, re-run a heavier spot-check sample) and re-score once, with no open-ended third try.

A 9 or higher is an expand, but only if the escape-rate line specifically cleared, because a high total carried by speed and adoption while material issues slip through is the most dangerous score on the card.

The go or no-go decision

A real decision is a threshold set in advance, not a feeling you arrive at. Set the bar before week 4 so you cannot move the goalposts to fit the result you want.

Loading diagram...

A clean go: the gate clears, the scorecard lands at 9 of 12 or higher, the escape-rate line cleared, and turnaround and net hours both beat baseline by a margin worth paying for. Expand to the next workflow; do not boil the ocean. Teams that scale AI well grow one proven workflow at a time.

A clean no-go: the gate fails, or the scorecard sits below 7 with marginal time savings and a worse escape rate. Not a wasted month. A pilot that produces a confident no with a documented reason is a successful pilot, and you now know what to demand from the next vendor.

The 7-to-8 band is its own answer, not a coin flip: run the single retune cycle, re-score, and let that second number decide. If it does not clear 9 after a focused tune, the tool is telling you it needs babysitting your team will not sustain.

The trap is the murky middle: decent numbers, no clear winner. Resist the urge to "extend the pilot to gather more data." Open-ended pilots are where the 95 percent go to die.

One honest exception: if your volume was genuinely too thin to read (a low-volume or seasonal department), a one-time, fixed extension to hit a real sample size is legitimate, as long as you set the new end date and threshold up front.

Absent that, if thirty days of real volume on your highest-frequency workflow did not produce a clear signal, the signal is weak, and weak signal on the easy workflow is a no-go on the hard ones. Decide, write down why, and move.

What this approach buys you

The discipline is the deliverable. A thirty-day pilot run this way produces three things a sprawling ninety-day trial rarely does: a defensible number for the budget conversation, a documented data posture you can point to if anyone challenges the privilege question, and a team that has used the tool on real work rather than poked at a demo.

Whichever way it lands, you come out with evidence instead of opinions, which is the whole point of a pilot rather than a launch.

If the workflow you pick is contract review, our in-house contract review playbook gives you the escalation tiers and fallback positions to baseline against, and the complete guide to legal AI for in-house counsel covers the broader buying criteria a pilot confirms. Once you have a go, the guide to rolling out legal AI to your team takes the proven workflow into wider use, the legal department KPIs to measure keep the gains visible past month one, and how AI is transforming in-house legal teams sets the wider context a single pilot sits inside.

FAQ

How long should a legal AI pilot last?

Thirty days of real volume on one workflow is enough for most in-house teams, and the constraint is a feature: it forces a single workflow, a single baseline, and a single decision. Outside guides land in a wider band (Clio and Axiom describe roughly eight to twelve weeks for a full ROI case), but that length is driven by tool count and scope. Narrow both and a month produces a defensible number. Extend only once, with a fixed new end date, and only if your volume was genuinely too thin to read.

What use case should an AI pilot for a legal team start with?

Start with the highest-volume, most standardized thing in your queue, which for most teams is inbound NDA triage or first-pass review of a recurring contract type such as vendor MSAs, order forms, or DPAs. You want enough reps to generate signal inside three weeks and a workflow where "right" and "wrong" are knowable. Do not pilot on bespoke, high-stakes M&A paper; the variance drowns the signal.

How do you measure the ROI of a legal AI pilot?

Baseline first, then compare three metrics: turnaround time, escape rate (material issues that slip past the AI), and net hours reclaimed after supervision and rework. Convert the hours to dollars at a loaded rate and set the dollar threshold before week four. This baseline-then-delta method is the one finance teams accept; Clio and Axiom both publish versions of it (Clio, 2026; Axiom, 2026).

How many AI tools should you pilot at once?

One. Running four trials in parallel splits your attention, denies each tool enough real volume to show its true behavior, and makes it impossible to isolate which tool produced which result. A pilot is an experiment, so change one variable. If you genuinely need a head-to-head, run the tools sequentially on comparable volume, never tangled on the same week's queue.

What metrics prove a legal AI pilot worked?

Three carry almost all the weight: median turnaround time from queue to routed, escape rate against your own human baseline, and net hours reclaimed per week. Speed alone is a trap, because a tool that is fast and wrong manufactures false confidence. Score escape rate as a gate, not a tiebreaker.

Why do most legal AI pilots fail?

The model is rarely the problem. Pilots fail on process: no baseline to compare against, too many tools at once, no named owner, and quiet cherry-picking where lawyers run the AI on easy contracts and do the hard ones the old way. MIT NANDA's mid-2025 report found roughly 95 percent of enterprise GenAI pilots delivered no measurable P&L return, and the diagnosis was the learning gap between a tool that demos well and one wired into real work (MIT NANDA, 2025).

Who should own a legal AI pilot?

One named person, not "the team." Someone has to own the baseline, route real work through the tool, log the misses, and own the go or no-go call. Without a name attached, the pilot evaporates the first week the queue spikes, which is every week.

Can you run a legal AI pilot in one week?

You can run a one-week smoke test, and some vendors structure their trials that way (GC AI publishes a day-by-day five-day script). It answers "does this tool basically work on my paper," which is worth knowing before you commit a month. What it cannot give you is a defensible number: one week rarely clears thirty contracts on a single workflow, and an escape rate estimated on a handful of spot-checks has a confidence band too wide to scale on. Use the week to disqualify obvious misses, then run the thirty days on the survivor.

How do you stop lawyers using unsanctioned AI during the pilot?

You mostly do not stop it by policy; you displace it by giving people a sanctioned tool that is faster than the consumer chatbot they are already pasting into. That is a real reason to sanction one vetted option early rather than leave the team in a no-approved-tool limbo, where the pressured move is shadow AI with no data posture behind it. Track whether unsanctioned use drops during the pilot as a secondary signal.

Run your pilot on a workbench built for it

Vaquill AI in-house workspace

Vaquill AI is an in-house legal workbench for contract review and drafting, self-serve with a 7-day trial and no procurement gauntlet, which is the self-serve posture a thirty-day pilot needs. Start your trial at app.vaquill.ai and see the in-house counsel solution for how the pieces fit together.

Legal AI that reads your documents and knows the law.
Ask a legal question, review a contract, or search thousands of your files. Every answer shows where it came from. 7-day free trial, no card.
Updated July 3, 202627 min read

New legal AI guides, weekly.

Arshita Anand

Arshita Anand

Co-Founder & CEO · Attorney

Arshita leads product and strategy at Vaquill, building the legal AI suite that solo, small-firm, and in-house US lawyers use to run a matter end to end.