
Retrieval first, then grounding, then verification. The model never answers from memory.
AI legal search vs keyword search: the short answer
An AI law search engine matches on meaning, so it finds the right opinion even when your words never appear in it. Keyword and Boolean search match on exact terms, so they win when you already know the statute, party, or phrase. Use semantic AI search to find a doctrine you cannot name yet. Use Boolean to pull every case once you know what to ask for. Serious research uses both, then verifies the output against the real source.
The gap is not small. In the Blair and Maron study (1985), experienced lawyers and paralegals who believed their Boolean searches had found 75% of the relevant documents had actually retrieved about 20% (David C. Blair and M. E. Maron, Communications of the ACM, March 1985). They were missing four out of five relevant cases and did not know it. That blind spot is the problem semantic search exists to fix.
Two tools, two jobs
You are researching a novel issue. You know roughly what happened to your client, but you do not know what the courts call it, so you type a few guesses into the search box and watch the results miss.
That guessing game is the exact moment an AI law search engine earns its keep, and it is also the reason so many lawyers now search for one.
The honest answer is not "AI good, keyword bad." Semantic AI search and Boolean keyword search solve different problems. One finds the line of reasoning when you do not know the magic words. The other returns every exact match when you already do.
Pick the wrong one and you either drown in noise or miss the case that decides your motion.
TL;DR
- An AI law search engine matches on meaning, so it surfaces relevant opinions even when your words never appear in them. Best for novel issues, doctrine you cannot name yet, and "is this still good law" questions.
- Boolean / keyword search matches on exact terms and operators. Best for a specific statute, party name, or term of art where you need every hit and zero hallucination risk.
- Fourth Amendment search cases (Terry, Katz, Carpenter) hold very different things under the same word "search." Keyword search makes you guess; semantic search follows the reasoning.
- For a precise statute hunt (42 U.S.C. 1983, 17 U.S.C. 107, 28 U.S.C. 1331), Boolean is faster and safer, as long as you avoid the bare-section-number trap.
- Whichever you use, verify the output. ABA Formal Opinion 512 (July 2024) puts the duty to check AI results and protect client confidences squarely on the lawyer.
In the Blair and Maron study, what share of relevant documents did lawyers actually retrieve?
Part of our legal AI verification and hallucination guide series.
For related verification / hallucination / vendor-trust coverage, see AI Case Law Search Explained: How Semantic Search Finds the Right Precedent and How AI Legal Research Works: RAG, Grounding, and Citations.
Why keyword search makes you guess
Keyword search is literal. It matches the characters you type, plus whatever stemming or synonyms the engine bolts on. That is a strength when your vocabulary is exact. It is a trap when it is not, because the law rarely files its best answers under the words a non-specialist would reach for.
Take the Fourth Amendment. The constitutional text turns on the word "search," and a cluster of Supreme Court cases interpret it: Terry v. Ohio, 392 U.S. 1 (1968); Katz v. United States, 389 U.S. 347 (1967); and Carpenter v. United States, 585 U.S. 296 (2018).
Same word, three very different holdings. Terry is stop-and-frisk. Katz gave us the reasonable-expectation-of-privacy test. Carpenter extended Fourth Amendment protection to cell-site location records.
Now imagine you are researching whether police needed a warrant to pull location data, and you search "cell phone." You will land on Carpenter. You will probably miss Katz, the case that built the entire framework Carpenter applies, because the 1967 opinion never says "cell phone." It could not. The phone in Katz was a glass booth on a sidewalk.
That is the failure mode. Keyword search rewards you for already knowing the vocabulary of the doctrine. When you do not, it quietly hides the cases you most need, and you do not see what you missed.
An AI law search engine works differently. It matches on meaning rather than exact characters, so a plain-English question like "did police need a warrant to get someone's phone location history" can surface the Katz line of reasoning even though your words never appear in the 1967 text. You describe the concept; it follows the doctrine. That is the core idea behind a modern legal research workflow: ask in plain English, get an answer grounded in real US opinions with citations you can open.
When Boolean wins
Flip the situation. You know exactly what you want. You need every case that cites a specific statute, or every opinion mentioning a named party, and you need a complete list with no model deciding what is "relevant" on your behalf. This is Boolean's home turf.
Say you are pulling civil-rights claims under 42 U.S.C. 1983 (often written "Section 1983"), or fair-use opinions under 17 U.S.C. 107, or anything resting on federal-question jurisdiction at 28 U.S.C. 1331.
A Boolean query returns the full set, deterministically, the same result every time you run it. No summarization, no paraphrase, no risk that the engine invented a holding. For a litigator building a citation table, that completeness is the whole point.
There is one trap worth flagging. Section numbers are ambiguous. A query for "1331" without the surrounding context (28 U.S.C., the section symbol, the right title) will drag in every opinion that happens to mention the number 1331, from page cites to docket numbers to dollar figures.
Quote the full citation, scope it to the right title, and you get a clean set. Skip that and you get noise. We cover the operator syntax, proximity, and wildcards in depth in the Boolean search developer guide.
So the rule of thumb:
- You know the words (statute, party, term of art, exact citation): reach for Boolean keyword search. Complete, repeatable, zero hallucination risk.
- You know the concept but not the words (novel issue, unnamed doctrine, "what do courts call this"): reach for the AI law search engine. It follows meaning, not characters.
Most real research uses both. Start semantic to find the doctrine, then go Boolean to pull every case in the line once you know what to ask for.
Semantic vs Boolean legal search: one query, two results
Take a real research question: did police need a warrant to get a suspect's historical cell-site location data? Here is what each engine returns when you run that same question.
Boolean query. You type "cell phone" /p warrant /p location. The engine returns every opinion where those terms sit within a paragraph of each other. You get Carpenter v. United States, 585 U.S. 296 (2018), at or near the top, because Carpenter uses all three words. You do not get Katz v. United States, 389 U.S. 347 (1967), the case that built the reasonable-expectation-of-privacy test Carpenter applies, because the 1967 opinion never says "cell phone." The phone in Katz was a glass booth on a sidewalk. The query cannot reach a concept its words do not name.
Semantic AI query. You type the plain-English question: "did police need a warrant to get someone's phone location history?" The engine matches on meaning, so it returns Carpenter and Katz and Smith v. Maryland, 442 U.S. 735 (1979), the third-party-doctrine case Carpenter limits. You get the whole line of reasoning. You describe the situation; the engine follows the doctrine.
The split is clean. Boolean gave you what you asked for. Semantic gave you what you needed. That is why a hybrid workflow beats either tool alone: semantic to surface the doctrine, Boolean to pull the complete set once the vocabulary is in hand.
The case Boolean can't tell you went stale
Here is a job neither pure keyword search nor a single semantic lookup handles on its own: telling you a case is no longer good law.
For forty years, Chevron v. NRDC, 467 U.S. 837 (1984), governed how courts deferred to agency interpretations of ambiguous statutes. It was one of the most-cited decisions in American law.
Then in 2024 the Supreme Court overruled it in Loper Bright v. Raimondo, 603 U.S. 369 (2024). A keyword search for "Chevron deference" will still return Chevron at the top of the list, ranked by relevance, with no flag that the doctrine it announced is dead.
A grounded AI law search engine can do better, because it reasons over the citation relationships, not just the text match. When the system retrieves Chevron, it can surface that Loper Bright overruled it and route you to the current standard.
That is the difference between an answer and a landmine.
It is also why grounding matters: the system checks the model's claims against real, current source passages instead of trusting the model's memory. We walk through that retrieval-and-grounding pipeline in how AI legal research works.
The same logic applies to statutes. Code does not sit still. Provisions get amended, repealed, and recodified, and the interpreting cases shift under them.
A US statutes and regulations layer that combines the U.S. Code and CFR with the opinions that interpret each section turns "what does 5 U.S.C. 706 require after Loper Bright" into one question, not a week of cross-referencing.
Where Agent Mode comes in
Semantic search answers one question well. Real research is a chain of them. You find the controlling case, then you need its progeny, then the circuit split, then whether your jurisdiction has adopted the minority view, then the statute it all hangs on. Each answer spawns the next query.
That is the gap an agent mode fills. Instead of you running each follow-up by hand, it runs the multi-step research autonomously: retrieve, read, decide what to check next, retrieve again, and assemble the result with citations attached.
Think of it as a junior associate who actually pulls every thread, except every claim traces back to a real opinion you can open and verify.
A worked example. Start with a plain-English question about an employer's vicarious liability for an off-duty employee. Agent Mode finds the governing standard, identifies the cases applying it in your circuit, checks whether any have been distinguished or overruled, and pulls the statutory hooks.
You get a research memo with the reasoning laid out and pin cites you can click, not a paragraph of confident prose you have to fact-check from scratch.
Semantic search is the entry point. Agent Mode is what turns a single good answer into finished research.
Trust: the part you cannot skip
None of this removes the lawyer's duty to verify. ABA Formal Opinion 512 (July 2024) is explicit: using generative AI does not change your obligations of competence, confidentiality, communication, candor, and supervision.
You have to understand the tool's limits, protect client confidences, and check the output before you rely on it. An AI law search engine that hands you uncheckable answers is a liability, not a feature.
This is why grounding is the whole ballgame. A serious answer layer is built with retrieval-augmented generation, not the model's training memory: the system fetches real opinions first, then answers from them, and attaches the citations so you can open the source and confirm the holding yourself.
A working corpus, for reference, is 8M+ US federal and state opinions plus the full U.S. Code and CFR. RAG plus open-the-source citations is what makes "verify the output" a thirty-second click instead of a research project.
For the hallucination side of this, including the sanctions that follow when nobody checks, see AI hallucinations and legal research sanctions.
A quick decision guide
When you reach for the search box, ask yourself one question: do I know the words, or just the concept?
| Situation | Reach for | Why |
|---|---|---|
| Specific statute, party, or exact citation | Boolean / keyword | Complete, repeatable, no hallucination risk |
| Novel issue you cannot name yet | AI semantic search | Matches meaning, surfaces unnamed doctrine |
| "Is this still good law?" | Grounded AI search | Reasons over citation relationships, flags overruled cases |
| Multi-step research chain | Agent Mode | Runs the follow-up queries autonomously, cites every step |
| Building a citation table | Boolean | Deterministic, returns every match |
The tools are not rivals. They are a workflow. Semantic search to find the doctrine when you are lost, Boolean to pull the complete set once you know the terms, grounding to keep the model honest, and Agent Mode to run the chain.
FAQ
What is the difference between an AI law search engine and keyword search? A keyword or Boolean search matches the exact characters you type, plus any synonyms the engine adds. An AI law search engine matches on meaning, so it can return a relevant opinion even when your words never appear in it. Keyword search rewards you for already knowing the doctrine's vocabulary. Semantic search lets you describe the situation in plain English and follows the reasoning.
Is semantic search better than Boolean search for legal research? Neither is strictly better. Semantic search wins on novel issues, unnamed doctrine, and "is this still good law" questions. Boolean wins when you need every case citing a specific statute or party and you want a complete, repeatable list with no model judging relevance for you. Most real research uses semantic to find the doctrine, then Boolean to pull the full set.
When should I still use Boolean keyword search? Use Boolean when you know the exact words: a specific statute (42 U.S.C. 1983), a named party, an exact citation, or a term of art. It returns the complete set deterministically, the same result every run, with no hallucination risk. That completeness is the point when you are building a citation table.
Can AI legal search tell me if a case has been overruled? A grounded AI search engine can, because it reasons over citation relationships, not just text matches. A plain keyword search for "Chevron deference" still returns Chevron v. NRDC at the top with no flag that Loper Bright v. Raimondo (2024) overruled it. A grounded system surfaces the overruling case and routes you to the current standard. You still verify before relying on it.
Does semantic search hallucinate case law? The retrieval step does not invent cases. It ranks real opinions from a real corpus by meaning. Hallucination risk comes from the generation step, when a model writes a summary from memory instead of from retrieved text. Grounding (retrieval-augmented generation) closes that gap by forcing the model to answer from fetched opinions and attach citations you can open and check.
Do I still have to verify AI legal search results? Yes. ABA Formal Opinion 512 (July 2024) is explicit that using generative AI does not change a lawyer's duties of competence, confidentiality, candor, and supervision. You have to understand the tool's limits, protect client confidences, and check the output before you rely on it. Grounding with open-the-source citations makes that check a thirty-second click.
What is semantic search in legal research? Semantic search turns your query and the opinions into numerical vectors that capture meaning, then ranks results by conceptual similarity rather than exact word overlap. So a search tool can treat "termination" and "dismissal" as related even though the characters differ. That is what lets you describe a fact pattern in plain English and still reach the right line of cases.
Find case law by the concept when you cannot name the keyword
Vaquill AI is a legal AI suite built for in-house counsel: plain-English legal research grounded in real US opinions, a US statutes and regulations layer, and agent mode for multi-step research, each answer carrying citations you can open. The honest limit: the public legal API covers US statutes and regulations only, while case-law research lives inside the product. Try a research question and check the citations yourself.
For more on AI-versus-Boolean search workflows, see AI Case Law Search Explained and the Boolean search developer guide.
New legal AI guides, weekly.
Further Reading
How to Verify AI Legal Citations Before You File (ABA 512 Checklist)
Read postABA Formal Opinion 512 (2024): Generative AI Duties for Lawyers
Read postAI Case Law Search Explained: How Semantic Search Finds the Right Precedent
Read postLegal Research With No Data Indexing or Human Review: A Confidentiality Checklist
Read postDoes OpenAI Train on Your Westlaw or LexisNexis Data?
Read post"We Do Not Train on Your Data": How to Verify the Claim
Read post
Product & Content
Legal AI suite for US working lawyers: research, drafting, document comparison, document matrix, matters, and citation-verified answers, in one tool.