Harvey AI Subprocessors: Where Your Client Data Actually Flows

Short answer: Harvey's published subprocessor list names ten entities. Six are infrastructure and model providers (Microsoft, OpenAI, Google Cloud, AWS, Anthropic, ElevenLabs), three are product-feature providers (RELX/LexisNexis for "Ask LexisNexis," SuSea/you.com and Parallel Web Systems for web browsing), and one is Harvey itself. Harvey enforces zero data retention with its model providers, does not train on customer data, and offers in-region processing. The real Harvey AI data flow question is not whether it trains on your data (it does not). It is how far your data travels down the subprocessor chain after it leaves Harvey.

A managing partner I know spent two weeks last quarter satisfying herself that Harvey does not train on her firm's data. She read the trust page, she read the security marketing, she confirmed the no-training language, and she signed.

The thing she never opened was the one document that actually answers the question she cared about: Harvey's subprocessor list. When I sent it to her, she read it twice and asked the same thing everyone asks. "Why is LexisNexis on here?"

That is the whole problem with how lawyers evaluate AI vendors in 2026. We fixate on Harvey subprocessors and data training as a binary ("do they train on my stuff, yes or no") and treat the answer as the privacy verdict. It is not.

The verdict lives one layer down, in the chain of companies your matter data touches after it leaves Harvey. Harvey, to its credit, publishes that chain. Most buyers never read it.

This post teaches you how to read it, walks through what Harvey actually discloses, and explains why the subprocessor list, not the headline trust claim, is the disclosure that matters.

TL;DR

  • The privacy question that matters for Harvey is not "does it train on my data" (it does not). It is "how far does my data travel after it leaves Harvey." The subprocessor list answers that.
  • Harvey's published list names ten entities. Beyond the expected cloud and model providers (Microsoft, OpenAI, Google Cloud, AWS, Anthropic, ElevenLabs), it names RELX/LexisNexis ("Ask LexisNexis"), SuSea/you.com, and Parallel Web Systems for web browsing and knowledge sources.
  • A subprocessor list is a data-flow map. Each name is a place your client's confidences can go, and each of those vendors has its own subprocessors. The chain is transitive.
  • Harvey's DPA commits to at least 30 days notice before a subprocessor change takes effect, if you subscribe to its notification list, and gives you 15 days to object on reasonable data-protection grounds. That window is shorter than most firms' vendor-review cycle, which is the catch worth understanding.
  • ABA Formal Opinion 512 ties your competence and confidentiality duties to actually understanding this flow. "I read the no-training clause" is not the same as competence here.
Quick check

How many entities does Harvey's published subprocessor list name?

Part of our legal AI verification and hallucination guide series.

For related verification / hallucination / vendor-trust coverage, see Harvey Subprocessors and Beyond: How to Read a Legal AI Vendor's Subprocessor List and "We Do Not Train on Your Data": How to Verify the Claim.

Start with what "subprocessor" actually means

A subprocessor is any third party your vendor hands your data to in order to deliver the service. When you upload a privileged memo to Harvey, Harvey is the processor. The foundation model that reads the memo, the cloud that stores it, the search provider that fetches a statute to answer your question: those are subprocessors. They process your data on Harvey's behalf, under Harvey's contract, but they are not Harvey.

This matters because your duty of confidentiality does not stop at the first vendor. Under the model rules and ABA Formal Opinion 512 (July 2024), a lawyer using a generative AI tool has to make reasonable efforts to understand where client information goes and who can access it.

The opinion is explicit that this extends to the vendor's own providers. You cannot discharge that duty by reading a marketing page that says "we don't train on your data." You discharge it by reading the document that names the third parties, which is the subprocessor list.

Think of the list as a map, not a legal formality. Every entry is a destination your data can reach. Your job reading it is to ask three questions of each name: what does this company do, what data of mine reaches it, and does it have its own subprocessors I should chase.

Most lawyers skip the list entirely because it looks like boilerplate. Harvey's is not boilerplate.

What Harvey actually discloses

Harvey legal AI platform homepage

Harvey publishes its subprocessor list at harvey.ai/legal/subprocessors (verified June 2026). It names ten entities, each with a stated purpose and processing location. Group them by function and the map gets readable fast.

The infrastructure and model layer is what you would expect from any AI-native legal product. Harvey lists each one as an "AI provider" or "cloud services provider," with processing in the USA, EU, Switzerland, and Australia (Anthropic in the USA, reachable via AWS Bedrock and Google Vertex for non-US customers):

  • Microsoft Corp (AI provider and cloud services provider; Azure)
  • OpenAI, LLC (AI provider; the GPT family, Harvey's original foundation-model partner)
  • Anthropic, PBC (AI provider; Claude models)
  • Google Cloud Platform and Amazon Web Services (AI provider and cloud; for non-US customers, the route to Anthropic via Vertex and Bedrock)
  • ElevenLabs Inc. (AI provider, including voice services)

None of that should surprise anyone who has read a single legal AI privacy post. The interesting part is the second group, the one my partner friend choked on:

  • RELX Inc., d.b.a. LexisNexis, listed for the "Ask LexisNexis" capability
  • SuSea, Inc. (you.com), for web browsing
  • Parallel Web Systems Inc., for Knowledge Sources and web browsing

Sit with the first one. RELX is the parent of LexisNexis, a direct competitor to Harvey in the broad legal-AI market and, more to the point, a company most firms treat as an arm's-length data adversary in contract negotiations. A Harvey query, depending on the workflow, can route to LexisNexis.

That is not a scandal. It is a product decision: Harvey integrated a LexisNexis lookup so users get authoritative legal content. But it is exactly the kind of fact that a subprocessor list surfaces and a trust page buries.

The web-browsing names matter for a different reason. you.com and Parallel are search and retrieval providers. When a tool can browse the open web to answer your question, the query (and sometimes the surrounding context) leaves the relatively contained Harvey-plus-foundation-model perimeter and fans outward to a general-purpose search vendor.

For a public-records research task that is fine. For a query that quotes a sealed settlement term, you want to know that browsing is in the picture and under what conditions it fires. Harvey's own web-search help article (last updated June 10, 2026) says web search is an optional feature a workspace admin must turn on, that "all Web Search queries are processed briefly and query content is deleted immediately after results are returned," and that "neither Parallel nor You.com store, train on, or allow human access to Harvey customer data." So browsing is off until someone enables it, and the two browsing vendors carry delete-after-processing obligations.

To be clear about the posture: per Harvey's security page, it "undergoes annual SOC 2 Type II and ISO 27001 audits" and "partners with top-tier security firms, including Schellman, NCC Group and Bishop Fox" for those audits and penetration tests. The same page states Harvey "requires Zero Data Retention (ZDR) by model providers" and "contractually prohibits model providers from training on customer data." Harvey also offers in-region processing in the EU and Switzerland or Australia for customers with data-localization needs. Its trust portal lives at trust.harvey.ai (gated behind a request flow, so no live link here).

The published list is genuinely more transparent than most. The point is not that Harvey is hiding something. The honest disclosure lives in the document nobody reads, and that document reveals destinations the marketing never mentions.

Loading diagram...

The transitive chain is the real exposure

Here is the part that separates a careful reviewer from a checkbox one. Every subprocessor on Harvey's list has its own subprocessors.

OpenAI runs on Azure and its own infrastructure. Anthropic runs on AWS and Google Cloud. you.com and Parallel have their own cloud and data vendors. So when you map Harvey's ten names, you have not finished the map. You have finished the first ring of it.

Your client's confidences can, in principle, traverse Harvey to OpenAI to Azure, or Harvey to Anthropic to AWS, or Harvey to you.com to whatever you.com uses. The chain is transitive, and the list you started with only shows you the first hop.

A March 2026 essay on dev.to, "The Invisible Third Party," makes this argument bluntly: subprocessor chains compound, each link adds surface, and standard notice windows are not built for the depth of the chain.

The same piece flags CLOUD Act exposure for data hosted on US Azure, AWS, or Google Cloud, meaning US legal process can in theory reach data on those platforms regardless of where the customer sits. You do not have to accept every claim in an opinion essay to take the structural point: the list is the start of due diligence, not the end of it.

This is not a reason to refuse AI. It is a reason to read down the chain at least one more level for any workflow that touches your most sensitive matters, and to decide which workflows you are comfortable letting browse the open web at all.

The 30-day notice clock, and why it is the catch

Harvey's Data Processing Addendum commits to giving customers notice "at least 30 days before the change takes effect" when it adds a subprocessor, provided you subscribe to its email notifications, and lets you object within 15 days "on reasonable grounds relating to the protection of DPA Data." On paper that sounds protective, and it is more than some vendors offer. Read it as an operating constraint, though, and the catch appears.

Thirty days is a fast clock for a law firm. If your firm has any real vendor-review process, security questionnaire, an InfoSec sign-off, a partner committee that meets monthly, then 30 days from "we are adding Vendor X" to "Vendor X is live in your data path" is not enough time to complete a genuine review.

Your practical options inside that window are usually: object and risk losing functionality, or accept by silence. Most firms accept by silence, because the alternative is friction with a tool the associates already depend on.

So the question to ask Harvey (and any vendor) in the DPA is not just "do you notify us." It is: what is the notice mechanism, can we designate a contact who actually receives it, and what are our rights if we object inside the 15-day window.

If objecting just means "you may terminate," that is a weak right when you are mid-matter and the tool is embedded in the workflow. None of this is unique to Harvey. It is the standard shape of the SaaS notice clause. Knowing the shape is the point.

How to read any subprocessor list in ten minutes

The skill generalizes. Next time you evaluate a legal AI vendor, do this:

  1. Find the list. It should be linked from the trust or security page and from the DPA. If a vendor will not publish one at all, that itself is a finding. Harvey publishes; many do not.
  2. Sort by function. Separate the infrastructure and model layer (expected) from everything else. The "everything else" bucket is where the surprises live: competitors, web-search vendors, analytics, support tooling.
  3. Flag the outbound-data names. Anything that browses the web, enriches data from external sources, or is itself a data broker deserves a second look, because those names fan your query outward.
  4. Chase one more hop for your most sensitive workflows. Pull the subprocessor lists of the names that matter (OpenAI, Anthropic, and the major clouds all publish theirs) so you understand the second ring.
  5. Read the notice clause for the window length, the mechanism, and your rights on objection.
  6. Match it to the matter. A public-records research query and a sealed-settlement analysis do not carry the same risk. Decide which workflows are allowed to reach which destinations.

Run that on Harvey and you come away genuinely informed, which is more than the no-training-clause reader can say.

For a cross-vendor view of how Harvey stacks up against Legora, CoCounsel, Lexis+ Protege, and Westlaw on data flow, see our companion map, where your legal AI data actually goes. That post answers "where does data go" across six vendors. This one teaches you to read the source document yourself, which is the durable skill.

FAQ

Who are Harvey AI's subprocessors? Harvey's published list names ten entities: Microsoft, OpenAI, Anthropic, Google Cloud, Amazon Web Services, and ElevenLabs on the model and infrastructure side; RELX/LexisNexis (for "Ask LexisNexis"), SuSea/you.com, and Parallel Web Systems (for web browsing) on the product-feature side; plus Harvey AI Corp itself. The current list lives at harvey.ai/legal/subprocessors.

Does Harvey AI train on my data? No. Harvey's security page states it "contractually prohibits model providers from training on customer data" and uses your data only to process your requests. The harder question is data flow, meaning which subprocessors your inputs reach and how long they hold the data, which the subprocessor list and DPA answer.

Where does Harvey AI store and process my data? Harvey lists processing locations of USA, EU, Switzerland, and Australia for its core providers, and offers in-region processing in the EU and Switzerland or Australia for customers with data-localization needs. Anthropic is listed in the USA, reachable via AWS Bedrock and Google Vertex for non-US customers.

Is Harvey AI secure for confidential client data? Harvey undergoes annual SOC 2 Type II and ISO 27001 audits and works with firms including Schellman, NCC Group, and Bishop Fox. It requires zero data retention from model providers and encrypts data in transit and at rest. Whether that is enough for a given matter depends on which workflows you allow to browse the open web and how far down the subprocessor chain you are willing to read.

Does Harvey use OpenAI, Anthropic, or LexisNexis? Yes to all three. OpenAI and Anthropic are listed as AI providers, and RELX/LexisNexis is listed for the "Ask LexisNexis" feature. The LexisNexis entry surprises buyers because Lexis is a Harvey competitor in the broader legal-AI market, but it is a product integration the list discloses honestly.

Does Harvey browse the web, and can I turn it off? Web search is optional and a workspace admin must enable it, per Harvey's web-search help article. When it runs, queries go to you.com or Parallel, query content is deleted immediately after results return, and both vendors are barred from storing, training on, or accessing Harvey customer data.

How much notice does Harvey give before adding a subprocessor? Harvey's DPA commits to at least 30 days notice before a subprocessor change takes effect, if you subscribe to its notification list, with a 15-day window to object on reasonable data-protection grounds. Subscribe a real person, not a shared inbox, so the alert does not get missed.

Where this leaves the buyer

The honest read on Harvey is that it has one of the better public privacy postures among the AI-native vendors and a subprocessor list more candid than most. The LexisNexis and web-browsing entries are not red flags; they are product integrations the list happens to be honest about. For the broader product picture beyond data flow, see our honest Harvey AI review.

The real lesson is methodological: stop treating "they don't train on my data" as the answer and start treating the subprocessor chain as the question.

One structural way to shrink the chain is to control more of the stack yourself. If your statutory and regulatory lookups run through an interface you operate rather than traversing a vendor's web-browsing subprocessors, that is one fewer destination on your map.

A scoped statutes-and-regulations API (U.S. Code, the CFR, and state codes) running over MCP on infrastructure you point at lets a statute lookup avoid fanning out through a third party's search vendor. Case-law grounding, for the record, sits inside the in-product research workflow (8M+ US court opinions), not a public API surface.

The takeaway is methodological. The fewer hands your client's data passes through, the shorter the list you have to defend to a client who asks, and the more useful a published subprocessor list becomes. The list is the disclosure. Read it.

Vaquill AI data processing and compliance posture for legal teams

For more on a research suite that publishes its data flow and subprocessors openly, see /features/legal-research.

Legal AI that reads your documents and knows the law.
Ask a legal question, review a contract, or search thousands of your files. Every answer shows where it came from. 7-day free trial, no card.
16 min read

New legal AI guides, weekly.

Vaquill AI

Vaquill AI

Product & Content

Legal AI suite for US working lawyers: research, drafting, document comparison, document matrix, matters, and citation-verified answers, in one tool.