Harvey Subprocessors and Beyond: How to Read a Legal AI Vendor's Subprocessor List

To vet a legal AI vendor, find its published subprocessor list, confirm it names the actual model provider, cloud, logging, and analytics companies (not categories), check the data residency and zero-data-retention posture of each, and confirm there is a dated change-notice mechanism. If the list is unpublished or names categories instead of companies, that is your answer.

TL;DR

  • A legal AI subprocessor list is the one document that tells you where your client's data physically goes after it leaves your vendor. Read it before you sign.
  • Most published lists are roughly 80% generic cloud and infrastructure plumbing and 20% legal-specific destinations. The legal-specific 20% is where the surprises live.
  • Sort every name into four types: model provider (reads your text), cloud/hosting (stores it), logging/observability (may retain it), and analytics/email (sees metadata). Each type carries a different risk.
  • Grade any list on 12 points (named LLM provider, per-vendor residency, sub-subprocessors, zero-data-retention, change notice, last-updated date) using the table below.
  • Harvey publishes a list of nine third parties plus itself (last updated April 10, 2026), with a per-vendor location column. That is more than most peers do. The full list is tabled below.
  • Harvey's zero-data-retention and no-training commitments do not live on the list. They live on a separate Subprocessor Update FAQ page. Read both, because the list alone does not carry that promise.
  • The five-question procurement email at the end sorts a vendor shortlist faster than any feature comparison.
4-question check
Question 1 of 4

How many third parties does Harvey's subprocessor list name, plus Harvey itself?

This is the general guide to reading any legal AI vendor's subprocessor list, with Harvey as the worked example. For the deep, Harvey-specific data-flow read, see Harvey AI Subprocessors: Where Your Client Data Actually Flows.

Part of our legal AI verification and hallucination guide series. For related vendor-contract coverage, see the DPA Review Field Guide for In-House Counsel, the Vendor Security Questionnaire Template, and Does Westlaw or LexisNexis Train OpenAI on Your Research?.

Privilege is easy to articulate and hard to defend after the fact. Under ABA Model Rule 1.6, a lawyer has a duty to make reasonable efforts to prevent inadvertent or unauthorized disclosure of information relating to the representation. That duty extends to the third parties your vendor uses.

When you upload a privileged memo to a legal AI tool, the foundation model that reads it, the cloud that stores it, and the search vendor that fetches a statute are all in scope. Each one is a place data sits, however briefly.

The stakes are not abstract. IBM's Cost of a Data Breach Report 2024 put the global average breach at $4.88M, and a leak through an unvetted fourth party is still your leak to disclose. A subprocessor you never mapped is a breach path you never priced.

The ABA made this national in July 2024. ABA Formal Opinion 512 says a lawyer using generative AI must understand the tool "to a reasonable degree," which includes how inputs are used and who else can access them. That is a direct instruction to read the subprocessor list, not the sales deck alone.

State bars say the same in their own words. The State Bar of Texas Professional Ethics Committee Opinion 705 (February 2024) treats generative AI use as triggering competence, confidentiality, and supervision duties, and expects lawyers to understand "the technology being used."

Florida Bar Ethics Opinion 24-1 (January 2024) requires lawyers to evaluate the AI provider's policies, including how the provider handles client data and which third parties touch it. California's Practical Guidance from the State Bar's Committee on Professional Responsibility and Conduct (COPRAC), issued November 2023, ties AI vendor diligence directly to the duty of confidentiality.

None of these opinions name a specific tool. All of them name the obligation to know where your data goes.

There is a contractual layer on top of the ethics one. If your firm has signed a CCPA-aligned DPA with a client, or a GDPR DPA with an EU-resident client, you have promised them visibility into your processors and subprocessors.

You cannot deliver that visibility if your own vendor will not deliver it to you. The subprocessor list is the artifact that makes that promise enforceable.

A well-maintained list lives at a stable, public URL and is linked from the vendor's trust, security, or legal pages. Check those three sections first. The common URL patterns are /legal/subprocessors, /subprocessors, /trust, or a "Subprocessors" link inside the privacy policy.

If you cannot find it in two minutes of clicking, search the vendor name plus "subprocessors" or "subprocessor list." If that still turns up nothing, the list is either unpublished or buried in a click-through DPA exhibit. Both are findings, not dead ends. Ask for it in writing and note how long it takes to arrive.

The format tells you something before you read a single name. A structured table on a stable page is a vendor that expects you to read it. A buried PDF, a paragraph inside the privacy policy, or "available on request" is a vendor that does not.

The four subprocessor types, and what each one risks

Every name on a legal AI vendor's list falls into one of four buckets. Sort them first, because the risk is different for each.

Where legal AI prompt data flows

Model provider (reads your text). This is the foundation-model company whose model actually reads your prompt: OpenAI, Anthropic, Google, Cohere, Mistral, or a self-hosted model. This is the loudest entry on the list because it is the one that processes the substance of your privileged input. The two questions that matter here are the data-retention posture (is it zero-data-retention) and the training posture (does the contract tier train on inputs).

Cloud and hosting (stores your text). The infrastructure your data sits on at rest: AWS, Azure, Google Cloud, Hetzner, and the like. The risk here is residency. AWS in us-east-1 is a different jurisdiction than AWS in eu-west-1, and "hosted in the US" at the top of the page does not tell you the region of each storage layer.

Logging and observability (may retain your text). Error tracking, request logging, and monitoring tools (Sentry, Datadog, and similar) often capture request and response payloads to debug failures. Even when the primary vendor deletes your data, its logging subprocessor may hold a copy for its own retention window. This is the destination buyers most often miss.

Analytics and email (sees your metadata). Product analytics, transactional email, and notification providers (Postmark, Resend, SendGrid, AWS SES, and analytics SDKs). Less sensitive than the model layer, but matter-related notifications, audit-log exports, and password resets can route subject lines and file names through these. The serious lists name them.

What Harvey actually publishes

Harvey legal AI homepage

Harvey publishes its subprocessor list at harvey.ai/legal/subprocessors (last updated April 10, 2026, verified June 2026). The page is short, which is a feature, and it names nine third-party entities plus Harvey itself, each with a stated purpose and a per-vendor location column, plus a sign-up to receive notice of new subprocessors.

Here is the full current list, tagged against the four types above so you can see the shape at a glance.

SubprocessorTypePurpose Harvey statesLocation(s) listed
Microsoft CorpModel + cloudAI provider and cloud services providerUSA, EU, Switzerland, Australia
OpenAI, LLCModelAI providerUSA, EU, Switzerland, Australia
Google Cloud PlatformModel + cloudAI providerUSA, EU, Switzerland, Australia
Amazon Web ServicesModel + cloudAI provider and cloud services providerUSA, EU, Switzerland, Australia
Anthropic, PBCModelAI providerUSA (non-US access via AWS Bedrock / Google Vertex)
ElevenLabs Inc.Model (voice)AI provider, including voice servicesUSA, EU
SuSea, Inc. (you.com)Product featureWeb browsingUSA
Parallel Web Systems Inc.Product featureKnowledge sources and web browsingUSA
RELX Inc., d.b.a. LexisNexisProduct featureAsk LexisNexisUSA
Harvey AI CorpThe vendorProvide the servicesUSA

Source: harvey.ai/legal/subprocessors, last updated April 10, 2026.

The expected names sit in the top block: Microsoft, OpenAI, Anthropic, Google Cloud, AWS, and ElevenLabs for voice. None of those should surprise anyone who has read a single AI privacy page. They are the model-provider and cloud layers.

The interesting names sit in the product-feature block. Harvey lists RELX Inc., d.b.a. LexisNexis for an "Ask LexisNexis" capability inside the product, SuSea, Inc. (you.com) for web browsing, and Parallel Web Systems Inc. for knowledge sources and web browsing. Harvey noted Parallel as a new web-Knowledge-Source subprocessor starting February 2, 2026.

The LexisNexis entry catches reviewers because LexisNexis is a direct competitor in the broader legal AI market. It is not a scandal. It is a product decision Harvey made to give users authoritative legal content inside the workflow, and the subprocessor list is the place that decision becomes visible.

Trust pages do not surface that. Subprocessor lists do.

Two things about this page deserve naming. First, the location column is real but coarse. It lists "USA, EU, Switzerland, Australia" for the big providers rather than pinning a single region per matter, so it tells you the possible jurisdictions, not the one your data actually lands in. That is a partial pass on residency, not a full one.

Second, the list page itself says nothing about zero data retention or training. That commitment lives one click away on Harvey's Subprocessor Update FAQ. There, describing its AWS and GCP model additions, Harvey states the data is "ephemerally" processed under zero data retention, with encryption in transit and at rest, no human review, no training use, and traffic encrypted using TLS 1.2 or higher. "Customer data is never used to train AI models unless explicitly authorized." That is the sentence you are actually buying, and it is not on the list you are told to read.

The honest read on Harvey: more transparent than the median vendor in the category, operating under zero-data-retention agreements with its foundation model providers, and the list still leaves the careful reviewer with one more layer of work. That is the right baseline to compare other vendors against.

The second-ring chase

Harvey names its first-ring providers but not OpenAI's, Anthropic's, or you.com's own subprocessors. The chain is transitive: OpenAI runs on Azure, Anthropic runs on AWS and Google Cloud, and each web-browsing partner runs on someone's cloud. You have to chase that second ring yourself.

Loading diagram...

The list itself is the test. If a vendor will not publish one, the test ends there. If they will, here is what to grade it on.

1. Is it actually published, or buried in a DPA? A subprocessor list linked from the trust page, the security page, and the DPA is a vendor that expects you to read it. A list available only on request, or only inside a click-through DPA exhibit, is friction by design. The friction itself is a signal.

2. Are the LLM providers named? Look for the actual model providers (OpenAI, Anthropic, Google, Cohere, Mistral, Meta) by name. "We use a large language model" or "we use leading foundation models" is not disclosure, it is marketing. If you cannot tell from the page which company's model reads your privileged memo, the page has failed its only job.

3. Is the data residency stated per subprocessor? "Hosted in the United States" at the top of the page is not enough. Each subprocessor has its own regions. AWS in us-east-1 is a different residency than AWS in eu-west-1. A list that names ten vendors but tells you the region of zero of them gives you a posture, not a fact. A list that gives you a multi-country column (as Harvey's does) is halfway there; the pinned region is the fact you still have to ask for.

4. Are the sub-sub-processors named, or at least pointed to? The chain is transitive. OpenAI runs on Azure. Anthropic runs on AWS and Google Cloud. Your vendor's web-browsing partner runs on someone's cloud. A list that links out to its named subprocessors' own subprocessor lists is doing the work for you. A list that names ten and leaves the second ring to you means at least one afternoon of clicking, and a vendor that has not made that easy has decided you will not do it.

5. Is the region of processing called out, not just storage? Storage at rest in the US is one fact. Processing in the US is a different fact. Inference can be routed to whichever model endpoint has capacity unless the vendor pins it. Ask whether the model call is region-locked or load-balanced across geographies, and look for that answer on the list.

6. Is the encryption posture per subprocessor disclosed? TLS 1.2 or 1.3 in transit, AES-256 at rest, managed KMS keys with automatic rotation. Some vendors offer bring-your-own-key (BYOK) for the storage layer. The list should at least gesture at the encryption tier of each major destination, especially the model providers.

7. Is there a last-updated date? A subprocessor list without a last-updated stamp is a document that may already be stale. The good ones are dated and ideally versioned. The best ones publish a changelog with the date a vendor was added or removed.

8. Is there a change-notification mechanism? RSS feed, email subscription, in-product banner, or webhook. The DPA usually commits to a notice window (often 30 days). The mechanism is how that window reaches the person inside your firm who can object inside it. If the only "notification" is an updated page that nobody monitors, the window is theatre.

9. Does the LLM provider have a zero-data-retention agreement? This is the single most important question for any model provider on the list. Both OpenAI and Anthropic offer enterprise customers zero-data-retention status, meaning prompts and outputs are not retained beyond the duration needed to return a response and are not used for any abuse-monitoring queue beyond what the contract specifies. Note that this promise often lives on a DPA or FAQ page, not the list itself, as Harvey's does.

Without this, your privileged inputs may sit in a 30-day model-provider log even if your direct vendor never sees them again.

10. Does the LLM provider train on inputs? Adjacent to the retention question but distinct from it. Enterprise tiers of OpenAI and Anthropic do not train on customer inputs. Consumer tiers do, by default. The fact that your vendor uses "OpenAI" tells you nothing on its own; the contract tier they use tells you everything. Ask for that to be on the page or in the DPA.

11. Is the backup and disaster-recovery vendor disclosed? Backups are a separate data path. A vendor whose primary storage is in one provider may run encrypted backups to a second provider, sometimes in a different region or geography. That backup destination is a subprocessor in every meaningful sense and belongs on the list.

12. Is the email and notification vendor disclosed? Less sensitive than the model layer, but firms have been surprised to discover that matter-related notifications, password resets, and audit log exports flow through a transactional email provider that copies subject lines into its own logs. The serious lists name the email vendor (Postmark, Resend, SendGrid, AWS SES) and what data crosses it.

A list that nails ten of those twelve is a vendor that has thought about this. A list that nails three is a vendor whose security page is mostly marketing.

The 12-point scoring table

Score each row pass or fail on the vendor's published list. The "why it matters" column is the one-line reason for a skeptical partner. The Harvey column scores its April 2026 list plus its Subprocessor Update FAQ.

#What to checkPass looks likeHarvey (Apr 2026)Why it matters
1Published, not buriedLinked from trust, security, and DPAPassFriction is a signal
2LLM provider named"OpenAI, Anthropic" not "leading models"PassTells you who reads the memo
3Per-subprocessor residencyRegion named for each vendorPartial (multi-country column)"Hosted in US" is not a region
4Sub-subprocessors pointed toLinks to each vendor's own listFailThe chain is transitive
5Processing region, not just storageInference region-locked, not load-balancedPartialStorage and processing differ
6Encryption posture per vendorTLS 1.2+, AES-256, KMS/BYOK notedPass (on FAQ page)Confirms the tier of each hop
7Last-updated dateA visible, recent datePassUndated means possibly stale
8Change-notice mechanismRSS, email, or webhook + notice windowPass (email sign-up)A snapshot is not a diff
9Zero-data-retention on LLMs"Under a ZDR agreement" statedPass (on FAQ page)Otherwise inputs sit in logs
10Training posture statedEnterprise tier, no training on inputsPass (on FAQ page)Tier, not brand, decides this
11Backup/DR vendor disclosedBackup destination namedFailBackups are a separate path
12Email/notification vendor disclosedPostmark/SES/etc. namedFailSubject lines leak metadata

What our subprocessor stack actually looks like

Vaquill AI data processing and compliance posture for legal teams

For a worked example of a "lawyer-readable" stack at the small-firm scale, here is what our own platform runs on.

We publish this to customers and their counsel on request, and the categories live on the security page.

  • Hetzner Cloud (United States, ISO 27001 certified) for application servers and compute. This is where the web app, API, and background workers run.
  • Supabase running on AWS for managed Postgres, authentication, and vault-style key management. AWS US region. Inherits AWS SOC 1 / SOC 2 / SOC 3 and ISO 27001.
  • Cloudflare R2 plus Cloudflare DNS and CDN (United States, SOC 2 Type II) for document object storage and edge delivery. R2 holds uploaded matter documents at rest.
  • Qdrant Cloud (United States) for the vector index used to retrieve relevant document chunks during research. Qdrant holds embeddings (numerical vectors) and metadata, not plaintext document content.
  • OpenAI (United States) for LLM inference on the GPT-5 family, under a zero-data-retention enterprise agreement: prompts and outputs are not retained beyond the response and are not used to train OpenAI models.
  • Anthropic (United States) for LLM inference on the Claude family as a fallback and for select workloads, also under a zero-data-retention agreement.
  • Dodo Payments (United States, PCI-DSS compliant) for USD billing. Card data never touches application servers; payment tokens do.

That is seven entities. It maps to the twelve-point checklist above as a useful comparison exercise: data residency stated, LLM provider named and ZDR posture disclosed, model layer training posture stated, storage and backup paths separated, payment data scoped out of the core perimeter.

The point of publishing it this way is that a careful reviewer can grade it in ten minutes, the same way you grade Harvey's.

Red flags in vendor subprocessor lists

Some patterns recur often enough across legal AI vendor lists that they deserve a name.

"We may use third-party services" without specifics. This is the most common evasion, and it is the easiest to call out. A list that names categories ("cloud hosting providers, foundation model providers, analytics") without naming the actual companies is not a subprocessor list. It is a placeholder.

Press for names in writing. If you do not get them, the vendor has decided you do not need to know, and you have decided whether that is acceptable for your matters.

The LLM provider is not named. This is the second-most-common evasion and the most consequential. If a vendor will not tell you whether your privileged input is being read by OpenAI, Anthropic, or a self-hosted model, you cannot evaluate their data flow at all. The model provider is the loudest disclosure on the list. Its absence is a finding.

No data-residency commitment. "Globally distributed" is not a residency commitment, it is the opposite of one. For US-based firms handling US matters, a US residency commitment, per subprocessor, in writing, is the minimum bar.

No notification mechanism on changes. A vendor that updates its subprocessor list silently is a vendor that has not built a process for telling you when your data path changed. The published list itself is a snapshot; what you actually need is the diff.

Zero-data-retention is missing for any LLM listed. If a model provider is on the list and neither the page nor the linked DPA says "under a zero-data-retention agreement," assume it is not. Ask. Get the answer in the DPA, not on a sales call.

A list that hits two or more of these is a list to push back on before signing.

The lawyer's evaluation checklist

AI vendor security questionnaire

Five questions to ask every legal AI vendor before the contract goes to the partner committee. They are short on purpose. You can run them in a single procurement email.

  1. Where is your published subprocessor list, and when was it last updated? If the answer is "we will send it to you," the list is not published. Note that.
  2. Which LLM provider reads my privileged input, on which enterprise tier, and is it under a zero-data-retention agreement? The provider name, the tier, and the ZDR status are three distinct facts and you want all three, in writing.
  3. What is the per-subprocessor data residency, and is processing region-locked, not just storage? Storage in the US with inference load-balanced to wherever the model has capacity is a different posture from end-to-end US processing.
  4. What is your notice window and notification mechanism for subprocessor changes, and what are my rights inside that window? The DPA usually says 30 days. The harder question is whether your firm's contact actually receives the notice, and whether your only right on objection is termination.
  5. Will you provide your subprocessors' own subprocessor lists, or links to them? This is the second-ring chase. A vendor that has the links ready saves you the afternoon. A vendor that does not has decided you will not do it.

Run those five at every legal AI vendor on your shortlist. The answers will sort the list faster than any feature comparison.

FAQ

What is a subprocessor in a legal AI vendor contract? A subprocessor is any third party your vendor hands your data to in order to deliver the service: the foundation-model company that reads your prompt, the cloud that stores your files, the logging tool that captures errors, the email provider that sends notices. If your vendor uses OpenAI's API, you are effectively using OpenAI, so your due diligence has to reach those fourth parties too.

How do I vet a legal AI vendor's subprocessor list? Find the published list, sort each name into model provider, cloud, logging, or analytics, then score it on the 12-point table above: named LLM provider, per-vendor residency, zero-data-retention posture, change-notice mechanism, and a last-updated date. Send the five-question procurement email at the end of this post to fill any gaps in writing.

Where do I find an AI vendor's subprocessor list? Check the vendor's trust, security, and legal pages first. Common URLs are /legal/subprocessors, /subprocessors, or a link inside the privacy policy. If two minutes of clicking turns up nothing, search the vendor name plus "subprocessors." An unpublished or request-only list is itself a finding.

Who are Harvey's subprocessors? As of its April 10, 2026 list, Harvey names Microsoft, OpenAI, Anthropic, Google Cloud, Amazon Web Services, and ElevenLabs (model and cloud providers), plus RELX/LexisNexis ("Ask LexisNexis"), SuSea/you.com (web browsing), and Parallel Web Systems (knowledge sources and web browsing), plus Harvey AI Corp itself. See the full table above, and check harvey.ai/legal/subprocessors for the current version. For the full data-flow read, see our Harvey data-flow deep dive.

Does Harvey have zero data retention? Harvey's Subprocessor Update FAQ states that its AWS and Google Cloud model additions process customer data "ephemerally" under zero data retention, with no human review and no training use, and that "customer data is never used to train AI models unless explicitly authorized." That commitment is on the FAQ page, not the subprocessor list itself, so capture both when you file your diligence.

Why is LexisNexis on Harvey's subprocessor list if it is a competitor? It powers an "Ask LexisNexis" feature inside Harvey, giving users authoritative legal content in the workflow. It is a product decision, not a scandal. The subprocessor list is simply the place that decision becomes visible, which is the whole point of reading one.

What is zero data retention and why does it matter for legal AI? Zero-data-retention (ZDR) means the model provider does not keep your prompts and outputs beyond the time needed to return a response. Without it, your privileged inputs can sit in a model-provider log (often for 30 days) even if your direct vendor never sees them again. Enterprise tiers of OpenAI and Anthropic offer ZDR; consumer tiers do not.

Is a SOC 2 report the same as a subprocessor list? No. A SOC 2 report attests to a vendor's internal controls. A subprocessor list names the outside companies that touch your data. You need both, and procurement teams routinely accept a SOC 2 badge while never requesting the list.

Do ethics rules require me to check subprocessors? Effectively yes. ABA Formal Opinion 512 (July 2024) tells lawyers to understand a generative AI tool to a reasonable degree, including how inputs are used and who can access them. State opinions in Texas (705), Florida (24-1), and California (COPRAC) tie AI vendor diligence to the duty of confidentiality. Reading the subprocessor list is how you meet that duty in practice.

How often does a subprocessor list change, and how will I know? Lists change every few months for active vendors, sometimes faster. Harvey, for example, added Parallel as a web-Knowledge-Source subprocessor starting February 2, 2026. The DPA usually commits to a notice window (often 30 days) before a new subprocessor goes live. Confirm the vendor has an actual mechanism (RSS, email, or webhook) that reaches the person in your firm who can object inside that window.

Read the list

The subprocessor list is the one document that connects your duty of confidentiality to the physical reality of where your data sits.

Harvey's is short, named, dated, and one of the better disclosures in the category, which is why it is worth reading even if Harvey is not on your shortlist. The same reading skill applies to every vendor you evaluate next.

For the parts of the chain you can control yourself, the simplest move is to shrink it. One way is to run statutory and regulatory lookups through an interface scoped to that purpose, which is one fewer destination on your map. Vaquill AI publishes its own subprocessor stack and runs a scoped statutes-and-regulations API (US Code, CFR, and 50-state codes), so you can grade us on the same twelve points you just used on Harvey. Case-law research is an in-app feature in modern suites like ours, not a public API.

The list is the disclosure. Read it before you sign.

For a research suite that publishes its data flow and subprocessors openly, see /features/legal-research.

Legal AI that reads your documents and knows the law.
Ask a legal question, review a contract, or search thousands of your files. Every answer shows where it came from. 7-day free trial, no card.
Updated July 3, 202626 min read

New legal AI guides, weekly.

Vaquill AI

Vaquill AI

Product & Content

Legal AI suite for US working lawyers: research, drafting, document comparison, document matrix, matters, and citation-verified answers, in one tool.