Build vs Buy Legal Data: Where Your Product's Law Should Come From

Title card for the Vaquill AI guide: Build vs Buy Legal Data: Where Your Product's Law Should Come From

Short answer: there are two ways to get the laws and regulations your product needs. You can collect them yourself from the government sources that publish them. Or you can license them from a supplier. Each route has real costs and real risks, and they fall in different places. Collecting costs staff time and upkeep. Licensing costs fees, plus dependence on someone else's coverage and terms. Which one fits depends on how many places you need law from, how fresh it must be, and who will answer for mistakes.

TL;DR

  • "Legal data" here means the text of statutes (laws passed by legislatures) and regulations (rules written by agencies).
  • Collecting yourself means one source per publisher. States differ, and they change their sites.
  • Licensing means one source for many publishers. You then depend on what the supplier covers and on its contract terms.
  • Either way, someone must check the text against the official source and someone must read the license.
  • The table and questions below help you map your own situation. The choice stays with you.

Collecting yourself means your team gets law straight from the offices that publish it, cleans it up and stores it. Licensing means a supplier does that work and you pay for access, usually through an API (a way for software to ask for data). Many teams mix the two. They may use free official files for a few places and a supplier for the rest. Each part of a mix carries the costs below.

For background on what counts as legal data, see our buyer's guide to legal data.

What collecting it yourself involves

Many publishers, no common format. There is no single national source for state law. Each state runs its own site. Even inside one state, the same law can sit on more than one site. Texas has one statutes site where the text appears only after a browser runs extra page instructions, and a second site with plain pages of the same codes. As read on October 5, 2026, a simple download from the first returns an empty shell. Our guide to scraping state sites goes through more of these cases.

Rules about automated access. Many sites publish a file called robots.txt, which tells automated visitors what they may fetch. As read on October 5, 2026, the California Legislature's file disallows all paths for general visitors and asks for a 10-second wait between requests. The Arizona Legislature's file asks for 120 seconds. At Arizona's pace, 10,000 pages take about two weeks. Some offices offer bulk files (whole data sets you can download at once) instead, so check before you write a crawler, which is software that visits many pages automatically. Some sites also show challenge pages to automated visitors.

Where robots.txt came from. Martijn Koster defined the method in 1994. The Internet Engineering Task Force (the group that writes internet standards) only made it a formal standard in September 2022, in RFC 9309. So the web ran on habit for close to thirty years. The standard is blunt about what it is: "These rules are not a form of access authorization." A robots.txt file is a request about how to crawl. It does not tell you whether a site's terms of use allow automated copying, so read both. The standard also does not define a rule for wait times between requests, so a site that asks for 10 or 120 seconds is making its own ask.

Editions, renumbering and change. Some codes are republished once a year. Others change all session long. Sections get renumbered or moved. Others are repealed. A citation (the label used to find a law) that worked last year may point somewhere else now. Our Texas guide shows how one state's numbering moves. You need a way to notice changes, and a way to keep old versions if your users ask what the law said on a past date.

Verification and legal review. Someone has to compare what you stored with the official text. That means sampling, checking odd cases such as tables and repealed sections, and deciding who signs off. Some publishers call their web copy official and some do not. The answer changes how much weight your copy can carry.

Where "official" comes from. This question has a short history. In June 2007 Minnesota's Revisor of Statutes, Michele Timmons, who was also a Uniform Law Commissioner, asked the Uniform Law Commission to study how to prove that digital legal materials are authentic. A drafting committee followed in 2009. In July 2011 the commission approved the Uniform Electronic Legal Material Act, according to the law librarians' association's own history. Then each state had to adopt it separately. California's law was enacted on September 13, 2012 and took effect on July 1, 2015, and the association's list of enactments runs to roughly two dozen states and territories. Before you treat a state's web copy as the official text, ask its publisher what status the copy has.

Copying and sharing rights. The law itself is generally free to copy. The extras may not be. The extras are things like annotations, which are notes added to the text. In Georgia v. Public.Resource.Org, No. 18-1150 (2020), the Supreme Court held that Georgia's code annotations could not be copyrighted, because the state's own commission wrote them. That ruling covers those annotations only. It leaves each site's terms of use for you to read, state by state. If you show law to your own customers, check each source.

Staff time. The first load takes work. Ongoing monitoring, fixing parsers (small tools that turn a web page into stored text) when they break, answering user reports and reviewing changes take work as well.

What you gain: direct control, no supplier fee, and a record you can trace to the office that published it. What you take on: the upkeep and the review.

What people who have tried it say

Open States is the best-known volunteer project that gathers state legislative data. It covers bills, votes and legislators, not the codes themselves, but the work is the same kind. Its contributor guide says "each state requires several custom scrapers designed to extract bills, legislators, committees, and votes from the state website. All together there are around 200 scrapers, each one essentially independent." Its page on state quirks adds two details. A few states ask for an API key before they share their best data. And California publishes database dumps (whole data sets, packaged for download), which the project loads in place of scraping pages. That is the bulk-file route from earlier, working as hoped.

On Hacker News in November 2024, a thread about a research corpus of state statutes drew a comment from showerst, who said they had thought about starting a similar effort for the codes. In their words: "I've been thinking of spooling up an openstates-esque project to scrape all 50 into a common format via the original source documents, but it's a big undertaking." Another commenter, singleshot_, explained why codes are fiddly: "state legislatures pass acts, but we want to deal with codes. An act might change the law in several different “places.”" So someone has to apply each act to the code by hand and keep the numbering intact. A third commenter, tomwhipple, noted that in Minnesota a public office, the Office of the Revisor of Statutes, does that assembly work, and its output is public.

What licensing it from a supplier involves

Coverage and gaps. Suppliers differ on which states, which kinds of law and which years they hold. An empty result may mean the law does not exist, or may mean the supplier does not hold it. Ask for a coverage report you can read, by state and type of law.

License terms. Read what you may do: store the text, show it to your customers, feed it to an AI feature, pass it on or resell it. Check what happens to those rights if your business model changes.

Continuity. Your product depends on a feed you do not run. Ask about notice before changes, how long old versions of the connection point (the interface your software uses) keep working, what you can export, and what the contract says if the supplier is sold or stops.

Checking the supplier's claims. Do not rely on a number in a sales deck. Pick five sections you know well from different states. Compare each to the official site. Look for a source link and a retrieval date on each record. Our guide to legal data provenance (where each record came from) lays out a test, and the post on refresh cadence explains how to check freshness.

Fees and staff time. You pay the supplier, and prices differ by supplier and by use. Someone on your side still reviews the contract, tests samples, handles integration and watches for changes.

What you gain: one agreement and one format in place of many, and a party that carries the upkeep. What you take on: fees, and dependence on the supplier's coverage and terms.

A plain decision table

Your situationFactors that matter most
You need law from one or two statesHow each state publishes; whether bulk files exist; who checks monthly
You need law from many statesNumber of publishers; upkeep staff; supplier gaps; how coverage is reported
Your product shows law text to customersTerms on sharing the text; official status of the source; text accuracy review
You use law only for internal researchInternal-use rights; how current the text is
You send alerts when law changesChange detection; dates on records; past-date versions
You have no in-house legal reviewerWho verifies text; how errors get reported and fixed
Your customers audit your suppliersSource records; security documents; contract exit terms

A worked example, with no verdict. A compliance product needs statutes and regulations from California, Texas and New York, shows the text to its customers, and must keep past versions. Read down the table: three publishers, terms on sharing the text, and dated versions. Those are the factors to price on either route. Test one real section, such as Texas Government Code section 402.042, against the state's plain-page site. Compare the wording, the source link, the "current through" date (the day the publisher says the text is up to date) and whether an older version is available. One route prices them as staff hours and review time. The other prices them as a fee, contract terms and test samples.

Questions to answer before you decide

  1. Which states and kinds of law do you need today, and which in two years?
  2. Must you show the law to your own customers, or only use it inside?
  3. How current must it be: a day, a month, a year?
  4. Do you need what the law said on a past date?
  5. Who will check text against the official source, and how often?
  6. Who reads the terms of each source or the supplier's license?
  7. What would a wrong or missing section cost you or your customer?
  8. If your source disappeared, how long could you keep running?

For a longer list of supplier questions, see ten questions to ask a legal data vendor.

This guide is part of US Law Data: The Complete Guide, a map of where US law comes from and how to use it.

What would you do?
Question 1 of 4

You want every section of one state's code. Its legislature's robots.txt file (the note that tells automated visitors what they may fetch) asks for a two-minute wait between page requests. What is the sensible first move?

FAQ

Is it legal to copy state statutes? Statute text is generally public. Copyright can still apply to extras such as annotations, and a site's terms of use may limit automated access. Read the terms of each source, and ask a lawyer if your product depends on it.

How long does it take to collect law from all 50 states? It depends on the formats, the access rules and how many bulk files are offered. Plan for ongoing upkeep as well as a first load, because the sources keep changing.

How do I check a supplier's coverage claim? Ask for a coverage list by state and type of law. Then test five sections you know against the official sites. A good supplier shows a source link and a date on each record.

What happens if my supplier changes its terms or closes? That depends on the contract. Ask about notice periods, how long old interfaces keep working, and what you can export. Put the answers in writing before you sign.

Can I use both routes together? Yes. Some teams use official bulk files where a state offers them and a supplier for the rest. Each source then needs its own checks for freshness, terms of use and accuracy.

Which kinds of law are hardest to keep current? Sources that change often, such as agency rules, and sources that renumber sections. Check how often each source updates and how the publisher marks the date it is current through.

To see what a supplier's data looks like in practice, the Vaquill AI API for US statutes and regulations is one example you can test with the questions above.

Connect our US primary law database.
Every US statute, regulation, constitution, and executive order via REST, MCP or SQL. 5M+ sections, section-level citations, and links to the official source. Plus a free open dataset.
Updated October 5, 202614 min read

New legal AI guides, weekly.

Priyansh Khodiyar

Priyansh Khodiyar

Co-Founder & CTO

Priyansh leads engineering and AI at Vaquill AI: the pipelines that pull statutes, regulations and court rules from every US jurisdiction's official publisher, and the REST API, MCP server and open dataset that serve them.