
Short answer: state legislature sites are built one at a time, so a scraper that works on one will break on the next. Expect different hosts, URL schemes, JavaScript rendering, session-law and codified layouts, yearly or two-year editions, and robots.txt files that range from open to closed with a long crawl delay. Before you write a parser, check whether the state offers a bulk download or an API. California, Illinois and New York each do, and Washington runs web services for bills and session law. For the rest, read robots.txt, identify yourself, throttle, cache, and store the source URL and edition with every section you save.
TL;DR
- There is no shared standard. Each state picks its own site, markup, URL pattern and edition schedule, and changes them without notice.
- Check for an official bulk file or API first. California publishes database dumps, Illinois exposes an FTP-style directory, and New York runs a keyed API. Washington runs free web services for bills and session law.
- A 200 response proves nothing. Texas's statute site returns the same 250,874-byte page for a statute and for robots.txt, because the text is drawn by JavaScript.
- Crawl delays are real: Arizona's robots.txt asks for 120 seconds between requests, California's asks for 10.
- Save the edition, the fetch time and the source URL with every section. Without them you cannot say which version of the law you hold.
You scrape a state statute site. Every request comes back with a success status, but your database fills up with empty sections. What is the most likely cause?
Part of our MCP and developer guide series.
If you want the federal side first, our US statutes API guide compares the official federal sources. This post covers the states.
Legal web scraping starts with the site's own rules
Before you fetch a page, read two things: the site's robots.txt and its terms of use. They differ a lot between legislatures.
- California. The leginfo.legislature.ca.gov robots.txt opens with
User-agent: *andDisallow: /, and it also setsCrawl-Delay: 10. The Legislature offers a different route for the same data, covered below. - Arizona. The azleg.gov file sets
Crawl-delay: 120for every user agent and disallows/search/and/xml/. Crawl-delay is not part of the robots standard, but well-run crawlers honour it. At one page every 120 seconds, 1,000 pages take about 33 hours and 10,000 take about 14 days. - Florida. The Online Sunshine site (leg.state.fl.us) disallows
/employees/for everyone and blocks two named bots. The Senate site (flsenate.gov) setsCrawl-delay: 10, disallows/api/, and lists/Laws/as allowed. - Minnesota. The revisor's file disallows a few search-result paths that generate wildcard parameters, and little else.
Terms of use are a separate question from robots.txt. Copyright is a third. The Supreme Court held in Georgia v. Public.Resource.Org (2020) that the annotations in the Official Code of Georgia Annotated cannot be copyrighted under the government edicts doctrine, because they were prepared by a contractor under the supervision of the state's Code Revision Commission, which the Court treated as their author. The case began when a nonprofit, Public.Resource.Org, posted Georgia's whole code online. After the state's Code Revision Commission sent cease-and-desist letters, it sued. Chief Justice Roberts, writing for the Court on 27 April 2020, put the principle in one line: "The animating principle behind this rule is that no one can own the law." (opinion on supremecourt.gov) That ruling is about annotations, and it leaves site terms and database rights to be read state by state. If your product depends on the answer, have a lawyer read the terms of each source.
Some sites also put a bot-detection layer in front of every page, including robots.txt. In one test from a single machine, a script that identified itself got a challenge page from nysenate.gov and block pages from ncleg.gov and legislature.mi.gov. Treat a response like that as the site saying "not this way." Look for an official feed, or write to the office that runs the site.
A short history of scraping state law
People have been scraping legislature sites for a long time, and the early mood was cheerful. In February 2009, Sunlight Labs posted "The Fifty State Project," a call for volunteers to build scrapers for every state. Its author, Josh Ruihley, framed the problem like this: "Data for state legislatures is available, but not yet accessible." He had tried Kentucky as a test, using Python and the Beautiful Soup library, and pulled every bill back to 2001 in about five hours. Then came the caveat: "not all states will be as simple as Kentucky." (The Fifty State Project)
The project grew into Open States. By its own account it was a Sunlight Foundation project from 2009 to 2016, ran independently from 2016 to 2021, and was adopted by Plural in 2021 (Open States, About). After Sunlight Labs closed, Miles Watkins wrote on 3 November 2016 that "maintaining the menagerie of scrapers into the future isn't easy either." (Why Open States?) Five hours for one state turned into a permanent maintenance job for fifty.
What developers say
The most useful voice I found is a Hacker News user, showerst, who says he has spent 15 years on legislative data and has made 2,400 commits to the Open States scrapers over nine years. He gave two answers that look like they disagree and are both right.
On bill tracking, in March 2026: "It doesn't scale like something like product or web search where you can just ignore broken pages, the penalty for missing things is too high." He also wrote that the sources "constantly change and break in goofy ways." (showerst, Hacker News, March 2026)
On the statutes themselves, in January 2026, he was calmer: "The nice thing about laws is that the host websites (or PDFs) don't change templates that often, so generally you can rescrape quarterly (or in some states, annually) without a ton of maintenance." (showerst, Hacker News, January 2026)
Put the two together and you get a useful rule. Bills move constantly, so a bill tracker has to run all the time and watch for breakage. A code changes slowly, so a statute fetcher can run on a schedule that matches the publisher's, and spend its effort on the checks in this post: did the text come back, which edition is it, and did the site's structure shift since last time.
What breaks, state by state
This table lists only what I fetched and read from each site.
| State | What I found | What it means for a scraper |
|---|---|---|
| California | robots.txt disallows / with a 10-second delay. downloads.leginfo.legislature.ca.gov lists pubinfo_YYYY.zip per session, daily zips and a readme. | The bulk files are the intended route. Tab-delimited .dat tables with .lob files for large records, including a LAW_SECTION_TBL for code sections. |
| Texas | statutes.capitol.texas.gov returns the same 250,874-byte app shell for a chapter URL and for robots.txt. The same chapter at tcss.legis.texas.gov/resources/GV/htm/GV.402.htm is one static HTML page with the text. | A plain fetch of the first host gets no law. The second host has the text, so the same chapter has two hosts and two behaviours. |
| Illinois | ilga.gov/ftp/ is a directory listing with separate ILCS and Public Acts folders. | Compiled statutes and enacting acts are different datasets. You need both to date a change. |
| New York | Open Legislation is a REST API with a free key. Its content types include NYS Laws, bills and committee agendas. | Use the API. The main site's front door is a challenge page for scripts. |
| Washington | Legislative Web Services are free and list services for legislation, documents, session law and "RCW cite affected". | Good for tracking which RCW sections a bill touches. The listed services cover bills and session law, not the code text. |
| Florida | The statutes site has a year selector from 1997 to 2025. | Each year is a separate edition. A section has one address per year. |
| Oregon | The ORS page is the "2025 Edition" and says it excludes the 2025 special session and the 2026 regular session. The 2027 Edition arrives online in early 2028. | An edition is a snapshot. The latest law sits in a different publication, Oregon Laws. |
| Minnesota | The revisor's page is titled "2025 MN Statutes" and notes that prior years are available. | Same pattern: key your records by edition year. |
| Arizona | robots.txt sets a 120-second crawl delay. | Fetching a whole code at that pace takes weeks. |
Eight things that break
1. The app shell. A server can answer 200 and send a page with no law in it. Texas shows it: the statute site draws content with JavaScript, and a plain HTTP client sees only the shell. Check for expected text in every response, such as the section number you asked for, and fail loudly when it is missing. A parser that finds nothing and reports success is the worst outcome.
2. Two hosts for one document. The same Texas chapter sits at two hosts with different behaviour. Store the exact URL you used, not just the citation.
3. Session law versus codified text. A legislature enacts session laws and then a revisor folds them into a code. Illinois publishes the two as separate folders. Washington's web services list session law next to "RCW cite affected". Oregon's edition leaves out the sessions that came after it. If you scrape only the code, you will be wrong about anything that changed since its last edition.
4. Editions. Florida offers a year selector, Minnesota names its year in the page title, Oregon publishes every two years. A section number is not a unique key. Use the number plus the edition.
5. Official versus unofficial. Oregon says its database is not the official text and the printed copy is. Georgia's official code is a published edition with annotations written under contract, which is why the Supreme Court case above exists. When it matters, say which source you hold and link it.
6. Robots files that your parser misreads. Florida's Senate file puts a blank line after User-agent: *, and it starts with a byte order mark. Python 3.12's standard urllib.robotparser read it as having no rules at all: no crawl delay, nothing disallowed. After I removed the blank lines and decoded the file as UTF-8 with the mark stripped, it returned a 10-second delay and blocked /Tracker/Login. Test your robots parser on real files.
7. Rate limits. Arizona's 120 seconds is an extreme case, but any state can add a limit or slow a response without telling you. Build the delay into the client.
8. Silent restructures. A site that adds a JavaScript layer or moves hosts does not announce it. Keep a canary for each source: a handful of known sections whose text hash you check on a schedule, with an alert when one comes back empty or very different.
Do it politely
The polite version of this job is short. It also keeps your project out of trouble with the offices you depend on.
- Read robots.txt first, and re-read it on a schedule.
- Identify yourself. Put a name and a contact address in the User-Agent so an office can write to you.
- Honour crawl delay and cap concurrency at one. Fall back to a couple of seconds when no delay is stated.
- Cache. Store the body and its
ETag, and send conditional requests so a page that has not changed costs the office almost nothing. - Prefer official bulk data and APIs. They are cheaper for you and for them.
- Save provenance. The URL, the edition and the time you fetched it go next to the text.
A small client that does these things:
import time
import urllib.robotparser
import requests
UA = "acme-statute-research/1.0 (+https://example.com/bot; legal-ops@example.com)"
def load_robots(base):
resp = requests.get(f"{base}/robots.txt", headers={"User-Agent": UA}, timeout=30)
text = resp.content.decode("utf-8-sig", errors="replace") # strip any BOM
rp = urllib.robotparser.RobotFileParser()
if resp.status_code == 200 and "<html" not in text[:500].lower():
# Python drops a rule group if a blank line follows its User-agent line.
rp.parse([line for line in text.splitlines() if line.strip()])
else:
rp.parse([]) # no usable robots file: apply your own limits
return rp
def polite_get(url, rp, last_hit, cache):
if not rp.can_fetch(UA, url):
raise PermissionError(f"robots.txt disallows {url}")
delay = rp.crawl_delay(UA) or 2
wait = last_hit["t"] + delay - time.time()
if wait > 0:
time.sleep(wait)
headers = {"User-Agent": UA}
if url in cache and cache[url]["etag"]:
headers["If-None-Match"] = cache[url]["etag"]
resp = requests.get(url, headers=headers, timeout=30)
last_hit["t"] = time.time()
if resp.status_code == 304:
return cache[url]["body"]
resp.raise_for_status()
cache[url] = {"etag": resp.headers.get("ETag", ""), "body": resp.text,
"fetched_at": time.strftime("%Y-%m-%dT%H:%M:%SZ", time.gmtime())}
return resp.text
Run against the real files, load_robots returns a crawl delay of 120 for azleg.gov, 10 for flsenate.gov and 10 for leginfo.legislature.ca.gov, where can_fetch is false for the code pages. After each fetch, check that the body contains the section number you requested before you save it.
What you are buying with a normalised API
A normalised API takes on the per-state work above and returns every state in one shape. The Vaquill AI API is one example. Here is a single request that covers California and Texas:
import os
import requests
resp = requests.post(
"https://api.vaquill.ai/api/v1/us/statutes/search",
headers={"Authorization": f"Bearer {os.environ['VAQUILL_API_KEY']}"},
json={
"query": "attorney general written opinion on question of law relating to official duties",
"corpusType": "STATE",
"state": ["ca", "tx"],
"limit": 4,
"fields": ["citation", "state", "sourceUrl", "lastAmendedYear"],
},
timeout=30,
)
for hit in resp.json()["results"]:
print(hit["actId"], hit["citation"], hit["sourceUrl"])
actId and citation always come back, whatever you list in fields. Two of the four results from a real run:
[
{
"actId": "STATE_CA_Cgov_T2_D3_P2_C6_A2_S12519",
"citation": "Cal. Gov't Code § 12519",
"state": "ca",
"lastAmendedYear": 2021,
"sourceUrl": "https://leginfo.legislature.ca.gov/faces/codes_displaySection.xhtml?lawCode=GOV§ionNum=12519."
},
{
"actId": "STATE_TX_Cgv_C402_S402.042",
"citation": "Tex. Government Code § 402.042",
"state": "tx",
"lastAmendedYear": 2013,
"sourceUrl": "https://tcss.legis.texas.gov/resources/GV/htm/GV.402.htm#402.042"
}
]
What that buys you, in concrete terms:
- One schema. The California section and the Texas section come back with the same fields. Your code does not branch on state.
- Stable ids.
actIdis the key. You take it from a response and use it to fetch the body at/us/statutes/section/{actId}/body. - Provenance.
sourceUrlpoints to the official page, so a reader can check the text against the legislature's own site. - Change tracking. The
changedSincefilter returns sections where a change was observed on or after a date, so a nightly job can ask what is new. It records when a change was seen, which is an upper bound on when it took effect.
Keep provenance visible in your own product too: show the source URL, the edition or amendment year, and the observed change date that your workflow relies on. With a normalised API, the work of fifty sites sits behind one contract. The cost to weigh against it is the full maintenance plan for your own fetchers: robots changes, edition boundaries, session-law backfills and monitoring, multiplied by the number of states you need.
If you are working on one state, our guides to California, Texas, Oregon and Illinois cover how each state numbers and publishes its law. For regulations, see state regulations through one API.
Sources read for this post
Every page below was fetched and read in October 2026, and the quotes and figures in this post come from them. The robots.txt files and the Texas response sizes are what those hosts returned to a script that identified itself.
- California:
leginfo.legislature.ca.gov/robots.txt,downloads.leginfo.legislature.ca.govand itspubinfo_Readme.pdf. - Texas:
statutes.capitol.texas.gov(including/robots.txtand/Docs/GV/htm/GV.402.htm) andtcss.legis.texas.gov/resources/GV/htm/GV.402.htm. - Arizona:
azleg.gov/robots.txt. - Florida:
leg.state.fl.us/robots.txt,flsenate.gov/robots.txtandleg.state.fl.us/statutes/index.cfm. - Minnesota:
revisor.mn.gov/robots.txtandrevisor.mn.gov/statutes/. - Illinois:
ilga.gov/ftp/. - New York:
legislation.nysenate.gov/static/docs/html/index.html. - Washington:
wslwebservices.leg.wa.gov. - Oregon:
oregonlegislature.gov/bills_laws/Pages/ORS.aspx. - Georgia v. Public.Resource.Org, No. 18-1150 (2020), syllabus on the Legal Information Institute site.
This guide is part of US Law Data: The Complete Guide, a map of where US law comes from and how to use it.
FAQ
Is it legal to scrape state statutes? State statute text is generally public, and the Supreme Court held in Georgia v. Public.Resource.Org that Georgia's code annotations cannot be copyrighted. Whether a site's terms of use restrict automated access is a separate question that varies by state. Read the terms of each source, honour robots.txt, and ask a lawyer if your product depends on it.
Which states offer bulk downloads or an API for statutes? From the sites I checked, California publishes PUBINFO database dumps, Illinois lists its compiled statutes in an FTP-style directory, and New York runs the Open Legislation API with a free key, which includes its laws. Washington runs free Legislative Web Services for bills, documents and session law, and for the RCW sections a bill affects. Always check the office's own page for the current offering.
Why does my scraper return an empty page from a state website? Many state sites draw their content with JavaScript, so a plain HTTP client gets only an app shell. Texas's statute site at statutes.capitol.texas.gov does this. The server still answers 200, which is why you need to check each response for the section text you asked for.
What does robots.txt crawl-delay mean? It is a request for the number of seconds to wait between requests. It is not part of the formal standard, but crawlers that mean to be polite follow it. Arizona's legislature asks for 120 seconds and California's asks for 10.
Why are statute numbers not enough to identify a section? Because states republish their codes on a schedule, a number can mean different text in different editions. Florida lists a separate statutes site for each year from 1997 to 2025, and Oregon publishes a new edition every two years. Save the edition year with every record.
Is the version on a state website the official law? Not always. Oregon says its online database is not the official text and the printed volume is. Check each source's own notice and record which one you hold.
How often should I re-fetch state statutes? Match the publication schedule. Editions change yearly or every two years, session laws change on the legislature's calendar, and some sites push updates daily. Use conditional requests so unchanged pages cost little, and keep a few canary sections to catch silent changes.
What does a legal data API do that a scraper does not? It returns every state in one schema with stable ids and a link back to the official page, and its maintainer handles the per-state fetching. Vaquill AI, for example, serves state codes behind one search call, and you keep the choice of which edition and which sources your product relies on.
New legal AI guides, weekly.
Further Reading
Attorney General Opinions by State: How to Find Them and What They Are Worth
Read postFederal vs State Law: Which One Applies to You?
Read postCalifornia AI Laws Explained: SB 53 and the Rest
Read postState Register vs. Administrative Code: What Each One Is
Read postNew York RAISE Act Explained: Who It Covers and When It Starts
Read postCalifornia Attorney General Opinions: How to Find, Read and Cite Them
Read post
Co-Founder & CTO
Priyansh leads engineering and AI at Vaquill AI: the pipelines that pull statutes, regulations and court rules from every US jurisdiction's official publisher, and the REST API, MCP server and open dataset that serve them.