
Short answer: USLM (United States Legislative Markup) is the XML schema the US government uses for the US Code, bills, public laws and the Statutes at Large. Akoma Ntoso (also called LegalDocML) is the international OASIS standard it is related to. If your input is American law, you will almost always meet USLM. USLM writes the address of a unit as a path, like /us/usc/t5/s1/a. Akoma Ntoso addresses a document with a layered URI like /akn/un/act/2011-11-22/10/eng@2013-12-19/!main and addresses units inside it with an eId. Learn the identifier rules first and the rest of the markup follows.
TL;DR
- USLM is maintained by the Government Publishing Office (GPO) on GitHub. It models the US Code, bills, public laws and compilations.
- Akoma Ntoso is an OASIS standard built around the library idea of a Work, an Expression, a Manifestation and an Item. The EU's legislative drafting tool is built on it.
- GPO calls USLM a derivative of Akoma Ntoso. The USLM User Guide is more careful. It says the two are designed to be consistent and share many names, and it does not call one a subset of the other.
- For a parser, the real differences are the root elements, the namespace, and how each format builds identifiers. Key your records on the identifier. Tree position changes.
- A JSON statutes API can sit in front of all this. The section id is then your key, and you skip the XML.
A colleague says US legislative XML and a foreign parliament's XML are "basically the same format", so one parser will cover both. What is the sound response?
Part of our MCP and developer guide series.
If you would rather read the Code than parse it, start with the US statutes API guide. Statutes at Large lookups covers the session-law side, and pulling US Code sections by citation shows the request flow.
Where the two formats come from
Akoma Ntoso came first. It is an OASIS Standard, approved on 29 August 2018 as version 1.0 of the Akoma Ntoso core. OASIS says the vocabulary defines 310 element names and 69 attribute names, and the namespace is http://docs.oasis-open.org/legaldocml/ns/akn/3.0. It was designed so that parliaments and courts anywhere could mark up documents the same way.
USLM is the US answer. GPO publishes the schema on GitHub, and its README says the Legislative Branch XML Working Group manages it. When I checked on 4 October 2026, the README said the main branch carries versions 2.0.14 through 2.1.0. The version of any given document is written as an attribute on its root element, so you can read it from the file.
How closely are the two related? The sources differ a little, and the difference is useful to know.
GPO's Beta USLM XML page on govinfo says USLM is "a 2nd generation XML schema and is a derivative of the international LegalDocML (Akoma Ntoso) standard," and that it extends the international standard to handle US legislative and regulatory documents. That page is dated December 2018.
The USLM User Guide in the GPO repository is more careful. Section 1.7 says USLM is "not defined to be either a derivative or subset of Akoma Ntoso," but that it is designed to be consistent with it, and that many element and attribute names match. It adds that an Akoma Ntoso rendition of a USLM document should be possible through a simple transformation.
Treat them as cousins. You can reuse your habits from one when you read the other. You cannot feed one into a parser built for the other.
A short history: two standards that kept meeting in the same rooms
The two formats were never strangers. On April 23, 2012, the Library of Congress's In Custodia Legis blog reported that OASIS had formed a new technical committee, the LegalDocML TC, to move a common legal document standard forward, with the specification "based upon the Akoma Ntoso-UN project's XML schema." The post added that Akoma Ntoso was already in use in several parts of the world, "including Brazil, Africa, and the European Union," and that in the United States "Grant Vergottini has applied Akoma Ntoso to the US Code" (In Custodia Legis, April 2012).
A year later the Library ran a contest. On July 16, 2013, it opened a legislative data challenge asking competitors to mark up four US bills in Akoma Ntoso, working from the raw text rather than converting existing US bill XML. The stated aim was to "help identify gaps in the Akoma Ntoso schema and examples of US bill text data that cannot be incorporated properly into the framework." The prize was $5,000. The judges included Kirsten Gullickson of the House Clerk's office, who co-chaired the Legislative Branch XML Working Group, alongside Monica Palmirani and Fabio Vitali, the chairs of the OASIS committee (In Custodia Legis, July 2013). That working group is the one the USLM README says manages USLM today.
The standard itself took a few more years. The OASIS committee page records that Akoma Ntoso 1.0 was approved as a Committee Specification effective June 6, 2017, and that the work "is based on the Akoma Ntoso-UN project." The committee notes that the specification keeps the Akoma Ntoso name and the akomaNtoso root element (OASIS LegalDocML TC). OASIS Standard status followed in August 2018, as above.
So you can read the overlap as a family resemblance. Both projects were working on how to name a piece of law so it survives amendment, and people from both sat on the same panels.
Where you meet each one in the US
You will see USLM in four places.
The US Code. The Office of the Law Revision Counsel (OLRC) publishes the Code in XML at uscode.house.gov/download. The User Guide (read 4 October 2026) says there is one XML file for each title, plus a zip of the whole Code, and that XML has been available since July 30, 2013.
Bills and resolutions. The GPO repository holds sample bills in USLM, from introduced versions through enrolled ones.
Public laws and the Statutes at Large. GPO's govinfo page lists Beta USLM XML for enrolled bills and public and private laws from the 113th Congress forward, and for the Statutes at Large (at the full volume level) from the 108th Congress forward.
Statute compilations. These are current-law texts of single acts, and govinfo lists them in USLM too.
Akoma Ntoso is rare in American law. It matters when your product handles foreign legislation. One documented example: the European Commission's open source drafting tool, LEOS, stores its content in Akoma Ntoso XML and supports the AKN4EU subschema, which the Interoperable Europe site says is meant to help Member States and EU institutions exchange texts. That is the one deployment I can source, so I will leave the list there.
What USLM looks like
Here is a real piece of USLM. It comes from the statute compilation sample COMPS-10510.xml in the GPO repository (Public Law 113-5). I removed the style attributes, dropped the subsection heading, and shortened the text.
<section identifier="/us/sComp/113/5/s1" styleType="OLC">
<num value="1">SECTION 1. </num>
<heading>SHORT TITLE; TABLE OF CONTENTS. </heading>
<subsection identifier="/us/sComp/113/5/s1/a" styleType="OLC">
<num value="a">(a) </num>
<content>This Act may be cited as the "<shortTitle>...</shortTitle>".</content>
</subsection>
<subsection identifier="/us/sComp/113/5/s1/b" styleType="OLC">
<num value="b">(b) </num>
...
</subsection>
</section>
Four things to notice. Levels are plain elements named for what they are: section, subsection, and in the Code, chapter, title and paragraph. Each level has a num whose value attribute holds a normalized label such as a for the printed (a) . The heading and content elements carry the words. And identifier is a path that doubles as an address.
The root element tells you the document type. In the GPO sample files the roots are statuteCompilation, pLaw (a public law) and resolution, and they declare the namespace http://schemas.gpo.gov/xml/uslm. The User Guide (section 1.5.2.2) names a different URL, http://xml.house.gov/schemas/uslm/1.0, and that guide text is older than the samples. Read the namespace from the file you received and do not hard-code either one.
Links inside the text use the same path language. This one is from the Public Law 115-11 sample:
<ref href="/us/bill/115/hjres/37">H.J. Res. 37</ref>
<ref href="/us/fr/81/58562">81 Fed. Reg. 58562</ref>
How a USLM identifier is built
The User Guide gives the reference format as item, work, language, portion, temporal and manifestation, and it borrows the ideas from the library model called FRBR. For most work you need three pieces.
- The work. It starts with the jurisdiction (
/us) and then the document. For the Code, that is/us/usc/t5for Title 5. - The portion. This extends the path to a unit inside the document:
/s1/afor subsection (a) of section 1. - The date. An
@and an ISO date ask for a past version, so/us/usc/t5/s1/a@2013-05-02is the subsection as it stood on May 2, 2013.
A trailing extension picks the format: .htm for HTML, .xml for XML, .pdf for PDF. The guide's example /us/usc/t5/s1/a.htm is the current subsection as HTML.
The guide also names four attributes that look alike and behave differently:
| Attribute | What it is | Does it change? |
|---|---|---|
@id | A globally unique handle, a GUID with an id prefix | No. It follows the element when it moves |
@temporalId | A readable name like s1201_a_1_A, built from where the element sits now | Yes. It is recomputed when the text is renumbered |
@name | A local name inside a parent, often parameterised with {num} | Yes, with the number |
@identifier | The URL-style path, set on the root and usually on every level | Reflects the current address |
The @id rule matters most. The guide says it must not reflect anything that can change, such as a section number, because numbers are renumbered. @identifier and @temporalId are addresses. @id is a handle. When you store records, know which one you are storing.
What Akoma Ntoso looks like
Akoma Ntoso borrows the library model more openly. The OASIS specification defines four levels, and each one gets its own address:
- Work: the abstract legal act, for example act 3 of 2005.
- Expression: a version of the work, such as a language or the text after an amendment.
- Manifestation: a format of an expression, such as XML or PDF.
- Item: the physical file.
The spec's own example gives one act three addresses, all from its FRBRWork, FRBRExpression and FRBRManifestation blocks:
<FRBRWork>
<FRBRthis value="/akn/un/act/2011-11-22/10/!main"/>
<FRBRuri value="/akn/un/act/2011-11-22/10"/>
</FRBRWork>
<FRBRExpression>
<FRBRthis value="/akn/un/act/2011-11-22/10/eng@2013-12-19/!main"/>
</FRBRExpression>
<FRBRManifestation>
<FRBRthis value="/akn/un/act/2011-11-22/10/eng@2013-12-19/!main.xml"/>
</FRBRManifestation>
Read the work address left to right: un is the jurisdiction token in the spec's example, act is the document type, 2011-11-22 is its date and 10 is its number. The expression adds the language and the version date (eng@2013-12-19). The manifestation adds the file extension. The !main marks the main component of the document, as opposed to an annex.
Inside the document, elements carry an eId that spells out the path with double underscores, such as sec_1__list_1__point_a. The spec also keeps a separate stable id for the work level, the wId, and a mappings block that records how an original wId maps to the current eId when numbers change. That is the same split USLM makes between @id and @temporalId.
What differs for a parser
Most of the differences are mechanical.
| Question | USLM | Akoma Ntoso |
|---|---|---|
| Namespace (read it from the file) | http://schemas.gpo.gov/xml/uslm in the GPO samples | http://docs.oasis-open.org/legaldocml/ns/akn/3.0 |
| Root | Varies by document: pLaw, statuteCompilation, resolution | akomaNtoso, with a document-type element inside |
| Unit address | @identifier, a path such as /us/usc/t5/s1/a | eId inside, FRBR addresses outside |
| Stable handle | @id (GUID) | wId plus a mappings record |
| Past versions | @ date suffix on the path | Expressions, addressed by lang@date |
| Section text | num, heading, content | num, heading, then content blocks |
Two practical tips follow. First, do not match elements by local name alone and hope. The User Guide warns that some USLM element names match XHTML names without matching XHTML meaning. Second, never rely on document order for lookups. Build a dictionary once, keyed on the identifier. This sketch does that for a USLM file with the standard library:
import xml.etree.ElementTree as ET
def find_by_identifier(path, ident):
for _, el in ET.iterparse(path):
if el.get("identifier") == ident:
return el
return None
sec = find_by_identifier("COMPS-10510.xml", "/us/sComp/113/5/s1")
print(sec.find("{*}num").get("value"))
for sub in sec.findall("{*}subsection"):
print(sub.get("identifier"))
The {*} wildcard (Python 3.8 and later) matches any namespace, so the same code works whichever URL the file declares. Run against the sample above, it prints 1, then /us/sComp/113/5/s1/a and /us/sComp/113/5/s1/b. For a full Code title, which can be very large, keep iterating and write each match into an index instead of returning on the first hit.
What developers say
Developers who work with these files tend to be practical about them. In March 2014 a Hacker News commenter described building a toolchain for Akoma Ntoso documents, and explained the choice: "it seems that every player is hellbent on creating their own custom XML format, so I chose something with international appeal and government backing" (esbranson, Hacker News, March 2014). In March 2026 the same commenter, in a thread about legislation in version control, summed up the US picture: "Both are also provided in XML in bulk, though the former has a modern USLM format and the latter has an archaic schema." That sentence was about the US Code and the Code of Federal Regulations (esbranson, Hacker News, March 2026).
The maintainers' own issue tracker is useful too. In September 2023 a developer who wanted a searchable database of laws for "army corps of engineers civil works" asked how to pull USLM into database tables, and a reply from the maintainers' side was honest: "Because XML has a fundamentally different architecture from relational databases (hierarchical versus tabular), mapping between the two is never easy." The advice was to use tools built for XML, with XQuery (GitHub, September 2023). The developer replied with thanks for "the awesome work at gpo." Both halves of that exchange are fair.
Another thread explains why the files look the way they do. Asked in January 2024 how the plain-text view of a bill is produced, a GPO maintainer wrote that for enrolled bills, public laws and the Statutes at Large the USLM XML "is currently produced during GovInfo processing as a result of a conversion from locator files into USLM XML," and that locator files are "an internal, legacy, format that is used by GPO as part of the print production process." The same answer said GPO was replacing its print composition system with an XML-based one (GitHub, January 2024). If some of your files look more like conversions than native documents, that is why.
A quick story: the link that went to a casino
Here is a small one that shows how fragile the edges of a standard can be. In July 2026 a developer filed an issue noting that section 1.7 of the USLM User Guide, the one about the relationship to Akoma Ntoso, linked to the standard's old website, which "now serves an Indonesian online-gambling ('slot gacor') site." They suggested pointing to the OASIS pages instead. A maintainer replied the same day that the guide had been updated (GitHub, July 2026). The takeaway is cheap: when you cite a standard in your own documentation, link to the standards body, and recheck the links now and then.
Why stable identifiers matter
Law moves. Sections are repealed, renumbered and transferred. If you store "the third paragraph of the second subsection," a reorganization will silently point it at the wrong text. Both formats were designed around this problem, which is why each one keeps a handle that does not change next to a label that does.
A citation is the other half. Lawyers write 42 U.S.C. § 1983. A USLM address for the same provision is shaped like /us/usc/t42/s1983. Your system needs a mapping between the human string and your own key, and that mapping has to survive renumbering and amendment.
Here is how one API handles it. Everything below about this API comes from its published OpenAPI spec, which I read on 4 October 2026. Vaquill AI gives every section an actId, for example USC_T42_C21_S1983: Title 42, Chapter 21, Section 1983. Note that it includes the chapter, which a lawyer's citation and the USLM path both leave out. The spec says to take the actId from a search or resolve response and not build it by hand, because the title and section can be derived but the chapter cannot. Section 1 of Title 26 is the spec's own example of why: USC_T26_S1 misses.
The same provision under each scheme, with the USLM path shaped as the User Guide describes:
| Scheme | Identifier for 42 U.S.C. § 1983 |
|---|---|
| Lawyer's citation | 42 U.S.C. § 1983 |
| USLM path (shape from the User Guide) | /us/usc/t42/s1983 |
API actId (from the spec) | USC_T42_C21_S1983 |
To go from a citation to an actId, call the resolve endpoint:
curl -G "https://api.vaquill.ai/api/v1/us/statutes/resolve" \
-H "Authorization: Bearer $VAQUILL_API_KEY" \
--data-urlencode "cite=42 U.S.C. 1983"
The spec's example response, trimmed, looks like this:
{
"resolved": true,
"inputCitation": "42 U.S.C. 1983",
"section": {
"actId": "USC_T42_C21_S1983",
"citationShort": "42 U.S.C. § 1983",
"title": "Civil action for deprivation of rights",
"corpusType": "USC",
"currentThrough": "2026-09-02",
"sourceUrl": "https://uscode.house.gov/view.xhtml?req=granuleid:USC-prelim-title42-section1983&num=0&edition=prelim"
},
"creditsConsumed": 2.0
}
The sourceUrl points to the government's published text. When a citation does not resolve, the shape is the same with the verdict flipped, and the spec says to treat it as unverified. The 2 credits are charged either way.
{ "resolved": false, "inputCitation": "42 U.S.C. 99999", "section": null }
The -G and --data-urlencode flags let curl encode spaces and a section symbol for you. If you build the URL by hand, percent-encode the cite value. The spec's own examples write the citation without the section symbol, as above.
What the API spec returns
To read the text, pass the actId to the body endpoint:
curl "https://api.vaquill.ai/api/v1/us/statutes/section/USC_T42_C21_S1983/body" \
-H "Authorization: Bearer $VAQUILL_API_KEY"
The spec's example response, trimmed:
{
"actId": "USC_T42_C21_S1983",
"html": "<p>Every person who, under color of any statute...</p>",
"plain": "Every person who, under color of any statute...",
"available": true,
"creditsConsumed": 6.0
}
The body endpoint, GET /api/v1/us/statutes/section/{act_id}/body, returns HTML and plain text by default. A format parameter also accepts markdown, content and operative, and structured=true adds a tree of subsections for pinpoint citations. For US Code sections it also splits the body into content (the operative text), sourceCredit and notes.
To get the subsection tree, add the flag to the same call:
curl "https://api.vaquill.ai/api/v1/us/statutes/section/USC_T42_C21_S1983/body?structured=true" \
-H "Authorization: Bearer $VAQUILL_API_KEY"
The section metadata endpoint does list link fields, including xmlUrl ("XML version of the statute section", which the spec marks as nullable, so it appears where the source provides one) and granuleId and packageId for govinfo provenance, plus a publisherKey that for US Code sections is the OLRC's usckey, which the spec says joins directly to the uscode.house.gov bulk XML. So if you want the original USLM, you can use publisherKey to find the unit in OLRC's files. The call that returns those fields costs 2 credits:
curl "https://api.vaquill.ai/api/v1/us/statutes/section/USC_T42_C21_S1983" \
-H "Authorization: Bearer $VAQUILL_API_KEY"
If you only want the words and the citation, you can skip the XML entirely.
Which format should you design for?
If your inputs are American, build for USLM first. The Code, bills, public laws and compilations all use it, and the identifier rules in this post cover them. If your inputs include foreign statutes, add an Akoma Ntoso reader and normalize both into one record of your own, with an identifier column you control. If you do not want to run XML pipelines at all, a section-by-citation API gives you the record directly. The statutes API guide compares the main options, and govinfo API versus a legal data API covers the trade-off in more detail. For the wider market, see the best legal data APIs for developers.
If you would rather skip the XML and get a section by citation as JSON, the Vaquill AI US primary law API lets you resolve a citation and read the text with one API key.
This guide is part of US Law Data: The Complete Guide, a map of where US law comes from and how to use it.
FAQ
What is USLM? USLM stands for United States Legislative Markup. It is an XML schema for US legislative documents, maintained by the Government Publishing Office. The US Code, bills, public laws, the Statutes at Large and statute compilations are all published in it.
Is USLM the same as Akoma Ntoso? No, but they are related. GPO's govinfo page calls USLM a derivative of the international LegalDocML (Akoma Ntoso) standard. The USLM User Guide says it is designed to be consistent with Akoma Ntoso and shares many names, but it is not defined as a subset. A parser for one will not read the other without changes.
What is LegalDocML? LegalDocML is the OASIS name for the standard also called Akoma Ntoso. The OASIS LegalDocumentML Technical Committee maintains it, and version 1.0 became an OASIS Standard on 29 August 2018.
What does a USLM identifier look like?
It is a path. /us/usc/t5 is Title 5 of the US Code, and /us/usc/t5/s1/a is subsection (a) of section 1. Add @2013-05-02 to ask for the version in force on that date, and .htm or .xml to pick a format.
Where can I download the US Code as XML? The Office of the Law Revision Counsel publishes it at uscode.house.gov/download, as one file per title plus a zip of the whole Code. GPO also posts USLM XML for bills, public laws and the Statutes at Large on govinfo.
Where is Akoma Ntoso used? Outside the US it is used in legislative drafting systems. One documented example is the European Commission's LEOS tool, which stores content in Akoma Ntoso and supports the AKN4EU subschema. In US legal data work you will mostly see USLM instead.
How does Vaquill AI relate to USLM XML?
Its published spec serves section text as HTML and plain text, with optional Markdown and a subsection tree. Its metadata includes an xmlUrl link field and, for US Code sections, a publisherKey that joins to the OLRC bulk XML. The text endpoint is a JSON response, so you can use a section without parsing any XML.
What is the difference between @id and @identifier in USLM?
@id is an immutable unique handle that stays with an element for its whole life. @identifier is a URL-style address that describes where the element is now. Store @id if you need a handle that survives renumbering, and @identifier if you need an address you can resolve.
New legal AI guides, weekly.
Further Reading
Regulations.gov API: Dockets, Documents and Comments Explained
Read postCompliance API: Two Meanings, and Where Regulatory Data Comes From
Read postOpen States API Guide: v3 Endpoints, Keys, Bulk Data, and Licensing
Read postLegiScan API Guide: Bills, Votes, Datasets, and Joining Them to Enacted Law
Read postCongress.gov API Guide: What It Covers and What It Doesn't
Read postUS Statutes API: The 2026 Guide to USC, CFR, and State Code Access
Read post
Co-Founder & CTO
Priyansh leads engineering and AI at Vaquill AI: the pipelines that pull statutes, regulations and court rules from every US jurisdiction's official publisher, and the REST API, MCP server and open dataset that serve them.