Abstract
Between 16 and 24 July 2026 the United States government came closer to restricting foreign open-weight AI models than at any point in the technology’s existence. The Treasury Secretary put sanctions and Entity List designations on the table. The Director of the Office of Science and Technology Policy named a specific Chinese laboratory and a specific American model. Axios reported an administration actively weighing a ban. Not one of those statements rests on an agreed definition of the thing being restricted.
“Open source” has no enforced meaning in artificial intelligence. The Open Source Initiative published a formal definition in October 2024; the largest company claiming the label rejected it, and no regulator or procurement authority anywhere has adopted it as binding. In the vacuum, the word has become a marketing asset, a regulatory exemption, and now the object of a federal policy fight — while the licence files it supposedly describes say something else entirely.
The paper reads the licences. It finds that of eleven major model families surveyed, three satisfy the OSI definition, and none of those three operates at frontier scale. It finds that Meta’s Llama 4 licence withholds rights from anyone domiciled in the European Union, inside a product marketed as open source, in the one jurisdiction that offers open-source models a regulatory exemption. And it finds that the characterisation now driving policy — that training on another party’s output is fair use when American labs do it and theft when Chinese labs do it — is real at the level of framing, weaker than advertised at the level of evidence, and being applied to conduct that three separate bodies of law treat three separate ways.
Key Findings
01 · Definition
The OSI’s Open Source AI Definition v1.0 has existed since October 2024 and has been adopted as a binding qualifying criterion by no regulator or public procurement authority located in the survey.
02 · Licence text
Llama 4’s licence states that rights are not granted to individuals domiciled in, or companies with a principal place of business in, the European Union.
03 · Scale
Of eleven major model families surveyed, three meet the OSI definition — OLMo 2, Pythia/GPT-NeoX, and, arguably, BLOOM. All three sit well below frontier scale.
04 · Litigation
The largest copyright settlement in US history — $1.5 billion, final approval 20 July 2026 — turned on how books were acquired, not on whether training is fair use. On the training question the same court ruled the other way.
05 · Exposure
If Chinese open weights were declared legally tainted, Cursor’s Composer 2 and Thinking Machines’ Inkling are both downstream by their own published account.
Method
Every load-bearing claim in the paper carries one of four evidence grades — verified (primary source retrieved and quoted directly), corroborated (two or more independent secondary sources agree), attributed (on record from one interested party, uncorroborated), and undetermined (could not be established; stated as a gap, never filled). The single claim that would be most damaging to the paper’s central subject is graded attributed and excluded from the findings. Every revenue and market-share figure circulating in podcast coverage of the dispute was excluded as third-party, unattributed and internally inconsistent.
The paper carries an author’s disclosure: the analysis was researched and drafted with substantial assistance from Claude, a model made by Anthropic — a party whose conduct §6 of the paper examines. The disclosure, the adversarial testing applied to the sections most favourable to that party, and the full source list are in the PDF.
Engage with this Working Paper
Substantive comment, technical critique, and named response are welcomed. The Working Paper is maintained as a public version-controlled document; the PDF above is the editorially-controlled v1.0 release. Every claim carries a grade; corrections that move a grade are welcomed and will be published.