Sources & Limits — Who Gave AI Its Idea of Goodness?
This note shows which sources support the essay, what they establish, and where the evidence stops. It is part of the essay’s free evidence layer.
Update (registered 8 September 2026; in the 9 September 2026 site release). Source 7 now reflects the 27 August 2026 rulings in the Anthropic–US government dispute, read directly from the court’s orders (Dkt. 250 and 251). This is a development in a live matter, not a correction of an earlier error: the essay’s original description was accurate as of its date. The current status carries an early-September reference date and remaining uncertainty about post-ruling appellate movement.
1. OpenAI publishes written rules for its models’ behaviour
Used for: OpenAI publishes a Model Spec setting out how its own rules interact with instructions from developers and users.
Best available evidence: The OpenAI Model Spec (official, versioned); announcement: “Introducing the Model Spec”; approach: “Inside our approach to the Model Spec”.
What it supports: The existence and content of a provider-authored behavioural specification.
Where the evidence stops: These are governance documents written by the company about its own intentions. They describe desired behaviour, not independently verified behaviour.
Source status: Official company documentation (provider-authored governance document).
2. Anthropic trains Claude on a written constitution
Used for: Anthropic trains Claude on a written constitution that it calls the final authority on the character it wants the model to have.
Best available evidence: Claude’s Constitution (official); explanation of the newer constitution: “Claude’s new constitution”.
What it supports: The existence, content and stated role of the constitution in training.
Where the evidence stops: As above: a provider-authored governance document, describing intent rather than independently audited outcomes.
Source status: Official company documentation (provider-authored governance document).
3. The US Department of Defense signed an agreement with Anthropic
Used for: A two-year agreement worth up to $200 million, with Claude integrated into workflows on classified networks.
Best available evidence: Anthropic’s announcement of 14 July 2025: “Anthropic and the Department of Defense to advance responsible AI in defense operations”. Instrument detail: a Chief Digital and Artificial Intelligence Office (CDAO) two-year prototype Other Transaction agreement with a ceiling of up to $200 million, one of several such awards to frontier-AI companies that month.
What it supports: The agreement’s existence, duration and ceiling. “Up to $200 million” is a ceiling, not guaranteed expenditure.
Where the evidence stops: The announcement describes the agreement’s scope, not how much was ultimately spent or what was delivered.
Source status: Official company statement; institutional procurement context.
4. Claude was deployed for national-security use in classified environments
Used for: Deployment in classified settings, and the fact that behaviour in that context could differ from civilian Claude.
Best available evidence: Anthropic’s announcement: “Claude Gov models for U.S. national security customers” (June 2025) — custom models “already deployed by agencies at the highest level of U.S. national security”, with access limited to classified environments, and adjusted handling for classified material.
What it supports: Actual deployment (not merely an announcement of intent), and that the military-context behavioural differences stem from Anthropic’s own customised model configuration rather than from a blanket government exemption. “Mission workflows” in the essay paraphrases the announcement’s application areas (strategic planning, operational support, intelligence analysis, threat assessment).
Where the evidence stops: The announcement is the provider’s own description; independent detail about what the deployed models do is not public.
Source status: Official company statement / product documentation.
5. Anthropic’s two stated red lines
Used for: Anthropic maintained publicly stated red lines against mass domestic surveillance and fully autonomous weapons.
Best available evidence: Anthropic’s public statements during the dispute, corroborated by the Congressional Research Service overview: Federal Government and Anthropic: Considerations for AI Innovation and Competition (CRS IF13217).
What it supports: The two positions as stated, precisely scoped to mass domestic surveillance (not all surveillance) and fully autonomous weapons (not all autonomy), held both publicly and as negotiation positions.
Where the evidence stops: No single canonical Anthropic “red lines” policy page was located; the positions are documented through official statements and institutional summaries rather than one dedicated policy document.
Source status: Official company statements; government/congressional research document for context.
6. The Pentagon’s position, and the supply-chain-risk designation
Used for: The Pentagon rejected the premise that a private supplier could restrict lawful military use, and designated Anthropic a supply-chain risk, restricting its role in relevant military contracts.
Best available evidence: The Secretary of Defense’s public direction (late February 2026) followed by formal notification on 5 March 2026, reported with direct official quotations. A Department official: “From the very beginning, this has been about one fundamental principle: the military being able to use technology for all lawful purposes” (Bloomberg, CNN); scope analysis for contractors: Mayer Brown.
What it supports: The designation’s existence, timing and scope: defense vendors must certify that they do not use Anthropic models as a direct part of Department contract work, which is narrower than a ban on all use, as Anthropic itself noted publicly. (This designation was later vacated by the 27 August 2026 summary judgment — see source 7.)
Where the evidence stops: The Department’s designation memo itself was not publicly available for inspection; the Pentagon’s position rests on officially quoted statements rather than a published primary document. This limitation is stated here deliberately.
Source status: Official statements quoted in contemporary reporting; legal/contractor analysis for scope, used because the primary document was not public.
7. Anthropic sued; a federal court struck down the measures, and appeals remain live
Used for: Anthropic sued over the government’s measures; a federal district court ruled them unlawful and struck down the supply-chain-risk designation, while a separate, narrower proceeding remained pending and the government continued publicly to treat the company as a supply-chain risk. The two proceedings have not resolved together.
Best available evidence: Two orders in Anthropic PBC v. U.S. Department of War, No. 3:26-cv-01996-RFL (N.D. Cal.), Judge Rita F. Lin, both filed 27 August 2026, read directly: the merits opinion (Dkt. 250) — a 59-page order on the cross-motions for summary judgment — for the reasoning and the disposition of the claims; and the separate Order of Final Relief (Dkt. 251) for the concrete legal consequences. The related appellate matters are a Ninth Circuit appeal (No. 26-2011) and a separate D.C. Circuit case (No. 26-1049). The late-August ruling and the government’s continued public position were also widely reported (NPR, Fortune, Reuters and others).
What it supports — reasoning (Dkt. 250): a grant of summary judgment for Anthropic on First Amendment retaliation, Fifth Amendment due process and the Administrative Procedure Act (including exceeding authority under 10 U.S.C. § 3252 and 5 U.S.C. § 558(b)); the court rejected the ultra vires separation-of-powers claim and entered judgment for the government as to agencies that took no relevant action or only interim measures. Its language includes “the challenged actions constituted unlawful retaliation in violation of the First Amendment” and “the empty invocation of national security is not a blank check to punish and retaliate against government critics”; page 59 states that “an order addressing relief will issue separately”.
Concrete legal consequences — Order of Final Relief (Dkt. 251, items 9–14): the participating defendants are permanently enjoined from giving effect to the challenged actions; the Supply Chain Designation is vacated, set aside and remanded under 5 U.S.C. § 706(2); the directive barring contractors from doing business with Anthropic is vacated and set aside; and the order expressly does not require the Department to use Anthropic’s products and does not prevent it from transitioning to other providers. No stay or bond was imposed, and non-participating defendants are excluded. The designation was thus narrower than a blanket ban (see source 6), and a ruling in one proceeding does not end the others. The caption names the “U.S. Department of War” (the renamed department); the essay’s “Department of Defense” refers to the 2025 agreement under that earlier name.
Where the evidence stops: The two orders are directly confirmed. On the position after 27 August 2026: a dated docket check on 9 September 2026 found no appeal of the judgment visible in the consulted snapshot and the D.C. Circuit case still pending. That rests on dated docket snapshots (a RECAP mirror), not a full live PACER search, and cannot capture filings made after the snapshot; the federal appeal window runs into late October 2026, so the absence of an appeal is a fact about the snapshot, not a closure. The agency’s own designation memorandum was not read as a standalone document (its content is known through the court’s account), and the government’s early-September statement that it still treats Anthropic as a supply-chain risk is a public position, not a court ruling. This is a live, time-sensitive matter, fixed only as of the reference date.
Source status: Two primary court orders directly inspected (N.D. Cal., 27 August 2026). Appellate status and the agency memorandum not independently confirmed; the government’s continued position via reporting. Live, time-sensitive matter.
Update policy
This page records the source basis used for the essay. Working papers, product documentation and live legal disputes may change; this note will be updated when a material source status changes. An update to this note does not silently rewrite the published essay — substantive corrections will be disclosed.
Last checked: 18 July 2026; sources 6–7 re-verified and expanded for the 27 August 2026 rulings, with both the merits opinion (Dkt. 250) and the Order of Final Relief (Dkt. 251) read directly, on 8 September 2026 (governance D-073 to D-077)