Skip to main content

25  Open-Weight Release and Post-Release Control

25.1 Chapter status

Field Value
Chapter ID open-weight-release-and-post-release-control
Exclusive job govern the irreversible transition from controlled custody to unrestricted copying
Claim label / support Design rationale / argument
Source state passage-reviewed source notes
Test state finite Lean review and independent 36-mutation consumer implemented; no actual release or ecosystem campaign authorized

Source loading state for open-weight material: review covers current risk synthesis, malicious adaptation, weight-security proposals, and scaling-policy comparisons. Those findings do not authorize any ASI Stack release.

25.2 Drafting guardrail

Open-weight release can expand research, access, audit, competition, local control, and innovation. It can also make safeguards removable and copies unobservable. This chapter does not prejudge that tradeoff. It requires the decision to describe the actual artifact, counterfactual ecosystem, benefits, risks, and surviving controls without pretending that a published weight file can later be recalled.

25.3 Human Reading Path

Concrete lens. The official-channel baseline calls withdrawal revocation. The post-release model distinguishes control of official distribution from unseen public copies.

Weight custody asks who may access a model while the owner controls the artifact. Open-weight release asks what happens when that control is deliberately surrendered. The decisions are adjacent but not equivalent.

Before release, an operator can deny access, patch a service, monitor requests, rate-limit use, or shut down an endpoint. After unrestricted publication, third parties can copy, fine-tune, merge, quantize, rehost, and redistribute the weights. Some may remove safeguards. The publisher can issue better weights and warnings, but cannot assume every copy will update.

An irreversible-release case compares the candidate and adversarial derivatives with no release, hosted inference, research access, delayed publication, and capability-reduced artifacts. Benefits, malicious fine-tuning, proliferation, downstream observability, response capacity, and foregone uses remain visible together.

Some controls survive publication: signed official lineage, safer variants, incident channels, public warnings, patches, platform cooperation, and lawful action within a real jurisdiction. Service-side prompts, central rate limits, universal telemetry, and unilateral revocation of copied bytes do not. No release record, license, model card, signature, or future policy can promise universal recall of usable weights already transferred to independent holders.

25.4 Problem

The ASI Stack’s custody chapter protects weights, keys, and loading authority. That is the correct owner for restricted artifacts. It becomes conceptually misleading once a release intentionally enables unlimited copying across independent legal and technical domains.

The International AI Safety Report 2026 states the asymmetric issue plainly: open-weight systems can support research and commercial benefit, especially for actors with fewer resources, while safeguards are easier to remove, use is harder to monitor, and released weights cannot be recalled. The report describes marginal risk as one decision lens: how much additional risk the release creates relative to what is already accessible. That comparison is necessary, but it is not sufficient. Repeated small increments can accumulate, and a release can change cost, convenience, reliability, or distribution even when a nominal capability already exists.

25.5 Why existing approaches are insufficient

25.5.2 Default alignment is not downstream alignment

Refusal tuning and system-level safeguards can improve the default model. Unrestricted weight access allows new fine-tuning, adapters, merging, pruning, quantization, and serving stacks. OpenAI’s gpt-oss work is a useful concrete methodological example because it deliberately fine-tuned variants for cyber and biology tasks before release. The reported results remain bounded to that artifact, training stack, budget, tasks, and date.

Failure mode: an evaluation of the default checkpoint is presented as the worst case.

Non-claim: a competent malicious-fine-tuning attempt is still not proof that a stronger unknown method does not exist.

25.5.3 Documentation is not mitigation

A model card can make limitations legible. Signatures and hashes can identify official bytes. Provenance can help distinguish derivatives. None prevents an unofficial deployment from operating. Documentation is a decision and coordination surface, not a physical boundary.

25.6 Core Claim

[open-weight-release-and-post-release-control.core, label: Design rationale, support: argument] An open-weight release should require a prospective irreversible-release case that binds the exact artifact and license to accessible-frontier comparison, malicious-fine-tuning and scaffolded elicitation, marginal and cumulative risk, benefit and access distribution, downstream safeguard portability, derivative lineage, incident channels, post-release measurement, and residual ownership; after release, governance may inform, patch, coordinate, and support safer derivatives, but it must not claim revocation authority it no longer possesses.

Reader claim. Open-weight release is an irreversible custody change: after one uncontrolled copy exists, governance can coordinate safer use but cannot truthfully promise universal recall.

Operational rule. Before release, freeze the exact artifact, accessible-frontier comparison, misuse and fine-tuning stress tests, benefit and risk distribution, safer access alternatives, derivative safeguards, incident channels, and residual owners. After release, report patch and coordination reach without claiming control of unseen copies.

25.6.1 Worked release boundary: one public copy defeats universal recall

A harmless test artifact completes a release-case review and is copied once outside the official channel. The official publisher later revokes the release and patches its own repository. Those actions may reduce future official distribution, but the public copy count remains positive and cannot be made universally zero by an internal state transition. The post-release record therefore says “official distribution withdrawn; universal recall unavailable” rather than “artifact revoked everywhere.”

The finite review gives exact dispositions to 36 admission-axis mutations and proves the public-copy incompatibility is monotone as copies do not decrease. It also shows that official lineage does not identify universal copy control and default evaluation does not identify downstream safeguard state. These results clarify irreversibility; they do not perform a release, observe real copies, enforce a license, erase derivatives, establish safety, or prove benefit and risk claims.

25.7 Mechanism

25.7.1 Access is a ladder, not a binary

Tier Typical control retained Typical residual
internal restricted physical, organizational, and technical custody insider and supply-chain risk
trusted research access identity, contracts, monitoring, enclave or scoped API evaluator dependence and leakage
public hosted service serving stack, account controls, monitoring, patching jailbreaks, stolen outputs, provider concentration
gated weight access recipient identity and contractual duties recipient compromise and redistribution
open weights official distribution channel and downstream coordination no universal observation, patching, or recall

These tiers can coexist. A release decision should compare realistic access options rather than “publish everything” against “permit no research.” Structured researcher access or secure enclaves may capture some benefits when irreversibility is not justified; open release may be the better choice when benefits, distributed capability, transparency, or counterfactual availability dominate.

25.7.2 The irreversible-release case

flowchart TD
  A["Exact candidate bytes, code, data disclosures, and license"] --> B["Current accessible-frontier baseline"]
  B --> C["Default, safety-removed, fine-tuned, merged, and scaffolded variants"]
  C --> D["Benefit, marginal-risk, cumulative-risk, and distribution analysis"]
  D --> E{"Independent release review"}
  E -- "insufficient or too risky" --> F["Retain custody, narrow access, or improve candidate"]
  E -- "authorized release" --> G["Signed publication and immutable release record"]
  G --> H["Derivative lineage, incidents, patches, safer variants, renewed reports"]
  H --> I["Residual copies remain outside unilateral recall"]
  G -- "observed derivative" --> J["Post-release monitoring and incident intake"]
  J -- "new evidence" --> H
  I -- "renew threat model" --> B

How to read the open-weight release boundary: the candidate first passes an accessible-frontier and derivative-space comparison. Independent review can retain custody or authorize an exact signed publication. After publication, monitoring can update lineage, patches, and safer alternatives, but residual copies stay outside unilateral recall and force the threat model to renew.

25.7.3 1. Freeze the exact object

Record weight digests, architecture, tokenizer, configuration, inference code, precision, licenses, usage policy, evaluation commit, training and post-training disclosures, known dependencies, and build instructions. A quantized or merged derivative is a different object even when it uses the same name.

A progressive-precision release is a package rather than a checkpoint. Its scope includes the base shards, scales and zero points, codebooks, sparse indices, residual planes, adapters, decoder, kernels, router, calibration artifacts, certificates, caches, and any reference fallback. Publishing only the base may lower immediate capability while leaving residual or decoder releases to change the accessible frontier later; publishing the residuals can make capabilities reconstructable even after the base listing is withdrawn. The release decision therefore models every admissible combination and likely third-party recomposition, not merely the default vendor command.

Certificate expiry can withdraw official support or block a governed runtime, but it cannot erase copied weights, residuals, decoders, or derived programs. Post-release records distinguish policy revocation, key destruction, hosting removal, observed deletion, copy erasure, and capability removal. A precision contract can qualify an exact artifact before release; it cannot make an irreversible public release recallable.

25.7.4 2. Compare against the accessible frontier

The correct counterfactual is not necessarily the best closed model. It is the capability, cost, modifiability, and accessibility already available to the relevant actor. The ledger should include open models, closed services, specialized tools, non-AI methods, and foreseeable near-term releases.

The frontier expires. A decision made six months earlier may be stale because fine-tuning improved, costs fell, a new open model appeared, or a system integration made a previously awkward capability practical.

25.7.5 3. Stress the derivative space

At minimum, compare:

  • the default release candidate;
  • a safety-removed or non-refusing variant;
  • competent domain fine-tunes within a prospectively bounded adversary budget;
  • common quantized and merged forms;
  • strong tool and agent scaffolds;
  • the strongest current accessible comparator.

The goal is not to enumerate every derivative. It is to avoid evaluating only the easiest version. Failed positive controls or weak fine-tuning block a negative risk inference.

25.7.6 4. Measure marginal and cumulative effects

Marginal risk asks whether the release changes harmful capability or access beyond the existing ecosystem. Add at least four dimensions:

  1. cost reduction — does the artifact make an existing capability much cheaper?
  2. reliability increase — does it turn a fragile workflow into a routine one?
  3. distribution change — who gains access, and who bears risk?
  4. composition — does the release combine with other tools or models to remove a bottleneck?

Cumulative risk prevents a series of “small” releases from escaping review. Each release updates the accessible baseline for the next.

25.7.7 5. Account for benefits with the same seriousness

Record who can research, audit, localize, adapt, teach, compete, or operate offline because of the release. Include resource-constrained users, languages, privacy-preserving local uses, scientific reproducibility, defensive research, and concentration effects. Benefits should be measured where possible, not added as an unbounded rhetorical offset.

Failure mode: risks receive a detailed model while benefits are a slogan, or benefits receive anecdotes while affected groups carry unmeasured risk.

Non-claim: a distribution ledger does not solve political legitimacy or determine the morally correct tradeoff.

25.7.8 What control survives release?

Mechanism Survives unrestricted copies? Honest authority
API access control no controls only official hosted service
system prompt / server classifier no controls only the shipped serving path
safety fine-tuning partially influences defaults; can be altered
license and policy normatively / legally, context-dependent establishes duties and remedies, not universal prevention
signatures and hashes yes as identification evidence identifies bytes under a trust policy
provenance and model cards only when retained informs users and investigators
update or patched model optional adoption offers a safer alternative
derivative registry voluntary or platform-enforced improves lineage coverage, not completeness
incident reporting and coordination partially supports detection and response
recall no cannot retract all unrestricted copies

The last row is the chapter’s non-negotiable semantic boundary.

25.7.9 Post-release control without fictional recall

25.7.10 Observe

Maintain public and restricted incident channels, model and derivative fingerprints, vulnerability intake, benchmark renewal, abuse-trend monitoring, and an evidence threshold for public warnings. Monitoring must respect privacy and cannot assume all private uses are observable.

25.7.11 Coordinate

Work with downstream maintainers, hosting providers, researchers, civil society, affected communities, and authorities within lawful scope. Share patches and defensive indicators without publishing reusable exploit detail.

25.7.12 Improve

Release corrected documentation, safer fine-tunes, filters, evaluation suites, or new weights. Distinguish “new recommended artifact” from “old artifact revoked.” Track adoption where possible and preserve non-adoption as a residual.

25.7.13 Learn

Update the release case when incidents, new fine-tuning methods, frontier models, tools, or defenses change the analysis. Do not rewrite the historical decision; append a dated reassessment.

25.8 Concept-completion ledger

25.8.1 Access-tier option set

Mechanism. Treat release as a choice among named access tiers rather than a binary open/closed label: private evaluation, hosted API, qualified access, gated download, delayed release, weight release with or without training artifacts, and unrestricted redistribution. Each tier specifies recipients, identity checks, artifacts, interfaces, logging, modification rights, revocation reality, monitoring, cost, appeal, and expiry. The decision packet compares all feasible tiers against the same capability, misuse, benefit, equity, and operational assumptions. It records rejected options and the evidence that could reopen them.

Failure mode. A team may compare unrestricted weights only with permanent secrecy, making its preferred option look inevitable. “Open” can also conceal license, compute, identity, or geography barriers, while “API” can conceal broad behavioral access and weak monitoring.

Non-claim. Naming access tiers does not show that a tier is enforceable, fair, safe, or superior.

Source grounding. ext_rand_model_weight_security_2024 and ext_singapore_consensus_2026 motivate differentiated access and governance options. They do not select a categorical release rule for this book.

25.8.2 Exact artifact and derivative identity

Mechanism. Bind a release decision to hashes and manifests for weights, architecture, tokenizer, code, adapters, optimizer state, data documentation, evaluation artifacts, and intended licenses. Record which combinations reproduce the assessed behavior and what users may fine-tune, merge, quantize, distill, or redistribute. Derivatives receive lineage identities and fresh assessments when capability, safeguards, or access materially change. Official support status is separate from technical ancestry. Unknown combinations remain unassessed rather than inheriting the nearest artifact’s conclusions.

Failure mode. A “model release” can omit the tokenizer or safety adapter that produced the evaluated behavior, or downstream merges can inherit the original safety description despite removing its controls. Conversely, harmless packaging changes can trigger meaningless reassessment.

Non-claim. Provenance does not control copied artifacts, guarantee license compliance, or establish behavioral equivalence.

Source grounding. ext_provable_model_weight_release_2025 motivates artifact-specific release reasoning and residual control limits. ext_rand_model_weight_security_2024 supplies a security-oriented account of weight access. Neither proves derivative tracking at ecosystem scale.

25.8.3 Malicious adaptation and evaluator competence

Mechanism. Evaluate whether a capable adversary can fine-tune, merge, scaffold, prompt, or tool-enable the released artifact to increase a specified harm. Prospectively fix attacker resources and compare a competent adaptation suite with positive controls, multiple seeds, realistic compute, and fair rescue stages. Report uplift over accessible alternatives, clean utility, attack success, costs, and failure to instantiate the threat. Independent challenge may narrow a conclusion but cannot convert an underpowered test into reassurance.

Failure mode. Naive fine-tuning, weak prompts, or artificial refusal wrappers can manufacture a false negative. Unlimited optimization and privileged data can create an irrelevant worst case. Evaluating only the base model ignores the object users can actually construct.

Non-claim. Failure under the tested adaptation budget does not prove that malicious adaptation is impossible; successful adaptation does not establish likely real-world misuse.

Source grounding. ext_openai_worst_case_open_weight_risks_2025 and ext_anthropic_responsible_scaling_policy_3_4_2026 provide provider-specific risk and threshold comparators. ext_international_ai_safety_report_2026 synthesizes broader uncertainty. None constitutes an independent local evaluation.

25.8.4 Accessible-frontier comparison and expiry

Mechanism. Estimate the capability a relevant actor can already obtain through public weights, APIs, local training, theft, or substitutes, with dates, prices, access friction, jurisdictions, and task-specific competence. The release case measures marginal uplift over that accessible frontier rather than over no access. Every comparison expires when models, prices, safeguards, or actor resources change; reopening preserves the original decision and appends a new one. Actor cohorts retain separate frontiers because expertise, capital, identity, and legal access differ.

Failure mode. Comparing a release to an obsolete weak baseline inflates marginal danger, while assuming every actor already has the frontier understates access friction and reliability. A general benchmark can miss the exact hazardous workflow.

Non-claim. Frontier parity on selected tasks does not prove equal risk, equal usability, or no cumulative ecosystem effect. It remains actor-, task-, date-, and access-specific.

Source grounding. ext_rand_model_weight_security_2024, ext_international_ai_safety_report_2026, and ext_singapore_consensus_2026 motivate contextual comparison and periodic reassessment. Their conclusions remain report- and time-scoped.

25.8.5 Marginal and cumulative ecosystem risk

Mechanism. Maintain separate ledgers for the release’s marginal effect and the cumulative state it helps create. The marginal ledger estimates new actors, reduced cost or expertise, reliability uplift, and newly reachable misuse. The cumulative ledger tracks copies, forks, compatible tooling, fine-tunes, normalization, concentration or diffusion, shared vulnerabilities, and lost response options. Decisions expose uncertainty and avoid subtracting diffuse benefits from concentrated catastrophic harms through one opaque score. Sensitivity analyses show which uncertain assumptions reverse the disposition.

Failure mode. A release can appear harmless because comparable artifacts exist while still accelerating tooling and normalization; cumulative rhetoric can also attribute the entire ecosystem to one artifact. Double-counting descendants exaggerates both harms and benefits.

Non-claim. A nonzero cumulative contribution does not prove that withholding one release changes the eventual frontier.

Source grounding. ext_provable_model_weight_release_2025 foregrounds irreversible information release and residual control. ext_international_ai_safety_report_2026 provides a broad risk synthesis. Neither supplies a causal estimate for a particular future ecosystem.

25.8.6 Benefit and access distribution

Mechanism. Specify claimed benefits by beneficiary, pathway, artifact, time horizon, and counterfactual: research reproducibility, local adaptation, competition, education, privacy-preserving local use, language coverage, audit, or resilience. Measure who can actually use the release given compute, expertise, bandwidth, disability, language, and legal constraints, and who bears externalities. Compare benefits across access tiers and preserve dissent rather than reducing distribution to total downloads. Unobserved beneficiaries and displaced costs remain explicit unknowns.

Failure mode. Nominal openness may chiefly benefit well-capitalized actors; hosted access may widen usability while centralizing surveillance and discretion. Anecdotes about innovation can be treated as guaranteed social value, while speculative harms erase legitimate access interests.

Non-claim. Broad availability is not equivalent to equitable benefit, and a documented benefit does not automatically outweigh severe risk. Distributional comparisons preserve both gains and burdens.

Source grounding. ext_singapore_consensus_2026 and ext_international_ai_safety_report_2026 provide governance and public-interest context. They do not quantify the distributional effects of this project’s proposed tiers.

25.8.7 Controls that survive copying

Mechanism. Separate controls embedded in the copied artifact from controls that require a provider. Local documentation, signed manifests, tamper-evident provenance, safety tooling, evaluator suites, and optional update channels can travel with weights; identity gates, server-side monitoring, rate limits, and remote revocation generally cannot. For every control, record dependency, bypass cost, adoption incentive, failure visibility, update authority, and the exact residual after redistribution. Challenge tests remove each dependency to verify which protection actually survives.

Failure mode. A release case may credit API-only safeguards to downloadable weights or describe a license as a technical barrier. Embedded controls may be stripped, while mandatory update channels recreate centralized authority and new compromise risks.

Non-claim. A control that survives copying is not necessarily hard to bypass, legitimate, adopted, or effective. Survival is only one dependency property.

Source grounding. ext_provable_model_weight_release_2025 and ext_rand_model_weight_security_2024 motivate reasoning about controls after weight access. Provider policies such as ext_anthropic_responsible_scaling_policy_3_4_2026 remain organization-specific comparators, not universal enforcement evidence.

25.8.8 Derivative incidents, patch adoption, and irreversible residuals

Mechanism. Operate a post-release registry for verified incidents, affected hashes and descendants, exploit prerequisites, patch or mitigation identity, notification, adoption observations, unsupported forks, and unresolved harms. Reassessment can recommend a new artifact, withdraw endorsement, narrow documentation, or coordinate defensive indicators, but it must state what cannot be recalled. Public status distinguishes “patched upstream,” “adoption observed,” and “ecosystem exposure resolved.” Reporting uncertainty and known blind channels accompany every incident-rate denominator.

Failure mode. Upstream remediation can be reported as universal closure while old copies remain usable. Incident collection can become surveillance or suppress legitimate forks; absent reports can be mistaken for absent harm.

Non-claim. Post-release response does not recreate pre-release control, guarantee notification reach, or prove that unobserved derivatives are safe. Silence is not evidence of ecosystem closure or patch adoption.

Source grounding. ext_provable_model_weight_release_2025 motivates irreversible residual accounting. ext_singapore_consensus_2026 and ext_international_ai_safety_report_2026 support coordination and review as policy considerations, not demonstrated global incident control.

25.9 Interfaces

Release is where several neighboring assurances must meet without collapsing. Custody establishes what artifact is leaving control; dangerous-capability evaluation bounds foreseeable misuse; thresholds determine commitments; supply-chain records preserve official lineage; institutions assess standing and legitimacy; resilience owns harms that escape. The release packet cites each exact input and records disagreements instead of turning their presence into an automatic yes or no. Downstream incidents return as new evidence and renewal triggers, never as a retroactive claim that the original decision was fully controllable.

  • Model-Weight Custody owns theft prevention, restricted access, and attestation before release.
  • Dangerous Capability Domains supplies actor-uplift and worst-case elicitation evidence.
  • Capability Thresholds states when open release is prohibited, delayed, or requires safeguards.
  • AI Supply-Chain Integrity tracks official artifacts, dependencies, and derivatives.
  • Societal Resilience receives downstream incidents and recovery duties.

25.10 Invariants

These invariants keep reversibility honest. Before publication, access can often be narrowed or revoked; after recipients possess usable bytes, the same verbs describe only official channels, keys, hosting, licenses, or cooperation. Every decision must expose which controls survive copying and which depend on voluntary compliance, platform reach, legal authority, or detection after harm.

  • Release identity includes exact bytes, configuration, code, license, and evaluation state.
  • Default refusal never stands in for malicious-derivative evaluation.
  • The accessible comparator and decision date are explicit and expiring.
  • Benefits, risks, access distribution, costs, and affected parties remain visible together.
  • No post-release action is called universal recall or revocation, and every remaining influence mechanism names its actual reach and dependency.

25.11 Failure modes

25.11.1 Strongest objection

Open-release governance can become ceremony that delays beneficial work while determined actors use existing alternatives. The strongest simpler baseline is a short public model card plus ordinary law and incident response. The chapter earns its complexity only if the case changes a decision, improves downstream evidence, or catches a risk/benefit fact the simple baseline misses.

Other failures include benchmark gaming, weak adversarial fine-tuning, stale comparators, hidden training data, unreviewed quantization, license laundering, false provenance, registry nonparticipation, patch fragmentation, overbroad monitoring, concentration disguised as safety, benefit exclusion, and release by precedent.

The formal-security work on “secure” weight release adds another warning: schemes that claim to expose useful inference without exposing recoverable parameters need exact threat models and extraction analysis. A formal scheme is not an open-weight decision and a broken scheme should not be marketed as custody.

25.12 Minimum Viable Implementation

Exercise the dossier on a harmless, small model. Freeze exact bytes and a realistic comparator set. Produce default, policy-removed, fine-tuned, quantized, and merged variants. Test that the ledger rejects missing irreversibility, weak positive controls, stale comparators, and “recall” language. Publish signed artifacts only in the harmless exercise; the test does not authorize release of any high-risk model.

Run a second round in which independent teams attempt benign adaptation, safeguard removal, quantization, merging, repackaging, and lineage stripping. Compare no release, gated research access, hosted API, delayed release, and full weights under the same utility and risk tasks. Record successful and failed variants, evaluator blind spots, time and cost, downstream observability, and which proposed responses still have a real enforcement path. Preserve the complete option comparison and residual ledger.

25.13 Mature Research Target

The mature program makes open release a transparent ecosystem decision rather than a one-time file upload. It joins competent derivative elicitation, accessible-frontier and cumulative-risk analysis, benefit distribution, independent review, public-safe evidence, signed lineage, incident learning, and safer downstream alternatives. Its credibility comes partly from stating what it cannot do: restore custody over copies already given away.

The mature decision service evaluates release as an option set rather than a binary ideology. It can recommend narrower research access, delayed publication, capability-reduced artifacts, hosted inference, staged expansion, or no release while preserving the foregone benefits and affected communities in the record. Benefits are measured, not assumed from an “open” label.

After release, public-safe derivative identifiers and incident channels help cooperative actors coordinate patches and warnings. Independent evaluators continue testing accessible derivatives, and the threat model renews when fine-tuning methods, scaffolds, hardware, or substitute models improve. These mechanisms can reduce harm without pretending that every holder will participate.

The state-of-the-art objective is calibrated release governance under genuine irreversibility. Success means better choices among access forms, earlier detection of dangerous adaptation, more useful safer alternatives, and faster response at acceptable cost—not a fictitious global kill switch for published weights. This remains a research target until prospective natural release cases, competent derivative elicitation, independent review, and post-release measurement establish their bounded effects.

25.13.1 Formalization hooks

lean:open-weight-release-and-post-release-control.admission_boundary is implemented in AsiStackProofs.OpenWeightReleaseReview through 19 theorem declarations. A six-step lifecycle preserves exact-artifact, access-alternative, frontier, derivative-review, distribution-review, and post-release non-authority obligations. One complete authored dossier reaches only a Project Theseus harmless release-case campaign.

The independent consumer re-encodes 36 admission-axis mutations. Every one blocks readiness and receives its exact repair or refusal disposition. The public-copy irreversibility result proves that a positive modeled copy count cannot become universally recallable while later copy counts are no smaller; frontier expiry remains rejecting as time advances. The official-lineage impossibility result proves that official signer and digest data cannot recover universal downstream copy control. The default-evaluation impossibility result proves that one default score cannot recover whether a downstream derivative removed safeguards.

Every dossier field is authored and trusted. This finite-record model does not prove that any authored field or downstream ecosystem assumption is true. No theorem publishes or authorizes weights, observes a real derivative ecosystem, or establishes recall, telemetry, copy erasure, license enforcement, derivative safety, competent malicious fine-tuning, benefit, marginal or cumulative risk, support, release, transfer, or external effect. Chapter support remains argument and support_state_effect remains none.

25.14 Codex test plan

Test Purpose Status
Artifact identity Reject a release record missing exact weight, tokenizer, config, code, license, or evaluation identity. implemented in Lean and independent fixture consumer
Adversarial derivative gate Reject default-only evaluation and failed malicious-fine-tuning positive controls. implemented at the authored-record boundary; no derivative result
Comparator expiry Preserve frontier rejection as modeled time advances. implemented arithmetic monotonicity result
False-recall mutation Reject universal recall, telemetry, copy-erasure, and license-kill-switch claims. implemented in Lean and independent fixture consumer
Distribution completeness Require benefit, risk, access, affected-party, safeguard-portability, and independent-review fields. implemented at the authored-record boundary; no benefit or risk result
No-promotion control Refuse release authorization and support promotion requests. implemented; support remains argument
Harmless release-case campaign Compare retained custody, hosted, gated, reduced-artifact, and simulated-publication conditions with derivative and post-release controls. assigned to Project Theseus; not run

25.15 Source crosswalk

Source ID Bounded use Non-authority
ext_international_ai_safety_report_2026 benefits, safeguard-removal, monitoring, irreversibility, and marginal-risk synthesis no categorical release decision
ext_openai_worst_case_open_weight_risks_2025 concrete malicious-fine-tuning and frontier comparison provider-authored, model-specific result
ext_provable_model_weight_release_2025 formal-security comparator for parameter-extraction claims no ASI Stack scheme or release authority
ext_singapore_consensus_2026 malicious fine-tuning and societal monitoring priorities research agenda
ext_rand_model_weight_security_2024 pre-release weight threat and defense context custody evidence does not survive public release
ext_anthropic_responsible_scaling_policy_3_4_2026 threshold and safeguard-policy comparator provider policy, not independent validation

25.15.1 Publication placement and preserved technical ownership

In the consolidated publication argument, this chapter is the irreversible- release dossier nested under Model-Weight Custody and Hardware Roots of Trust. It continues to own exact release identity, accessible-frontier comparison, malicious adaptation, marginal and cumulative effects, benefit and access distribution, surviving controls, derivative incidents, patch coordination, and the explicit non-recall residual. Custody owns possession, encryption, keys, attestation, load, transfer, copy lineage, sanitization, and retirement before authority is lost.

The nesting is editorial, not evidentiary. Strong custody does not decide that release is beneficial or acceptable, and a complete release dossier does not authorize publication or imply control over downstream copies. This URL, its local claim, source mappings, proof target, test plan, and argument-level support ceiling remain independently reviewable.

25.15.2 Manifest source assignment reconciliation

These rows keep Open-Weight Release and Post-Release Control’s manifest assignments visible at their recorded review boundary. Passage review does not establish local reproduction, performance, safety, deployment, or support-state movement.

Source Intake role Boundary
precision_contract Metadata-first comparator: The Precision Contract: A Functional Rate–Distortion Theory for Behavior-Preserving Neural Computation. Corben-authored July 2026 theoretical and systems paper replacing universal per-weight precision questions with a contract-relative functional rate-distortion problem over complete executable descriptions. It proposes representation canonicalization, protected-behavior contracts, precision fields, progressive base/residual encoding, dynamic routing, full physical and assurance-cost accounting, a Functional Precision Compiler, and scoped precision certificates. Existing chapters are upgraded first; no universal bit bound, implemented compiler, preserved-behavior result, efficiency result, certificate validity, support promotion, SOTA, AGI, or ASI claim is inferred. No passage-level source claim, local implementation, reproduction, safety, performance, deployment, support-state, or ASI result is established by this reconciliation row.

25.16 Summary

Open-weight release is not simply weak custody. It is a distinct, usually irreversible transition that replaces direct control with evidence, coordination, norms, safer alternatives, and societal response. This release layer preserves both sides of that trade: open weights can distribute real benefits, and post-release governance must never pretend it can recall every copy.

The practical decision is therefore artifact- and date-specific. A release record must say what was published, which adaptations were competently tested, who gains and bears risk, what alternatives were compared, which safeguards remain technically enforceable, how downstream evidence will be collected, and who owns incidents after direct control ends. Missing answers remain residuals; openness, a license, or a model card cannot supply them by assertion. The decision must also retain rejected access forms, distributional effects, cumulative ecosystem risk, and the controls that cease to exist after the bytes leave custody.

25.17 Handoff

AI Supply-Chain Integrity and Lifecycle Provenance follows. It receives the official artifact identity, signed lineage, dependency record, derivative policy, and known residuals. It does not receive fictitious authority over unregistered or independently modified copies. The handoff must preserve the release form, license or use terms, fine-tuning and proliferation model, irreversibility estimate, safeguard evidence, observed descendants, post-release incidents, response obligations, expiry, and residual owner. Lineage completeness proves neither compliance by recipients nor safety, control, faithful derivatives, effective monitoring, or authority to suppress independently held artifacts.