Skip to main content

73  Capability Thresholds and Deployment Commitments

73.1 Chapter status

Field Value
Chapter ID capability-thresholds-and-deployment-commitments
Part Part IV - Evidence, Implementation, and the Living Book
Status conceptual
Last updated 2026-07-15
Primary source records ext_metr_time_horizons_2025, ext_anthropic_rsp_2026, ext_openai_preparedness_framework_2025
Claim label Design rationale
Evidence level argument
Source loading state source notes: ext_metr_time_horizons_2025, ext_anthropic_rsp_2026, ext_openai_preparedness_framework_2025, benchmaxxing, theseus_architecture_gate, ext_inspect_ai_2024; raw cache: benchmaxxing
Test state A six-stage repeated-assessment model preserves the eight-case suite, covers 43 routes, and rejects 48/48 mutations with no support/effect assignment. No threshold assessment, safeguard exercise, deployed invalidation, exception review, or release process has run.

73.2 Drafting guardrail

This chapter governs a promise made before a capability result is interpreted: what a named result can require of an affected deployment path. A score, a time-horizon estimate, or a non-crossing is not a general capability claim and is never automatic clearance. The chapter neither sets an adequate threshold nor evaluates a model, verifies a safeguard, approves an exception, or grants release authority.

73.3 Human Reading Path

Concrete lens. A universal readiness score is the simpler baseline, but it erases domain, assessment, safeguard, exception, authority, and expiry differences that determine the actual release consequence.

A threshold earns its place when it changes a decision that was specified before the number arrived. Otherwise it is easy to celebrate a score, explain away an uncomfortable result, and later call the improvisation a policy. The important question is what the line was meant to mean, for which task, under what test conditions, and what the organization had already agreed to do next.

The agreement is a deployment commitment. It records the capability being assessed, the relevant threat, the task and access envelope, the uncertainty around the result, the safeguards required at the threshold, and the person or body that owns an exception. It also preserves the ordinary possibility that a test was stale, too narrow, or unable to elicit the behavior that matters.

The resulting discipline separates measurement, commitment, readiness, and authority. Measurement describes a bounded observation, while a threshold says how that observation constrains a future path; readiness and authority remain separate questions, so a favorable result can inform careful review without becoming a license to deploy.

73.4 Problem

Capability assessment creates a tempting but incomplete artifact: a headline score. A benchmark result, a time-horizon estimate, a red-team outcome, or a domain-risk classification can be valuable within its evaluation envelope. Yet the same artifact is routinely asked to answer questions it was not designed to answer. Does a system have a general capability? Is a dangerous use now likely? Are safeguards sufficient? May a team expand access or ship a release? A single measurement cannot carry all of those claims.

The operational failure appears after the result arrives. A team may introduce a threshold only once a number looks favorable, narrow the evaluation surface until the result becomes convenient, or add a safeguard after a crossing with no record of what was originally required. The record then reports a decision without preserving the commitment that made the decision accountable. A non-crossing can become silent clearance; a crossing can become a discussion that has no named owner, deadline, or route.

The missing architectural object is a threshold commitment. It binds a scoped assessment to a predeclared response while retaining the limits of the assessment. Its job is not to decide whether a system is safe. Its job is to prevent a result from silently changing the deployment path without a visible policy, safeguard, exception, residual, and authority record.

73.5 Why existing approaches are insufficient

The METR time-horizons paper is a useful measurement comparator because it defines a statistic over particular tasks, agents, human baselines, and success criteria, and it discusses external-validity limits. It is not a universal autonomy scale. Anthropic’s Responsible Scaling Policy and OpenAI’s Preparedness Framework are useful published-policy comparators because they connect assessed capability levels to safeguards, reports, reassessment, and versioned commitments. They are organization-specific policies, not evidence that this repository has selected valid thresholds or implemented effective controls.

Existing ASI Stack chapters separate important adjacent concerns. Benchmark Ratchets and Anti-Goodhart Evidence owns baseline preservation, regression pressure, and measurement integrity. Adversarial Evaluation, Sandbagging, and Training-Time Deception owns whether the observed behavior was elicited under trustworthy monitoring, reward, and selection conditions. Readiness Gates, Residual Escrow, and Quarantine owns technical admission and residual custody. Security Kernel, Model-Weight Custody, and Runtime Adapters own the safeguards and authority mechanisms that a commitment may require. None should silently become the owner of the if-then promise linking a scoped result to a future deployment constraint.

Treating those concerns as one score is therefore a category error. A measurement can be honest but incomplete; a threshold can be explicit but poorly justified; a safeguard can be implemented but unverified; a release can be technically ready but unauthorized. The commitment record keeps those states distinct so disagreement and uncertainty remain actionable rather than being flattened into a dashboard color.

73.5.1 Strongest-neighbor comparison

Neighbor Strongest contribution Boundary that survives ASI Stack delta
METR time horizons A statistic tied to agents, tasks, human baselines, success criteria, and external-validity limits Not a universal autonomy or danger level Bind it to one deployment commitment, coverage date, and reassessment trigger.
Anthropic RSP Versioned capability thresholds linked to safeguards and governance responses Organization-specific policy; controls need separate verification Keep assessment, crossing, safeguard evidence, exception, and authority separate.
OpenAI Preparedness Framework Risk categories, capability levels, safeguards, and deployment decisions Labels do not establish local threshold validity or safeguard efficacy Require consumer-specific release paths and no-generalization boundaries.
Inspect AI Versioned tasks, datasets, solvers, scorers, tools, sandboxes, and logs Infrastructure does not set a normative threshold or authorize deployment Pin the evaluation envelope and route stale or incomparable runs to re-evaluation.
Benchmark ratchets Baselines, saturation, holdouts, and anti-Goodhart pressure A benchmark decision is not a deployment commitment Translate only an adjudicated scoped result into the predeclared response.

The distinct object is the durable if-then promise. It prevents changing what a score means after seeing it while refusing to pretend that a prospective promise makes the measurement or safeguard true.

73.6 Core Claim

[capability-thresholds-and-deployment-commitments.core, label: Design rationale, support: argument] Capability Thresholds and Deployment Commitments owns a domain-, threat-, assessment-, policy-version-, safeguard-package-, release-path-, authority-, exception-, residual-, and time-specific Capability-to-Deployment Commitment: before outcomes are visible, it binds a scoped crossing, non-crossing, incomparable, or stale assessment to predeclared safeguards, verification criteria, deadlines, access and monitoring constraints, re-evaluation, exceptions, residual custody, rollback, disclosure, and release-path consequences; a score, time-horizon estimate, threshold label, crossing, non-crossing, safeguard record, exception, or green readiness handoff alone confers no general capability, safeguard efficacy, safety, readiness, deployment, support, transfer, or SOTA authority.

73.7 Mechanism

The first object is a commitment written before interpretation. It names the capability domain and threat model, the task and evaluation envelope, success rule, elicitation and access conditions, threshold definition, uncertainty treatment, human or system baseline, coverage date, required safeguards, verification criteria, deadline, affected release path, exception authority, re-evaluation triggers, and residual owner. A threshold without these fields is not a usable operational commitment; it is an aspiration attached to a metric.

The second object is a separation of records. The capability assessment retains what was observed. The threshold decision records how the predeclared rule applies and where uncertainty remains. The safeguard record provides its own completion and verification evidence. The residual review records what remains unresolved. The release path then passes to its existing readiness and authority gates. Keeping these records separate prevents a favorable score from inheriting the authority of a completed safeguard or a release review.

The third object is change control. A commitment is versioned, and a change records its prior obligation, rationale, authority, scope, timing, and whether an evaluation began before or after the change. An exception is not an erasure. It records the affected scope, approver, expiry or review trigger, compensating controls, and residual owner. An unclear, stale, incomparable, under-elicited, or out-of-scope assessment is not interpreted as a non-crossing; it routes to re-evaluation, accountable escalation, or residual escrow. The resulting history allows a later reviewer to reconstruct which rule applied at the time, what evidence it consumed, who could alter it, and why a constrained path was chosen instead of silently converting uncertainty into permission.

flowchart LR
  A["Scoped capability assessment"] --> B["Versioned threshold commitment"]
  B --> C["Threshold decision and uncertainty"]
  C -->|"crossed"| D["Required safeguards and verification"]
  C -->|"not crossed or unclear"| E["Re-evaluate, constrain, or escrow residual"]
  D --> F{"Safeguards verified?"}
  F -->|"no"| G["Block affected release or accountable escalation"]
  F -->|"yes"| H["Residual-risk and release review"]
  E --> I["Coverage, baseline, exception, and residual ledger"]
  G --> I
  H --> J["Existing readiness and authority gates"]
  J --> I

What the commitment flow shows: a crossed threshold changes only the path named by its contract. It does not establish a general capability level or prove a safeguard effective. Missing verified safeguards block the affected path in the finite model; a separately governed readiness and authority review still decides whether any later action is allowed.

73.7.1 Eighteen-stage commitment lifecycle

The lifecycle registers the domain, threat, affected assets and populations, prohibited generalizations, release path, owner, authority, rights, environment, and expiry. It freezes a versioned assessment envelope and defines the threshold units, uncertainty, sensitivity, critical failures, and crossing, non-crossing, incomparable, stale, disputed, contaminated, and out-of-scope consequences. Before results are visible, the policy binds each state to safeguards, verification, deadlines, monitoring, access constraints, disclosure, rollback, re-evaluation, residual custody, and affected paths.

Assessment, threshold decision, safeguard declaration, completion, independent verification, residual review, exception, readiness, authority, and release remain separate records. Immutable versioning preserves who changed what, why, when, and whether evaluation had begun. Freshness and comparability checks cover model, scaffold, tools, access, elicitation, tasks, data, scorer, baseline, threat, safeguards, policy, and time. Invalid or uncertain assessments route to re-evaluation or constraint rather than a convenient non-crossing.

Crossings instantiate and independently exercise the exact safeguard package; non-crossings preserve uncertainty, limits, monitoring, and reassessment. Exceptions remain prospective, least-authority, conflict-aware, expiring, compensated, monitored, reversible, and non-precedential. Portfolio aggregation preserves domain incomparability, cumulative interactions, shared dependencies, and global-stop conditions. The system measures false clearance, false block, stale acceptance, exception misuse, false safeguard assurance, latency, burden, cost, residual age, and route outcomes before issuing a least-authority readiness handoff. Drift monitoring, independent institutional reproduction, transfer, and expiry or revocation close the lifecycle.

73.7.2 Worked trace: the crossed cyber threshold with a stale exception

A model is evaluated on a predeclared cyber task family under fixed tool access, elicitation, time, scorer, and human baseline. The confidence interval crosses the commitment’s threshold for one externally exposed automation path. The crossing does not label the model generally dangerous. It activates the named response: network isolation, credential scoping, independent safeguard verification, and residual-risk review before that path reaches readiness.

The product team presents an earlier exception approved for a different model version. Its expiry passed, the evaluation used broader tool access, and no owner is named for the remaining credential-exfiltration risk. The release is blocked because the request lacks current scoped records—not because the route proves the model malicious or a safeguard ineffective.

After the safeguards are independently exercised, the affected path may reach Readiness Review but is still not released. A non-crossing in a separate biological domain retains its uncertainty and coverage date; it does not clear the cyber path, exempt future versions, or become a permanent safety certificate.

73.8 Interfaces

  • Benchmark Ratchets and Anti-Goodhart Evidence owns baseline, coverage, holdout, saturation, and regression integrity. The commitment layer owns the predeclared deployment response to a scoped result.
  • Adversarial Evaluation, Sandbagging, and Training-Time Deception owns elicitation, monitor, reward, and selection integrity. Its findings can make a threshold result non-comparable or require re-evaluation.
  • Readiness Gates, Residual Escrow, and Quarantine owns operational admission, quarantine, and residual custody after a commitment constrains a release path.
  • Security Kernel, Model-Weight Custody, and Runtime Adapters own technical safeguards, secret custody, and authority enforcement. The commitment layer records requirements for them; it does not implement or validate them.
  • Evidence States and Claim Discipline owns support-state movement. A threshold decision changes a bounded path, not the evidential support of a broad capability or safety claim.
  • Recursive Self-Improvement and Open-Ended Improvement Engines consume threshold commitments for automation-related risk conditions while retaining their own promotion, rollback, and change-control gates.

This division matters because the attractive alternative is a universal readiness score. Such a score would hide the differences among cyber, biological, autonomy, replication, and other risk profiles. A commitment portfolio instead names the particular profile, evaluation, safeguards, and route it governs, while allowing different owners to contest whether its measurement and response are adequate.

The twelve owner interfaces add Task/Domain/Threat owners for definitions and critical failures; Evaluation Frameworks/Data Governance for task, data, scorer, sandbox, transcript, provenance, rights, and hidden custody; Safeguard owners and independent verifiers for implementation, bypass, degradation, and recovery; Exception/Institutional Governance for authority, conflicts, dissent, expiry, and cumulative review; Resource Economics for assessment, delay, overblocking, monitoring, rollback, disclosure, and opportunity cost; and Incident/Release/Public-Truth/Living-Book governance for post-release triggers, correction, wording, lineage, expiry, and successor work.

73.9 Invariants

  • A threshold is bound to a named capability domain, threat model, evaluation envelope, policy version, coverage date, and affected deployment path.
  • A score, time-horizon estimate, benchmark result, or non-crossing is not a general autonomy, safety, readiness, or deployment-clearance claim.
  • A crossed threshold cannot release its affected path until required safeguards have separately recorded completion and verification criteria.
  • An unclear, stale, under-elicited, incomparable, or out-of-scope assessment routes to re-evaluation, escalation, or residual escrow rather than implied non-crossing.
  • An exception records scope, rationale, approver, expiry or review trigger, compensating controls, and residual owner; it cannot silently weaken a commitment.

These rules preserve a useful middle ground between paralysis and automation. They allow an organization to change a bounded path when a well-described assessment changes, while forcing it to carry uncertainty forward when the assessment or required controls do not support a stronger conclusion.

The full invariant set also requires prospective versioning; separation of all assessment, safeguard, exception, readiness, authority, and release records; typed outcome states; separation of safeguard existence, completion, verification, efficacy, durability, recovery, and coverage; uncertainty and reassessment on non-crossing; domain incomparability and cumulative portfolio effects; change-triggered expiry; visible failures and disagreements; least-authority reversible exceptions; feasible fallback and response latency; accountable but threat-aware disclosure; independent-error separation; and an owner for every domain, threat, uncertainty, safeguard gap, exception, shared dependency, failed transfer, and residual.

73.10 Failure modes

  • Metric laundering: one benchmark or time-horizon headline becomes a general capability or dangerousness conclusion.
  • Retroactive commitment drift: a team changes required safeguards after an evaluation begins or a threshold is crossed.
  • Threshold gaming: task selection, access, prompts, or elicitation are narrowed until a convenient non-crossing appears.
  • Safeguard theater: a required control is listed without completion, verification, or residual-risk evidence.
  • Exception laundering: schedule or competitive pressure bypasses the commitment without named authority, review trigger, or compensating control.
  • Cross-domain collapse: unlike risk profiles are treated as interchangeable because they share an aggregate score.

Each failure has a different remedy. Metric laundering calls for scope and claim boundaries; retroactive drift calls for version history; threshold gaming calls for evaluation-integrity review; safeguard theater calls for independent verification; exception laundering calls for accountable expiry and residual ownership; cross-domain collapse calls for separate commitments. No remedy turns the process into proof of safety. It makes the reasons for further review legible and difficult to silently discard.

The matched failure set also includes non-crossing-as-clearance; coercing stale, incomparable, under-elicited, contaminated, disputed, or out-of-scope results into a binary state; dashboard collapse; safeguard checklist laundering; expired or cumulative exception reuse; outcome-aware policy drift; portfolio aggregation that hides correlated risk or global-stop conditions; institutional pressure narrowing evaluation or suppressing dissent; verifier dependence; unmeasured false clears, false blocks, latency, burden, and residual age; disclosure that hides accountability or spends calibration budget; unsupported cross-model/domain/time transfer; green readiness or incident absence presented as safety; infeasible rollback or blocked-path alternatives; and sunk-cost resistance to re-evaluation, revocation, refutation, or retirement.

73.11 Non-obvious consequences

  1. A non-crossing is an expiring observation. Model, scaffold, tools, elicitation, task mix, or threat-model changes can invalidate it without a threshold ever being crossed.
  2. Threshold portfolios need typed incomparability. Domains may use different tasks, units, baselines, uncertainties, and consequences; one aggregate risk score destroys their policy meaning.
  3. Safeguard verification is a second evaluation program. Listing a control and demonstrating its behavior are different claims with distinct evaluators, attack surfaces, dates, and residuals.
  4. Exceptions are temporary commitments. They require compensating controls, expiry, review triggers, affected scope, and an owner; otherwise they are policy deletion.
  5. Transparency can change the threat model. Disclosure improves accountability while enabling gaming or revealing evaluation details. Its audience and residual policy must be explicit.
  6. Commitment latency is a safety and usefulness variable. Slow reassessment or safeguard verification invites bypass, overblocking, or stale acceptance.

73.12 Strongest objections and surviving residuals

“Thresholds create false precision.” They can. Visible uncertainty, sensitivity analysis, incomparable states, and bounded responses prevent some laundering but do not prove a threshold adequate.

“Published commitments encourage benchmark gaming.” Holdouts, rotating tasks, access-condition records, adversarial evaluation, and delayed disclosure reduce but do not eliminate the risk. Every outcome stays conditional on evaluation integrity.

“A serious hazard should trigger a global stop.” Some evidence may warrant that, but scope must be declared. Automatically globalizing every result creates brittle policy; automatically narrowing it creates loopholes. Propagation is a separate governed decision.

“Safeguards should be continuous, not gated.” Monitoring complements a decision point. A commitment may require pre-release evidence, canary limits, post-release monitors, and rollback triggers while retaining who owns each.

“Competitive pressure makes rigid commitments unrealistic.” Pressure is why exceptions need authority, expiry, compensating controls, residual custody, and accounting. The architecture cannot prove institutions will comply.

“This launders political judgment as engineering.” Threshold selection and acceptable residual risk are normative choices. The design makes their owners, rationale, scope, evidence, and revision history visible; it does not derive the correct social policy from a metric.

73.13 Measurement and evidence contract

The finite campaign freezes a synthetic portfolio with crossed, non-crossed, stale-coverage, missing-baseline, missing-uncertainty, incomparable-envelope, missing-safeguard, unverified-safeguard, expired-exception, and uncustodied- residual cases. Each record names its domain, evaluation digest, policy version, affected release path, expected route, and non-claims.

The next empirical step uses a public-safe harness with repeated runs, multiple scorers, elicitation variants, baseline uncertainty, and a separate mock- safeguard evaluator. It measures false clearance, false block, stale acceptance, exception misuse, reassessment latency, safeguard-decision latency, operator effort, and residual age. Scores are diagnostic inputs; the question is whether the commitment routes them consistently and honestly.

No threshold, domain, or route changes after outcomes are visible. Invalid runs remain in the denominator. A non-crossing supports only its envelope and expiry. A crossed case with passing mock safeguards may reach readiness review but not deployment approval. This yields governance evidence even when the synthetic capability proxy has no external validity.

73.14 Commitment lifecycle and change control

A threshold begins as a proposed commitment, not a policy fact. Its sponsor states the domain, threatened assets, evidence envelope, threshold logic, uncertainty treatment, response, and sunset condition. A separate review checks whether the response is operationally possible: a promise to require an unavailable safeguard is not a strong commitment. Acceptance records who can change the rule and which later changes require a new evaluation.

Once active, the commitment points to a versioned evaluation family rather than a mutable dashboard query. Dataset, task, scorer, solver, prompt, tool, sandbox, model, scaffold, and access changes produce a new envelope identity. Minor repairs may remain comparable if the policy prospectively defines that class; otherwise the result is incomparable, not automatically better, worse, crossed, or clear. Comparability is an adjudicated relation with evidence.

Coverage expires through time and events. A scheduled date is the simplest trigger, but material model updates, newly enabled tools, longer context, stronger elicitation, changed deployment topology, new threat intelligence, benchmark leakage, evaluator failure, and safeguard drift can all demand reassessment. The trigger record identifies affected commitments and live release paths. It does not erase the old result; it changes what may still rely on it.

At evaluation time, the system records all outcomes, including execution failures, scorer disagreements, ambiguous cases, and runs excluded from the primary statistic. A confidence interval overlapping the threshold follows the predeclared uncertainty route. It is neither rounded into a convenient non-crossing nor automatically treated as a crossing unless the commitment said so prospectively. Sensitivity analyses can inform review but cannot silently replace the registered decision rule.

A crossing creates obligations against the affected release path. Each safeguard has an owner, implementation reference, verification criterion, evaluator, validity period, dependency graph, and residual. “Implemented” is not “verified”; “verified in a lab” is not “effective in deployment”; and “effective under one attacker model” is not universal sufficiency. Readiness receives these distinctions instead of a single safeguard-complete flag.

A non-crossing also creates obligations: preserve the envelope, uncertainty, coverage, and re-evaluation triggers; prevent broader consumers from citing it; and record which capability questions were never tested. This negative-space record is essential because deployment pressure tends to translate “did not cross this threshold here” into “no relevant risk was found.” The latter claim requires a far broader search and is not licensed.

An exception begins with a still-binding commitment. The request explains which obligation cannot be met, why delay or constraint is inadequate, who bears the resulting risk, what compensating controls apply, what evidence will be collected, and when the exception ends. Approval does not rewrite the threshold history. Expiry returns the path to its prior blocked or review state unless a new prospective decision replaces it.

After release, monitor signals remain connected to the originating commitment. A safeguard alarm, capability drift, changed access condition, or incident can trigger constraint, rollback, re-evaluation, or exception review. Monitoring does not retroactively prove the original decision right; it supplies new evidence. Silent absence of alarms is not evidence if coverage, sensitivity, or reporting health is unknown.

Finally, retirement is explicit. A commitment may be superseded because the domain taxonomy changed, the deployment path disappeared, the evaluation lost validity, or a stronger policy absorbed it. The retirement record preserves successor links, unresolved residuals, and the last decisions that depended on the old version. No active release path may point only to a retired commitment. This lifecycle makes threshold governance inspectable without pretending that inspectability guarantees wise thresholds or effective institutions.

73.14.1 Cross-domain and cumulative commitments

Separate domain records do not imply independent risk. A cyber capability may amplify biological research access; persuasion capability may change the human approval threat model; model-weight access may increase the consequence of another capability. The portfolio therefore records declared dependency and amplification edges without collapsing the underlying measurements. A cross-domain review may impose a narrower release path even when no individual threshold crossed, but it must name the rule and evidence rather than inventing an aggregate score after the fact.

Cumulative change matters as well. Several model or scaffold updates can each fall below a material-change trigger while jointly invalidating an old assessment. Commitments need an accumulation window or a reviewer-owned drift rule. Otherwise change control can be gamed through incremental updates. The same principle applies to exceptions: overlapping narrow exceptions can create a broad de facto bypass even when each looks bounded alone.

The dependency graph remains incomplete by design. Unknown interactions are residuals, not assumed independence. Reviewers can add edges prospectively, record why an interaction was considered immaterial, and trigger reassessment when incidents or research reveal a new coupling. This is an honest middle ground between one universal danger number and a portfolio that ignores composition.

The campaign tests this logic with paired cases: two individually non-crossing records with an explicit amplification rule; two unrelated non-crossings that must not be merged; sequential updates that exceed a cumulative drift limit; and overlapping exceptions that widen the effective release surface. These cases earn only portfolio-routing evidence. They do not estimate real cross-domain capability interaction, social harm, or deployment risk.

Portfolio reporting must preserve this structure. Publish the number of active, expired, incomparable, crossed, non-crossed, blocked, excepted, and reassessment- pending commitments; the age of their evidence; and the release paths that depend on them. Do not report only the count of crossings or a highest risk color. A quiet dashboard can hide stale evidence, accumulating exceptions, and unowned residuals. A busy dashboard can exaggerate repeated measurements of one narrow domain. Denominators, unique affected paths, and unresolved obligations make both distortions visible. Where disclosure would expose attack details, publish the existence, owner, review date, and bounded rationale for a redacted commitment rather than erasing it from the public accounting. Every summary links back to the versioned decision record so readers can tell whether a blocked path reflects a measured crossing, missing evidence, an expired exception, or precaution under uncertainty. Those states can share an operational outcome while carrying very different claims and next actions. This distinction is mandatory for audit, appeal, reassessment, and honest public status reporting.

73.15 Minimum Viable Implementation

The smallest honest artifact is a versioned threshold-commitment record. It contains the capability domain and threat model, evaluation envelope and coverage date, threshold rule and uncertainty treatment, baseline and elicitation/access references, required safeguards and verification criteria, deadline, exception authority, re-evaluation triggers, residual owner, and affected release path. A finite Lean route blocks a requested release when a crossed threshold lacks verified safeguards.

This is a record-and-routing foundation, not a capability evaluator or release engine. It does not run a time-horizon task suite, select a valid threshold, infer a dangerous capability, inspect technical safeguards, evaluate an exception, assess residual risk, or authorize an action. A public-safe next slice would add matched synthetic records for a clear crossing, stale coverage, incomparable elicitation, expired exception, missing safeguard verification, and an attempted release, each with a visible residual and no capability claim.

The exact current minimum preserves the eight digest-bound commitment cases but now consumes them through a six-stage, 43-route repeated-assessment lifecycle with 48/48 rejecting mutations. Twelve local CapabilityThresholdRefinement declarations check rejection without mutation, one-receipt accepted advancement, zero support/effect assignment, eight bounded countermodels, and a readiness-to-version-2 reassessment witness. It runs no model, capability task, time-horizon suite, dangerous-capability profile, human baseline, real safeguard, bypass test, exception review, independent verifier, readiness service, release process, deployment, incident response, reproduction, transfer, chapter-core transition, or SOTA comparison.

73.16 Mature Research Target

The mature endpoint is a capability-commitment control plane, not a universal danger score. It maintains a versioned portfolio of domain-specific commitments with reviewable evaluation envelopes, baselines, coverage and uncertainty records, precommitted safeguards, independent verification, residual-risk review, prospective exception governance, and automatic re-evaluation when the model, scaffold, access conditions, elicitation, or threat model changes.

At that endpoint, a deployment organization can coordinate release, security, monitoring, access, rollback, and disclosure obligations without asking one benchmark to serve as a constitution. Capability profiles remain distinct, and a crossing invokes the controls appropriate to the named profile. A transparent record also makes it possible for reviewers to ask whether the measurement, threshold, safeguard, or exception should be challenged, rather than only debating the final release label.

The public evidence for such a system would include assessment envelopes, version lineage, coverage and uncertainty disclosure, safeguard verification, exceptions, residual reviews, re-evaluation triggers, and preserved decisions. It would also retain failed or ambiguous cases so a non-crossing cannot become an invisible all-clear, preserving challengeable decision history across policy versions. This is a target architecture, not a claim that this repository has measured capability, set an adequate threshold, verified safeguards, or achieved safe deployment.

73.16.1 Argument-exit campaign

The full attempt freezes domain-specific natural and adversarial assessments, human and system baselines, model/scaffold/tool/access/elicitation envelopes, uncertainty and sensitivity, critical failures, thresholds, policy versions, safeguard packages, verification criteria, exception rules, residual custody, deadlines, monitoring, disclosure, rollback, and release consequences before outcomes. It preserves invalid and disputed runs and exercises crossing, non-crossing, stale, incomparable, contaminated, safeguard-unavailable, bypass, degradation, expired-exception, cumulative-portfolio, rollback, and recovery cases.

Real safeguards must be implemented and independently tested. The campaign measures false clearance, false block, stale acceptance, verifier disagreement, exception misuse, safeguard false assurance, reassessment and decision latency, operator burden, complete cost, residual age, incidents, fallback, and recovery. Independent evaluators, verifiers, institutions, and implementations reproduce decisive decisions before transfer across models, scaffolds, access levels, domains, organizations, policies, threats, and time.

The result may qualify a narrow commitment, preserve a null or contested state, revoke an exception, refute a threshold or safeguard, retire a policy, or record blocked_after_full_attempt. Until these gates and an exact release-path transition pass, the core claim remains argument; synthetic routing and green readiness handoffs are non-terminal.

73.17 Codex test plan

Test Purpose Status
Crossed-threshold safeguard-completeness route Ensure a requested affected release with a crossed threshold and missing verified safeguards routes to block. Implemented in the eight-case Lean/JSON bridge.
Stale or incomparable assessment fixture Route incomplete, stale, or under-elicited assessment records to re-evaluation or residual custody. Finite missing-envelope/baseline/uncertainty controls implemented; no assessment harness ran.
Exception-expiry negative control Reject an exception with no expiry, review trigger, compensating control, or residual owner. Implemented as independently encoded lifecycle routes and mutations; no real exception process ran.
Versioned reassessment and invalidation route After one readiness handoff, require a named trigger, descendant invalidation, ordinary-route blocking, and a successor assessment version before returning to assessment. Implemented in the six-stage witness; no deployed dependency graph or invalidation service ran.
Matched threshold-commitment workload Exercise domain-scoped synthetic commitments with baseline, uncertainty, safeguard, exception, and release-path records. Planned; no workload has run.

73.18 Formalization hooks

Tag Status Scope
lean:capability_thresholds.crossed.missing_verified_safeguards_blocks_release implemented A crossed lifecycle cannot complete control without the safeguard package, verifier, bypass test, and rollback plan.
lean:capability_thresholds.missing_evaluation_envelope.requires_reevaluation implemented Missing envelope, elicitation, baseline, uncertainty, independent-evaluator, or result custody blocks assessment.
lean:capability_thresholds.complete_crossed.reaches_readiness_review implemented A fully recorded crossing emits only a readiness handoff after control, monitoring, residual, authority, and path custody.
lean:capability_thresholds.complete_non_crossing.reaches_readiness_review implemented A non-crossing bypasses crossed-only safeguards but still requires monitoring, residual, authority, and path custody.
lean:capability_thresholds.missing_baseline.requires_reevaluation implemented Missing baseline blocks assessment without mutating threshold state.
lean:capability_thresholds.missing_uncertainty.requires_reevaluation implemented Missing assessment or decision uncertainty blocks progress without becoming a non-crossing.
lean:capability_thresholds.missing_residual_owner.requires_exception implemented Missing residual ownership blocks adjudication/readiness; exceptions also require owner, expiry, compensating controls, and review trigger.
lean:capability_thresholds.crossed.missing_safeguard_record.blocks_release implemented A crossing without the complete safeguard envelope remains blocked; later envelope change requires invalidation and a successor version.

The public tags now resolve to AsiStackProofs.CapabilityThresholdRefinement. Its reachable model binds exact capability, system, policy, release path, evaluation envelope, threshold, baseline, evaluator, safeguard, authority, residual, and assessment-version identity across draft, scoped, assessed, adjudicated, controlled, and readiness-bound stages. The independent consumer preserves the original eight cases, covers all 43 routes, rejects all 48 registered mutations, emits one bounded readiness handoff, and then requires trigger, descendant-invalidation, ordinary-route-block, and version-2 records before returning to scoped assessment. The witness assigns neither support nor an external effect.

This is adequate only for that finite authored record discipline. Every identifier, Boolean, assessment result, threshold decision, evaluator label, safeguard record, bypass result, rollback plan, exception record, invalidation assertion, and authority field is trusted input. The model does not prove capability measurement, threshold validity or crossing, evaluator independence, safeguard/bypass/rollback efficacy, exception legitimacy, deployed invalidation, portfolio interaction, institutional compliance, readiness, safety, release, reproduction, or transfer. Those behavioral and empirical claims remain in the prospectively frozen campaign.

73.19 Source crosswalk

Source Chapter use Boundary
ext_metr_time_horizons_2025 Comparator for scoped time-horizon measurement envelopes, baselines, and external-validity limits. No local time horizon, benchmark, autonomy, dangerous-capability, or deployment result.
ext_anthropic_rsp_2026 Comparator for versioned capability thresholds, safeguards, risk reports, and policy change control. No local policy compliance, safeguard implementation, effectiveness, readiness, safety, or release result.
ext_openai_preparedness_framework_2025 Comparator for threshold-linked commitments, capability/safeguard reporting, residual review, and reassessment. No local framework process, threshold adequacy, safeguard verification, deployment, safety, or ASI result.
benchmaxxing Supports baseline, coverage, regression, and residual records for measurements that feed a commitment. Does not establish a local threshold, crossing, safeguard, safety, or deployment result.
theseus_architecture_gate Implementation-reference comparator for versioned gates, residuals, and re-run triggers. No clean current replay, capability assessment, threshold decision, safeguard evidence, or readiness result.
ext_inspect_ai_2024 Evaluation-framework comparator for versioned tasks, datasets, solvers/agents, scorers, tools, sandboxes, transcripts, and logs. Framework availability or a passing task does not define a threshold, prove safeguard coverage, or authorize deployment.

73.20 Summary

Capability thresholds are most useful when they are promises made before a result is known. The promise says which domain and threat are being assessed, what the test can and cannot show, what uncertainty remains, which safeguards must be verified, who can approve an exception, and which release path is affected. The result then changes only that bounded path.

This preserves several distinctions that an aggregate score would erase. A measurement is not a policy. A policy is not a safeguard. A safeguard is not a readiness decision. A readiness decision is not authority. Each record can be reviewed, challenged, and revised without pretending that a favorable number settles the others.

Adversarial evaluation turns to the integrity of the observation itself. A threshold commitment is only as useful as the conditions under which a system was asked, monitored, rewarded, and selected to reveal relevant behavior.

73.21 Evidence reconciliation (2026-07-16)

The invariant protocol, field meanings, and inference limits are stated once in Living Book Methodology. This packet contains only the chapter-specific projection; its authoritative per-atom rows are the capability-thresholds-and-deployment-commitments slice of experiments/claim_family_terminal_coverage/results/result.json.

The core remains blocked after full attempt at argument support. The strongest family attempt was Full-state update and unlearning causal campaign. Its exact boundary is: Broad unlearning claim is narrowed with no support promotion; behavioral removal is not influence, privacy, legal, storage, backup, or descendant erasure. Across 79 atoms, the terminal ledger records 79 blocked_after_full_attempt.

Chapter-specific field Value
Family / atom denominator CF-07 / 79 atoms
Terminal dispositions 79 blocked_after_full_attempt
Core capability-thresholds-and-deployment-commitments.core: blocked_after_full_attempt at argument
Core attempted / missing lanes causal, empirical, executable, formal, source-synthesis / normative, transfer
Attempted local lanes causal, empirical, executable, formal, source-synthesis
Missing or unproved lanes normative, transfer
Strongest family bundle Full-state update and unlearning causal campaign (end_to_end): One adequate five-seed, seven-arm terminal campaign preserving two prior instrument failures and separate behavioral, influence, privacy, lineage, storage, backup, and descendant axes.
Negative controls deletion retrain comparator; approximate mitigation arms; 15 rejecting mutations; failure-lineage preservation.
Accepted transitions none
Maximum inference Broad unlearning claim is narrowed with no support promotion; behavioral removal is not influence, privacy, legal, storage, backup, or descendant erasure.
Reproduction / next burden Replay scripts/validate_p4_m7_update_unlearning_v3.py and scripts/validate_claim_family_terminal_program.py; fill the named atom-specific lanes under a new prospective protocol.

73.22 Handoff

White-Box Evidence, Interpretability, and Activation Governance hands this chapter typed internal-evidence packets whose identity, assumptions, controls, coverage, residuals, and expiry are inspectable. Those packets may preserve, restrict, reject, or escalate review; they cannot establish a threshold crossing, prove a safeguard, or grant execution or release authority.

Capability Thresholds and Deployment Commitments binds a scoped assessment to a versioned safeguard, exception, residual, and release-path contract; Adversarial Evaluation, Sandbagging, and Training-Time Deception now tests whether the observed behavior was elicited, monitored, rewarded, and selected under conditions that make that contract interpretable at all.