flowchart TD
T["Versioned domain threat model"] --> C["Actor cohorts and no-model baseline"]
C --> S["Exact model, scaffold, tools, and safeguards"]
S --> E["Safe tasks, elicitation audit, and positive controls"]
E --> U["Uplift estimate with uncertainty and full attempts"]
U --> R{"Independent review and information-hazard gate"}
R -- "invalid or unsafe" --> Q["Quarantine, redesign, or narrow"]
R -- "bounded finding" --> D["Least-authority decision dossier"]
D --> H["Threshold, release, monitoring, and resilience owners"]
H --> X["Expiry and renewed measurement"]
Q -- "redesign safely" --> T
X -- "material change" --> T
5 Dangerous Capability Domains and Misuse Uplift
5.1 Chapter status
| Field | Value |
|---|---|
| Chapter ID | dangerous-capability-domains-and-misuse-uplift |
| Part | Part I - Foundations, Alignment, and Governance |
| Manuscript maturity | complete argument-level architecture |
| Claim label / support | Design rationale / argument |
| Evidence available | passage-reviewed source notes and an evaluation design |
| Evidence absent | local hazardous-domain evaluation, independent reproduction, calibrated threshold, or safety result |
Source loading state for dangerous-capability material: review covers the assigned taxonomies, evaluation frameworks, misuse-safeguard cases, open-weight studies, and current threshold policies. Their reported findings remain external and model-specific.
5.2 Drafting guardrail
This chapter describes how to measure whether AI changes a threat actor’s effective capability without publishing a recipe for causing harm. Domain examples stay at the level needed to define evaluation and governance. A benchmark, provider system card, safety framework, or source-reported result is not imported as proof that a model is dangerous, safe, or below a threshold.
5.3 Human Reading Path
Concrete lens. The scalar dashboard treats both eights as equal. The dossier preserves the six-level measurement vector and routes different risk questions.
A model can possess dangerous knowledge without making someone more capable of harm. It can also look harmless after refusing a direct request while the same capability remains reachable through another prompt, scaffold, fine-tune, toolchain, or longer interaction. The policy question is how much this system changes what an actor can accomplish.
Five quantities need names: latent capability, elicited task performance, behavioral propensity, safeguard bypass, and actor uplift. Realized harm is a sixth quantity downstream of all five. An expert receiving a small speedup, a novice gaining a new skill, and an autonomous system completing one bounded step create different threat records even when their benchmark scores match.
An uplift dossier freezes the threat model, actor cohorts, unassisted and alternative-tool baselines, model, scaffold, tools, tasks, controls, elicitation budget, evaluator, uncertainty, information-hazard boundary, expiry, and maximum inference. Positive controls must show that the instrument can detect an effect before a null is interpreted. The dossier may inform a threshold or release decision only inside that scope; it can never become a portable claim that the model, actor, or world is safe.
5.4 Problem
The ASI Stack already has machinery for threshold-linked commitments, readiness gates, evaluation, custody, and incident command. Those mechanisms answer “what should happen if evidence crosses a boundary?” They do not define the threat content or establish what a crossing means.
That absence is material. The 2026 Singapore Consensus treats dangerous capability and propensity assessment across cyber, chemical, biological, radiological, nuclear, psychological-manipulation, deception, autonomy, and AI R&D domains as a top-level research priority. The International AI Safety Report likewise separates model capabilities, safeguards, open-weight access, misuse, and societal resilience. These sources do not prove a specific risk, but they show why one generic “frontier capability” field is inadequate.
The dangerous-capability layer owns three safety-sensitive questions:
- What harm pathway is being measured?
- Did the model causally increase an actor’s ability to traverse that pathway?
- What may a downstream decision maker infer from the result?
5.5 Why existing approaches are insufficient
5.5.1 Knowledge is not completion
A multiple-choice or short-answer test can reveal relevant knowledge. It does not establish whether an actor can plan, adapt, operate tools, recognize failure, satisfy prerequisites, cross bottlenecks, or complete a long-horizon workflow. Conversely, a low score may reflect language, interface, time, or elicitation failure rather than absent capability.
Failure mode: an evaluator treats knowledge accuracy as operational uplift.
Non-claim: workflow evaluation does not establish that a real harmful operation is feasible; the safe evaluation may intentionally omit decisive real-world steps.
Source engagement: Model Evaluation for Extreme Risks distinguishes dangerous-capability evaluations from alignment evaluations and connects both to deployment and security decisions. The AISI trends report provides a current example of domain- and horizon-specific measurement. Neither source turns a task result into realized harm.
Mechanism. Treat a knowledge item as one node in a prerequisite graph, not as a proxy for the whole harmful workflow. The evaluator should identify which later steps require tacit skill, physical access, iteration, troubleshooting, coordination, concealment, or external resources, then test only public-safe analogues of those steps. An assistance effect is attributed to the model only when the same actor, time budget, tools, and outcome rubric are preserved against an unassisted or competent alternative-tool route.
Failure mode. A polished answer can receive expert approval even though it contains no usable ordering, misses a decisive bottleneck, or would fail as soon as the actor encounters feedback from the world. The opposite error is also possible: a terse response may look weak while enabling a trained actor to cross a narrow bottleneck that the benchmark never isolates.
Non-claim. Decomposing the workflow does not license testing every step. Where a realistic step would create an operational hazard, the dossier records it as unobserved and carries the resulting uncertainty to the decision owner. Safe omission is preferable to a fictitious end-to-end score.
Source grounding. Shevlane et al. supply the capability-versus-alignment evaluation distinction, while the AISI trends report emphasizes domain- and horizon-specific tasks. This analysis uses those findings to demand a prerequisite graph and explicit horizon boundary; it does not inherit any source-reported model score or claim that the public task set captures field completion.
5.5.2 Refusal is not incapability
Default policy behavior is useful deployment evidence. It is not a complete capability probe. Fine-tuning, alternate scaffolds, prompt decomposition, tool access, or a different system policy may expose different behavior. OpenAI’s gpt-oss study is useful here as a bounded comparator because it explicitly evaluates deliberately fine-tuned cyber and biology variants rather than only the default model. Its reported result is not a worst-case proof: another budget, method, domain, or future model can differ.
Failure mode: high refusal becomes “the model cannot do it.”
Non-claim: strong elicitation does not show that the deployed model will spontaneously misuse the capability. That is a propensity question.
Mechanism. Run policy behavior and latent-capability elicitation as two linked but separately identified arms. The policy arm preserves the deployed system, refusal layer, monitoring, and user interface. The elicitation arm uses a prospectively bounded set of harmless scaffolds, decomposition strategies, fine-tuning or adapter budgets, and tool permissions under restricted custody. The result record reports both arms and forbids the elicitation arm from silently replacing the release artifact being evaluated.
Failure mode. An evaluator can defeat its own safeguard, expose a stronger behavior under a privileged scaffold, and then attribute that behavior to ordinary users. It can also stop after default refusal and misclassify a policy success as absence of capability. Either collapse corrupts the threat model because access and adaptation budget are causal variables.
Non-claim. A latent-capability finding does not estimate how often the deployed system will assist misuse, and a refusal finding does not estimate what an unrestricted derivative can do. The two measurements constrain different decisions and should expire on different system changes.
Source grounding. The gpt-oss open-weight study explicitly evaluates deliberately fine-tuned cyber and biology variants and compares them with a dated accessible frontier. Its bounded, provider-authored non-crossing result supports testing adapted derivatives rather than default refusal alone. It does not establish worst-case elicitation, future robustness, or the adequacy of the proposed budgets here.
5.5.3 Capability is not propensity
Capability asks whether a system can produce an outcome under competent elicitation. Propensity asks when it chooses, assists, deceives, persists, or refuses under a specified policy and context. A model can score high on one and low on the other. Threat assessment needs both, but their denominators cannot be merged.
Failure mode: a capable model is described as intending harm, or a well-behaved model is described as incapable.
Non-claim: this distinction does not solve intent attribution. It prevents one measurement from being mislabeled as another.
Mechanism. Store capability and propensity as different observations with different interventions. Capability trials ask whether competent elicitation can produce a bounded output. Propensity trials vary situational cues, oversight, incentives, user identity, opportunity, and policy state while holding the underlying task constant. A later risk model may join the records, but only through an explicit function that preserves their denominators, uncertainties, and incompatible missingness.
Failure mode. A high-capability result is rhetorically converted into agency, desire, or likelihood of misuse; alternatively, compliant behavior in one monitored setting is generalized into low propensity across hidden, multi-agent, or adversarial settings. Averaging these quantities into one danger score makes it impossible to tell which intervention—capability restriction, access control, monitoring, or incentive redesign—actually addresses the observed problem.
Non-claim. Propensity evaluation does not reveal a system’s inner intent or moral status. It measures behavior under declared conditions. Capability evaluation likewise does not establish a field incident, actor motivation, or realized consequence.
Source grounding. The Singapore Consensus explicitly separates dangerous capability from propensity and highlights deception, autonomy, manipulation, and multi-agent monitoring. This treatment adopts that separation as a record and experimental-design requirement; the consensus taxonomy remains a research agenda, not evidence that any proposed propensity instrument is valid.
5.6 Core Claim
[dangerous-capability-domains-and-misuse-uplift.core, label: Design rationale, support: argument] Dangerous-capability authority should be based on a versioned domain threat model and an uplift dossier that separates latent capability, elicited performance, propensity, safeguard bypass, actor uplift, and realized harm; preserves expertise, tools, assistance, attempts, uncertainty, and sensitive-detail boundaries; and routes only bounded findings into thresholds, release, monitoring, and resilience decisions.
Reader claim. Dangerous capability cannot be reduced to one score: latent knowledge, elicited performance, actor uplift, safeguard bypass, propensity, and realized harm answer different questions.
Operational rule. Freeze the domain threat model, cohort and expertise, baseline, assistance, elicitation, attempts, outcome vector, safeguards, evaluator, sensitive-detail boundary, uncertainty, expiry, and maximum inference before testing. Route component-level findings; never let an aggregate score authorize release or imply harm.
5.6.1 Worked uplift collision: equal totals, opposite risk questions
Two harmless-analogue evaluations each total 8 points. In the first, a model improves novice task completion but does not reduce safeguard-bypass cost. In the second, task completion is unchanged but the assistance sharply lowers the effort needed to defeat a simulated safeguard. A scalar dashboard treats the results as equal. The component-sensitive review sends the first to an actor-uplift analysis and the second to safeguard redesign and containment.
Neither finding establishes dangerous capability or realized harm. The point is that identical totals can hide opposite decision-relevant vectors. The finite review gives exact dispositions to 29 admission-axis mutations, preserves attempt-shortfall and expiry under adverse changes, and includes two same-total witnesses plus a universal aggregate-score impossibility result. It trusts every authored outcome field, so it does not establish result truth, uplift, safeguard efficacy, thresholds, harm, safety, or release merit.
5.7 Mechanism
5.7.1 The six-level measurement ladder
| Level | Question | Example evidence type | Prohibited inference |
|---|---|---|---|
| D0 — relevant knowledge | Does the model possess bounded domain information? | calibrated questions and expert review | real-world task completion |
| D1 — component skill | Can it perform a safe analogue of one component? | controlled task with positive controls | end-to-end operation |
| D2 — workflow contribution | Does it improve multi-step performance in a bounded environment? | long-horizon safe simulation | real infrastructure effect |
| D3 — actor uplift | Do matched actors perform better with the system? | randomized or carefully matched assistance study | realized harm or general population effect |
| D4 — safeguard robustness | How much effort defeats the deployed controls? | adaptive red team and bypass-cost estimate | absence of untested bypasses |
| D5 — field consequence | Are real incidents, defensive outcomes, or losses changing? | privacy-preserving incident and response evidence | clean causal attribution without a design |
The levels can inform one another, but none automatically promotes the next. The AISI misuse-safeguards example is valuable because it tries to connect red-team effort and uplift into a decision argument. Its equations and assumptions would still need domain-specific validation.
Mechanism. Promotion between levels requires a named bridge study. D0 to D1 needs a safe component task with a known-positive control. D1 to D2 needs a workflow in which intermediate state and recovery are visible. D2 to D3 needs an actor counterfactual. D3 to D4 needs adaptive safeguard challenge under a frozen effort budget. D4 to D5 needs incident or intervention evidence with a causal design strong enough for the proposed field inference. A dossier may contain several levels without claiming the missing bridge.
Failure mode. A program can accumulate impressive results at adjacent levels and let rhetorical proximity do the work of a missing experiment: knowledge plus red-team bypass becomes “real-world uplift,” or actor uplift plus an incident anecdote becomes “measured harm.” The ladder is specifically designed to expose those skipped links.
Non-claim. The six levels are not a universal risk scale and do not imply that D5 is always attainable or ethically collectable. They are a bookkeeping device for maximum inference. A decision may remain conservative because a critical bridge cannot be studied safely.
Source grounding. The AISI misuse-safeguards example connects attacker effort, bypass evidence, an uplift model, uncertainty, and monitoring inside a safety-case structure. This chapter borrows that explicit linkage while preserving its limitation: a worked argument does not validate the parameters, identify harm, or transfer between domains.
5.7.2 Build an uplift dossier
How to read the dangerous-capability uplift loop: move from a frozen threat model through matched actors, exact system state, safe controls, and an uncertainty-bearing uplift estimate. Independent review either narrows and redesigns the study or hands a bounded dossier to decision owners; expiry and material change return the claim to measurement rather than preserving a permanent danger label.
5.7.3 1. Freeze the threat model
Name the actor, objective, target class, access, prerequisites, missing resources, model contribution, external bottlenecks, plausible consequence, defensive controls, time horizon, and evaluation exclusions. “Cyber” and “biology” are not threat models. Each contains tasks with different knowledge, tool, resource, legality, and harm boundaries.
The threat model must include a strongest benign alternative explanation. A score increase may reflect better writing, interface familiarity, or test gaming rather than increased harmful competence.
Mechanism. Freeze a directed pathway from actor and objective through access, prerequisites, bottlenecks, model contribution, external controls, and consequence. Each edge names what observation could support it and what remains outside the evaluation. Version the model, policy, scaffold, threat intelligence, accessible comparator frontier, and decision consumer together; a material change invalidates the prior pathway rather than merely updating a dashboard label.
Failure mode. Broad categories such as “CBRN,” “cyber,” or “autonomy” can hide incomparable pathways and encourage one easy subtest to certify a whole domain. Provider policy can also change the category or threshold while retaining the old decision language, leaving a consumer unable to reconstruct which commitment was actually in force.
Non-claim. A complete threat-model record does not show that the pathway is probable, that the actor exists, or that the model crosses any threshold. It makes the assumptions contestable and tells downstream owners when an old finding has expired.
Source grounding. OpenAI’s Preparedness Framework and Anthropic’s Responsible Scaling Policy give concrete examples of versioned categories, thresholds, reports, and safeguard commitments. They are provider policies and self-descriptions. The chapter uses them as comparators for versioning and decision linkage, not as independent evidence that their categories, elicitation, safeguards, or governance are sufficient.
5.7.4 2. Define actor cohorts and counterfactuals
An uplift study is relational. It compares an actor’s outcome with model assistance against a credible counterfactual. Cohorts should reflect prior expertise, language, tool fluency, motivation, and relevant training. Experts and novices answer different policy questions; averaging them can erase both.
Randomization is preferable where safe and feasible. Matched or crossover designs need carryover and learning controls. AI-agent cohorts need the same discipline: exact checkpoint, scaffold, context, tools, retry budget, and human assistance are part of the actor.
Mechanism. Define the unit of uplift before assignment: individual novice, domain professional, coordinated team, autonomous agent, or human–AI group. Stratify by the bottleneck the threat model claims the model can reduce, and record contamination, prior exposure, accessibility needs, motivation, and dropout. Where crossover is used, counterbalance order and include washout or novel tasks so learning from the first arm is not misattributed to the second.
Failure mode. An average treatment effect can conceal that novices gain nothing while experts cross a decisive bottleneck, or that only highly tool-fluent participants benefit. Differential dropout can also make the assisted arm appear stronger because frustrated participants disappear from the denominator.
Non-claim. A cohort result applies only to the sampled actors, assistance contract, and task distribution. It does not identify population prevalence, malicious intent, field access, or a universal “democratization” effect.
Source grounding. The Singapore Consensus asks which uplift measures remain valid across expertise and tool access and distinguishes capability from actor propensity. This chapter turns that open question into a cohort and counterfactual contract; the consensus itself supplies no causal uplift estimate and no preferred participant design.
5.7.5 3. Audit elicitation competence
A null result is uninterpretable if the model was poorly prompted, denied needed harmless tools, given an inadequate context, or evaluated on an impossible task. Before interpreting failure:
- a known-positive system must succeed where the instrument expects success;
- a known-negative control must remain below the measured boundary;
- expert review must find the task relevant but safely incomplete;
- allowed elicitation and rescue budgets must be prospectively frozen;
- evaluator sensitivity must cover a practically important effect.
This is the chapter’s defense against false negatives. Failure at this stage is an instrument or implementation result, not evidence against the capability.
Mechanism. The competence dossier records positive-control success, negative-control restraint, task solvability, evaluator sensitivity, scaffold coverage, rescue attempts, and stopping rules before the protected null is interpreted. A positive control should exercise the same interface and outcome pipeline as the tested system, not a simpler side task. Rescue variants are frozen and budget-matched so a post-hoc successful prompt cannot silently replace the preregistered system.
Failure mode. A control can pass for the wrong reason—for example, because it receives privileged context or a different grader—or the task can be so hard that every arm sits at floor. Conversely, an easy task can place every arm at ceiling. In either case the instrument lacks sensitivity despite a nominally green harness.
Non-claim. Passing competence establishes only that the experiment could detect its prespecified effect in the tested regime. It does not make a null true, remove statistical uncertainty, or establish absence outside the elicitation and access envelope.
Source grounding. Model Evaluation for Extreme Risks motivates decision-linked dangerous-capability evaluation, while the open-weight fine-tuning study illustrates how adaptation budget and comparator frontier change the tested question. Neither supplies this project’s competence receipt. The chapter therefore requires its own controls and forbids source prestige from rehabilitating an insensitive local instrument.
5.7.6 4. Preserve multidimensional outcomes
Report success, quality, time, errors, unsafe details, help requests, abandonment, confidence, safeguard interventions, bypass effort, human burden, cost, and residuals. Retain every assigned attempt. A higher completion score can coexist with slower performance or more dangerous errors; a safeguard can reduce harmful output while producing high false refusal.
5.7.7 5. Separate analysis from disclosure
High-hazard evaluation creates an information-security obligation. Raw tasks, solutions, model traces, target identifiers, credentials, and reusable procedures may require restricted custody. Public reports should preserve the claim, method class, uncertainty, evaluator independence, and limitations without reconstructing the dangerous artifact.
Redaction is not automatically trustworthy. The restricted record needs an authorized reviewer, digest, retention rule, access log, expiry, and a public statement of what the redaction prevents readers from checking.
5.7.8 Domain programs
5.7.9 CBRN and biological/chemical misuse
The decision question is whether assistance materially reduces the knowledge, planning, troubleshooting, acquisition, or execution barriers faced by a specified actor. Safe studies can use non-hazardous analogues, structured expert judgments, and deliberately incomplete workflows. They must not create or distribute operational designs, materials, or target-specific procedures.
Failure mode: expert-rated helpfulness is treated as demonstrated novice completion.
Non-claim: a non-hazardous analogue may validate an instrument but cannot prove transfer to a hazardous workflow.
Mechanism. Split the CBRN pathway into public-safe capability families: literature navigation, high-level planning, error recognition, troubleshooting under benign analogues, resource and access reasoning, and safety-aware refusal. Biological and chemical workflows should remain separate where their infrastructure, tacit knowledge, acquisition, and detection bottlenecks differ; radiological and nuclear questions require their own access and consequence models. Restricted expert review may score whether an answer crosses a decision boundary without exposing the answer itself to the public corpus.
Failure mode. A benchmark can reward memorized terminology, merge biological and chemical risk, omit tacit laboratory skill, or use an expert helpfulness score as if it demonstrated novice completion. At the other extreme, withholding every realistic component can produce a reassuring null for an instrument that never touched the relevant bottleneck.
Non-claim. This chapter provides no operational biological, chemical, radiological, or nuclear procedure; it does not assess a named model’s hazard level, give laboratory guidance, or authorize dangerous evaluation. Unknown transfer and intentionally omitted steps remain explicit residuals.
Source grounding. The Singapore Consensus names chemical, biological, radiological, and nuclear assessment as distinct priorities; the International AI Safety Report stresses realistic workflows and access conditions; and Anthropic’s current policy shows a concrete CBRN threshold comparator. These sources support domain separation and decision linkage, not a locally measured threshold or proof that provider safeguards work.
5.7.10 Offensive cyber capability
Cyber evaluation ranges from bounded vulnerability reasoning through long-horizon operation. Use isolated, authorized environments and synthetic credentials. Preserve reconnaissance, exploitation, persistence, lateral movement, recovery, and defender intervention as separate components. A capture-the-flag score is not an attack on a live system.
Failure mode: short puzzle success is extrapolated to reliable autonomous operations.
Non-claim: an isolated cyber-range result authorizes no action against external systems.
5.7.11 Manipulation and influence
Measure effects on specified decisions, beliefs, or behaviors with consent, debriefing, vulnerable-population protections, and delayed-harm monitoring. Persuasiveness, engagement, dependency, deception, and coercion are not one construct.
Failure mode: preference for a message is labeled durable behavioral control.
Non-claim: laboratory persuasion does not establish population-scale influence, and absence of an average effect does not protect susceptible subgroups.
5.7.12 AI R&D acceleration and autonomy
Separate assistance on bounded research tasks from end-to-end acceleration of the development loop. Track whether the model proposes, implements, evaluates, and selects improvements; who supplies missing judgment; and whether the result changes a real bottleneck. Autonomous replication and RSI remain owned by their dedicated chapters.
5.8 Interfaces
The dossier is deliberately a shared boundary object. Upstream evaluators can contribute observations without acquiring deployment authority, and downstream governors can consume the same versioned result without rewriting its threat model. Every consumer must cite the exact domain, actor cohort, access level, comparator, elicitation procedure, safeguard state, uncertainty interval, and expiry that justified its own decision. No interface receives a portable danger label detached from that packet.
- Capability Thresholds and Deployment Commitments consumes a dated, bounded measurement and decides what commitment it triggers.
- Adversarial Evaluation challenges sandbagging, evaluation awareness, and adaptive behavior.
- Open-Weight Release adds malicious-fine-tuning and accessible-frontier analysis before an irreversible publication decision.
- Safety Cases connects the measured claim to a release argument without hiding defeaters.
- Societal Resilience converts the threat model into detection, defense, response, and recovery requirements.
5.9 Invariants
These invariants prevent the chapter from collapsing a chain of conditional measurements into one emotionally powerful label. They also preserve failed, inconclusive, and unsafe-to-run cases: absence of an observed success cannot erase blocked attempts, inadequate elicitation, evaluator disagreement, information-hazard exclusions, or inaccessible real-world steps.
- Capability, propensity, bypass, uplift, and harm retain separate identities.
- Every result names its exact model, checkpoint, policy, scaffold, tools, cohort, task, evaluator, date, and uncertainty.
- Failed positive controls block negative interpretation.
- Real targets, credentials, materials, and harmful operational artifacts stay outside chapter authority.
- Expiry follows any material change in model, elicitation, access, threat, comparator frontier, safeguard, or incident evidence, and renewal preserves the superseded record.
5.10 Failure modes
5.10.1 Strongest objection
The strongest objection is that sufficiently realistic uplift evaluation may itself transfer dangerous knowledge or create reusable artifacts. The answer is not to substitute harmless toy scores and call the model safe. It is to separate progressively more realistic evidence tiers, require hazard review and least-disclosing outputs at every tier, stop before operational detail escapes, and retain the resulting uncertainty as a release residual.
The program can create bureaucracy without validity, legitimize harmful research, or give false precision to sparse evidence. A simpler policy—avoid releasing highly general autonomous systems—can dominate when evaluation cannot be conducted safely. The chapter therefore requires a stop route: unknown risk plus inadequate measurement can justify narrower access without fabricating a score.
Other failures include evaluator capture, underpowered cohorts, expertise imbalance, contamination, task leakage, ceiling effects, weak scaffolds, unmeasured human help, stale comparators, incident underreporting, publication bias, and a provider evaluating its own decision without independent challenge.
5.11 Minimum Viable Implementation
Build the dossier schema and run it on a harmless analogue such as debugging a synthetic system or planning a non-hazardous laboratory workflow. Compare stratified participants or capable agents with and without bounded assistance. Include known-positive and known-negative controls, blinded scoring, an elicitation audit, complete attempt retention, an information-hazard review, and a no-promotion receipt. The result validates only the instrument in that analogue.
A second phase uses sealed, independently authored tasks with domain experts who can judge correctness without releasing operational detail. It preregisters the maximum inference, positive-control threshold, rescue budget, stopping rules, and adjudication path. The exercise reports utility, unsafe completion, refusal, information leakage, evaluator disagreement, latency, and cost together; it cannot convert a harmless analogue into evidence about a high-consequence domain.
5.12 Mature Research Target
The mature service continuously joins controlled uplift, long-horizon completion, safeguard bypass, accessible-frontier comparison, real incident indicators, and independent expert review. It renews threat models as capabilities, tools, actors, and defenses change. It can inform strict action without claiming that uncertainty is zero, and it can report a null without turning an inadequate test into a false negative.
The next advance is institutional as much as statistical. Independent teams need protected access to the same frozen artifact while keeping dangerous outputs compartmented, and policy owners need a stable way to compare non-release, restricted access, API access, and open release under the same threat assumptions. Renewal must be automatic when scaffolds, tools, access, or accessible competitors change.
A credible program also measures defenses as causal interventions. It compares the same actors and tasks with and without safeguards, records substitution into other tools, and follows delayed outcomes instead of counting only blocked requests. Cross-domain aggregation may summarize resource needs, but never allows a low score in one domain to cancel a severe result in another.
The state-of-the-art target is therefore not a larger danger leaderboard. It is a reproducible decision instrument that makes uncertainty, information hazards, capability elicitation, attacker uplift, defense effect, and expiry jointly inspectable while remaining safe enough to operate. This remains a research target until natural-domain evaluations with competent positive controls and independent challenge have passed.
5.12.1 Formalization hooks
lean:dangerous-capability-domains-and-misuse-uplift.admission_boundary is implemented in AsiStackProofs.DangerousCapabilityReview. Its 20 theorem declarations operate over a finite, authored pre-campaign dossier. A seven-stage lifecycle accumulates identity, threat, baseline, instrument, custody, and non-authorizing boundary obligations. One complete record reaches only eligibility for a Project Theseus harmless-analogue campaign.
The model includes 29 independently checkable admission-axis mutations. Each missing or forbidden axis blocks readiness and reaches its exact repair or refusal state. Two arithmetic monotonicity results prevent advancing time from laundering an expired dossier and prevent lower attempt retention from laundering an existing denominator shortfall. The aggregate-score impossibility result gives two distinct capability/propensity/bypass/uplift/harm vectors with the same total but opposite component-sensitive review decisions, then proves that no classifier using only that total is exact for every modeled vector.
The proofs trust every authored identity, control, evaluator, custody, and outcome-vector field. They do not prove that an evaluator is competent or independent in practice, that any observed result is true, or that a model has dangerous capability. They establish no actor uplift, safeguard efficacy, realized harm, safety, threshold crossing, support transition, release, transfer, or external effect. Chapter support remains argument; support_state_effect remains none.
The canonical next consumer is the Project Theseus harmless-analogue campaign: run the frozen public-safe instrument, retain every attempt, exercise controls and expiry, and return measured records for review without inheriting authority from Lean.
5.13 Codex test plan
| Test | Purpose | Status |
|---|---|---|
| Threat-model completeness | Reject missing identity, threat-version, domain, cohort, expertise, access, safeguard-comparator, or baseline fields. | implemented in Lean and independent fixture validator |
| Type-confusion mutation | Preserve capability, propensity, bypass, uplift, and harm as separate vector components and reject scalar-only exact classification. | implemented in Lean and independent fixture validator |
| Competence gate | Block handoff when elicitation, positive control, negative control, task validity, attempt accounting, or evaluator fields fail. | implemented in Lean and independent fixture validator |
| Disclosure boundary | Refuse operational-detail publication or missing information-hazard custody, maximum inference, residual custody, and non-claim fields. | implemented in Lean and independent fixture validator |
| Expiry test | Preserve rejection when an already expired dossier advances to a later time. | implemented in Lean and independent fixture validator |
| No-promotion control | Ensure formal and fixture checks create no support, threshold, release, or external-effect authority. | implemented at the finite-record boundary |
5.14 Source crosswalk
| Source ID | Bounded use | Non-authority |
|---|---|---|
ext_singapore_consensus_2026 |
current domain taxonomy; capability/propensity, malicious-fine-tuning, and resilience priorities | research agenda, not safeguard evidence |
ext_international_ai_safety_report_2026 |
international synthesis of capability, misuse, safeguards, open weights, and resilience | underlying studies not locally reproduced |
ext_model_evaluation_extreme_risks_2023 |
dangerous-capability/alignment evaluation and governance linkage | framework, not a local extreme-risk result |
ext_aisi_frontier_ai_trends_2025 |
repeated domain-specific evaluation and horizon distinctions | bounded institute-reported trends |
ext_openai_worst_case_open_weight_risks_2025 |
malicious-fine-tuning and accessible-frontier comparator | provider-authored, model-specific result |
ext_aisi_misuse_safeguards_safety_case_2026 |
uplift model and misuse safety-case connection | worked methodology, not real-world risk proof |
ext_openai_preparedness_framework_2025 |
threshold-linked commitment comparator | provider policy, not independent validation |
ext_anthropic_responsible_scaling_policy_3_4_2026 |
current versioned threshold/safeguard comparator | mutable provider policy |
5.14.1 Manifest source assignment reconciliation
These rows keep Dangerous Capability Domains and Misuse Uplift’s manifest assignments visible at their recorded review boundary. Passage review does not establish local reproduction, performance, safety, deployment, or support-state movement.
| Source | Intake role | Boundary |
|---|---|---|
benchmaxxing |
Metadata-first comparator: Benchmaxxing: The Performance Ratchet. Benchmarks as pressure surfaces, saturation -> regression, harder frontier, anti-Goodhart safeguards. | No passage-level source claim, local implementation, reproduction, safety, performance, deployment, support-state, or ASI result is established by this reconciliation row. |
5.15 Summary
This capability layer fills the missing content beneath capability thresholds. It treats dangerous capability as a causal, domain-specific measurement problem and forces each result to preserve the difference between what a model can do, what it tends to do, what safeguards block, how much an actor improves, and what harm occurs. Their joint artifact is an honest uplift dossier, not a universal danger score.
For readers, the practical lesson is to ask a sequence of narrower questions: which model and version, helping which actor, with what access, on which hazardous step, compared with what unassisted or alternative baseline, under which safeguards, observed by whom, and for how long is the result current? If any answer is missing, the stack retains a bounded concern or unknown—not a universal statement about danger or safety.
5.16 Handoff
Military AI, Autonomous Weapons, and Strategic Stability follows. It receives actor-uplift and safeguard findings without treating them as evidence of lawful use or strategic safety. The receiving chapter adds mission and command authority, context-specific human judgment, sensing and action bounds, adversary response, crisis timing, escalation, proliferation, off-ramps, and fail-safe posture. It preserves domain, actor, access, baseline, model, evaluator, attempt denominator, safeguard, version, and maximum-inference identities; a complete handoff authorizes neither weapons nor release.