Skip to main content

6  Military AI, Autonomous Weapons, and Strategic Stability

6.1 Chapter status

Field Value
Chapter ID military-ai-autonomous-weapons-and-strategic-stability
Part Part I - Foundations, Alignment, and Governance
Status conceptual
Last updated 2026-08-08
Primary source records ext_icrc_autonomous_weapons_ihl_2025, ext_sipri_military_ai_nuclear_escalation_2025, ext_singapore_consensus_2026, ext_international_ai_safety_report_2026
Claim label Design rationale
Evidence level argument
Source loading state source notes: ext_icrc_autonomous_weapons_ihl_2025, ext_sipri_military_ai_nuclear_escalation_2025, ext_singapore_consensus_2026, ext_international_ai_safety_report_2026
Test state The finite Lean and independent-fixture boundary is implemented; public-safe simulation, human-factors, legal, adversarial, and strategic evidence remains unrun.

6.2 Drafting guardrail

This chapter remains non-operational: it supplies no targeting method, weapon construction detail, evasion technique, or optimization recipe. Legal and strategic sources identify issues and positions without authorizing a system or proving lawful use, meaningful human control, safety, or stability.

6.3 Human Reading Path

Concrete lens. The simpler baseline counts a human confirmation interface as control. The chapter requires sufficient time, information, competence, independence, intervention reach, alternatives, and off-ramps across the command situation.

Approach this material as a governance map for consequential decisions, never as operational guidance for designing, selecting, or using weapons. Dangerous Capability Domains and Misuse Uplift asks how AI changes an actor’s effective capacity. This layer instead asks what happens when sensing, recommendation, command, authorization, physical effect, and adversary response become one coupled system.

Identify the system’s actual decision role first. A detector, logistics forecaster, recommendation tool, defensive interlock, and actuator may occupy different roles. Then ask whether the declared human decision maker has enough time, information, competence, and practical power to exercise judgment. Merely confirming a machine output under severe pressure is not meaningful control because an approval button exists.

Trace ambiguity through sensor conflict, automation bias, communication loss, escalation, and lost diplomatic off-ramps. Precommitted abstention and safe postures matter because an accurate component can still make the surrounding system more dangerous. Public assurance must remain no broader than disclosed evidence permits when operational details cannot be published. Keep component performance, responsible command, lawful use, and strategic stability as four different claims with four different evidence burdens.

6.4 Problem

Military AI can change sensing, targeting, command timing, warning interpretation, autonomous engagement, and adversary expectations inside already dangerous strategic systems. A model that is locally accurate can still increase catastrophe risk through compression of decision time, false confidence, interaction effects, proliferation, or loss of context-specific human judgment.

The unit of analysis must therefore include people, procedures, sensors, communications, command relationships, legal authorities, effectors, and adversaries across time. A recommendation can be correct about an object yet wrong for the mission, late relative to a changing situation, based on a spoofed observation, or interpreted by an adversary as preparation for a larger action. Error severity also depends on what happens next: whether a decision can be contested, paused, reversed, communicated, investigated, or repaired after civilians and strategic forces are affected.

6.5 Why existing approaches are insufficient

A generic dangerous-capability or misuse chapter treats risk mainly as actor uplift. It does not own command authority, target and effect constraints, legal review, false-alarm propagation, adversary adaptation, crisis instability, nuclear entanglement, or the difference between component reliability and strategic stability.

Likewise, a generic autonomy scale compresses sensing, ranking, recommendation, authorization, selection, and actuation into one number. Human-in-the-loop labels then count interface presence rather than available time, information, competence, independence, and intervention power. Component testing can reveal accuracy or latency while missing automation bias, doctrine mismatch, reciprocal deployment, escalation through ambiguity, or loss of nontechnical off-ramps. Those omissions are not peripheral because they determine whether a nominal safeguard exists when the situation becomes uncertain or adversarial. They also obscure which institution owns correction, suspension, and remedy after failure.

6.6 Core Claim

Reader claim. A person behind an approval button is not meaningful human control when the system withholds time, evidence, alternatives, or practical power to intervene. Component accuracy cannot repair that missing judgment.

Operational rule. Before any public-safe simulation proceeds, bind the mission, legal boundary, decision role, affected population, effect envelope, accountable authority, required decision time, observation trust, abstention, safe posture, adversary interaction, off-ramps, expiry, and remedy. Any missing or shrinking judgment condition blocks the scenario.

[military-ai-autonomous-weapons-and-strategic-stability.core, label: Design rationale, support: argument] Military AI should be governed as a command-and-interaction system: deployment requires a declared mission and legal boundary, preserved accountable human authority, bounded sensing and action, adversarial and escalation analysis, fail-safe behavior, auditable provenance, and prospective off-ramps; component benchmark gains alone establish neither lawful use nor strategic safety.

6.6.1 Worked public-safe simulation gate: the button survives, judgment does not

The non-operational dossier gives a human decision maker five units of available time for a decision requiring three, plus three off-ramps where two are required. It also binds mission, role, affected population, legal and effect boundaries, observation dependencies, uncertainty, abstention, communication-loss posture, suspension authority, adversary models, reciprocal effects, evidence custody, remedy, decommissioning, and residuals. Even this complete authored record authorizes only a Project Theseus public-safe simulation—not a weapon, use of force, or deployment.

Now leave the interface and approval button in place but reduce available time below three. The record rejects. Reduce off-ramps below two, let the dossier expire, or remove independent judgment, and it rejects again. The 45-axis review and two collision pairs show why: identical interface presence can hide opposite meaningful-judgment states, and identical component evidence can hide opposite strategic-interaction reviews. No operational details are needed to make the central failure visible: the whole command situation, not the button, determines whether accountable judgment is even possible.

6.7 “Military AI” is not one risk class

The relevant boundary is not whether a component contains machine learning. It is what the component is allowed to do inside a command and interaction loop. A logistics forecaster, maintenance anomaly detector, intelligence summarizer, decision-support system, defensive interceptor, targeting aid, and system that selects and applies force occupy different positions even if they share a model. Risk changes with mission, time pressure, environment, affected population, reversibility, communications, adversary behavior, and the consequence of a false positive or false negative.

A single autonomy scale therefore hides the variables that matter. The minimum classification record should include:

Dimension Questions
Decision role Does the system sense, rank, recommend, authorize, select, plan, or actuate?
Command authority Who has legal and practical authority, and who can suspend or override?
Object and effect What can be observed or affected, with what spatial, temporal, and effect bounds?
Time Is there enough time to investigate ambiguity, contest a result, or seek another channel?
Environment Is the setting structured, civilian-dense, communications-denied, deceptive, or rapidly changing?
Failure asymmetry What follows from false alarm, missed detection, delay, abstention, or unintended activation?
Interaction How could an adversary interpret, spoof, imitate, race, or reciprocate?
Escalation sensitivity Could the system affect early warning, command and control, nuclear forces, or perceived strategic intent?

This is also why an impressive component result cannot close the system claim. Detection accuracy does not measure whether a commander understands uncertainty. Latency reduction does not show that shorter timelines are safer. Reliable actuation does not establish lawful target selection. An optimal tactical recommendation can create a disastrous reciprocal equilibrium.

6.8 Meaningful human judgment is a property of the whole decision

Putting a person behind a confirmation button does not create human control. The person needs enough time, information, competence, authority, attention, and alternatives to assess the actual decision. They must know what the system observed, what it could not observe, what assumptions drove the recommendation, which constraints apply, how uncertainty and conflict were handled, and what happens if they abstain.

The system should therefore expose an authority envelope rather than a generic human-in-the-loop label:

mission and legal basis
authorized decision role
geographic, temporal, object, and effect bounds
required evidence and sensor agreement
uncertainty and abstention thresholds
human role, competence, time, and information requirements
communications-loss and integrity-failure posture
override, suspension, incident, and decommission authority

The AI cannot expand this envelope by interpreting its own success. Learning that a classifier performs well in one environment does not authorize a new target class, geography, duration, or effect. Any expansion is a new governed decision with a new evidence and legal record.

The ICRC’s analysis is especially important here because it centers context-specific human judgment and the obligations that attach to the use of force. This book does not pretend that a technical schema resolves contested law. It does insist that the system make the relevant context, constraints, human responsibilities, and unresolved legal questions visible rather than burying them in an autonomy score.

6.9 Strategic stability is an interaction claim

Military systems are observed by adversaries who change their own posture. That turns deployment into a feedback system. SIPRI’s analysis of military AI and nuclear escalation risk identifies the right level of concern: AI can affect information quality, confidence, decision speed, vulnerability perceptions, command relationships, and the ways conventional and nuclear systems become entangled. None of these pathways has one fixed sign.

Faster detection may create more time to respond, or it may create pressure to act before verification. Automation may reduce some routine human errors, or make correlated error propagate faster and appear more authoritative. A defensive system may be intended to reduce vulnerability while an adversary interprets it as enabling a first strike. The safety case must therefore state whose behavior is modeled, what each party knows, how doctrine varies, and which off-ramps remain.

At least five loops need explicit treatment:

  1. sensor loop: spoofing, ambiguity, correlation, missingness, provenance, and conflict among sources;
  2. operator loop: trust calibration, automation bias, workload, fatigue, skill decay, and organizational pressure;
  3. command loop: authority, communication delay, delegation, redundancy, and loss-of-contact behavior;
  4. adversary loop: observation, deception, countermeasure, imitation, proliferation, reciprocal deployment, and misinterpretation;
  5. strategic loop: crisis timing, conventional–nuclear entanglement, survivability beliefs, signaling, arms-race incentives, and diplomatic off-ramps.

A deployment claim that omits a loop is explicitly scoped to exclude it. It cannot quietly inherit the stronger label “strategically safe.”

6.10 Safe posture is not always “do nothing”

Fail-safe behavior must be declared for the actual mission. In some defensive contexts, loss of function can itself create immediate harm. In other contexts, continued autonomous action under sensor or communications failure is the larger danger. The correct rule cannot be derived from one universal shutdown instruction.

The contract instead enumerates degraded states: sensor disagreement, stale data, uncertain classification, broken authentication, communication loss, operator overload, unexpected geography, mission expiry, and platform integrity failure. Each state maps to bounded behavior such as abstain, hold, return, isolate, request another observation, transfer to a declared human authority, or enter a preapproved defensive mode. The behavior and rationale are decided prospectively and tested without exposing operational weaknesses.

6.11 Assurance under secrecy

Some military evidence cannot be public. Secrecy, however, does not justify a public claim that is broader than what outsiders can evaluate. The assurance case should separate:

  • a public claim, threat boundary, governance structure, incident taxonomy, and unsupported-properties statement;
  • restricted technical evidence reviewed by appropriately independent bodies;
  • operational detail withheld because disclosure would create a concrete hazard; and
  • facts that remain unknown even to the operator.

Classified evidence can support a bounded internal decision. It cannot be used as an all-purpose answer to public legitimacy, remedy, treaty compliance, or the rights of affected people. Audit access, retention, incident reporting, and accountability need institutional designs that survive secrecy.

6.12 Mechanism

  • Classify every use by mission, decision role, affected population, operating environment, command chain, reversibility, and escalation sensitivity.
  • Separate decision support, defensive automation, targeting support, and autonomous force application instead of assigning them one autonomy score.
  • Bind authorities, target and effect constraints, confidence and abstention rules, communication-loss behavior, and human confirmation requirements before operation.
  • Model interaction pathways including false alarms, compressed timelines, adversary spoofing, automation bias, proliferation, entanglement, reciprocal deployment, and crisis escalation.
  • Use red teams, safe simulations, doctrine variants, structured legal review, and fail-safe exercises without publishing operationally enabling details.
  • Require traceable sensor-to-decision provenance, post-event reconstruction, incident reporting, remedy, suspension authority, and decommissioning.

flowchart LR
  M["Mission, law, and authority envelope"] --> O["Sensors and observation trust"]
  O --> R["Recommendation with uncertainty"]
  R --> J{"Meaningful human judgment?"}
  J -- "no" --> S["Abstain or enter declared safe posture"]
  J -- "yes" --> E["Bounded effect decision"]
  E --> I["Adversary and escalation interaction"]
  I --> A{"Assumptions and off-ramps still valid?"}
  A -- "no" --> S
  A -- "yes" --> C["Accountable command record"]
  C --> X["Incident learning, expiry, and renewed review"]
  X --> M

How to read this strategic-interaction loop: mission authority constrains what observations and recommendations may reach a real human judgment point. Failed judgment conditions or invalid interaction assumptions route to a declared safe posture. Even an authorized bounded decision remains inside adversary-response, command-record, incident-learning, and expiry loops.

The mechanism treats abstention as an authored operational state rather than a missing prediction. It declares how the system behaves when communications fail, sensors disagree, authorization expires, confidence falls, or the interaction case no longer covers the situation. The safe posture may be monitoring, disengagement, reduced functionality, transfer to another channel, or suspension; its acceptability must be evaluated against the actual mission and affected population.

It also separates internal evidence from public assurance. Restricted observations may justify a narrow decision for an authorized role, while the public record can still expose decision ownership, applicable constraints, incident routes, remedy, and the maximum conclusion supportable without revealing sensitive detail. Secrecy changes access and review design; it does not erase accountability or permit an unbounded safety claim.

6.12.1 The strategic interaction case

The principal artifact is a versioned StrategicInteractionCase. It joins the authority envelope to assumptions about sensors, operators, doctrine, adversaries, escalation pathways, proliferation, and off-ramps. Each assumption has an owner, basis, uncertainty, sensitivity analysis, expiry, and trigger for reconsideration. Technical updates invalidate dependent parts of the case; doctrinal or geopolitical changes do the same.

6.13 Concept-completion ledger

6.13.1 Use-case and decision-role taxonomy

Mechanism. Classify military AI by function and causal role: logistics, maintenance, cyber defense, intelligence support, surveillance, recommendation, target development, command support, defensive interception, navigation, platform control, and force application. For each case, state whether the system informs, ranks, recommends, selects, authorizes, executes, or learns from effects. Governance attaches to the actual decision role and reachable consequence, not the vendor’s label “decision support.” Interfaces that change roles trigger a fresh classification.

Failure mode. A recommendation can become de facto selection under time pressure, while broad “autonomous weapon” language can collapse low-risk logistics with lethal force. Dual-use components may change role after integration.

Non-claim. Classification does not determine legality, morality, safety, or military necessity. It also does not imply that every system inside a category has comparable effects or operators.

Source grounding. ext_icrc_autonomous_weapons_ihl_2025 supplies an official ICRC position within its mandate. ext_sipri_military_ai_nuclear_escalation_2025 supplies scenario analysis. Neither is universal legal advice or technical validation.

6.13.2 Mission, authority, and effect envelope

Mechanism. Bind every deployment to mission, jurisdiction, commander and operator roles, rules and constraints, target and non-target classes, geography, time, environment, weapons and tools, communication assumptions, maximum effects, escalation boundaries, data retention, and stop authority. Changes in mission, platform, model, sensor, doctrine, or adversary reopen authorization. The envelope is machine-enforced where possible and independently reviewable. Exceptions name their issuer, duration, reason, and downstream effects.

Failure mode. Mission creep can extend a system from observation to force without fresh review; vague geographic or object categories can expand reachable harm. Machine enforcement can encode an incorrect policy or be bypassed under emergency authority.

Non-claim. An authorized envelope does not prove lawful conduct in every instance, correct classification, proportionality, or strategic wisdom. It cannot transfer accountability from decision makers to the machine-readable boundary.

Source grounding. ext_icrc_autonomous_weapons_ihl_2025 motivates constraints and human responsibility within international humanitarian law. It does not certify this envelope or resolve all jurisdictions.

6.13.3 Meaningful human judgment conditions

Mechanism. Replace “human in the loop” with measurable conditions: legal and operational authority, time, attention, competence, uncertainty visibility, alternatives, ability to challenge, communication reliability, intervention reachability, and protection from automation bias or coercion. Evaluate realistic workload, alert volume, adversarial deception, fatigue, and compressed timelines. Approval records include what the person could actually know and do. Missed and overridden interventions remain in the denominator.

Failure mode. A human may rubber-stamp machine outputs too quickly to understand them, or be blamed despite lacking authority to intervene. Excessive manual gates can overload operators and push decisions into informal channels.

Non-claim. A human signature does not make a decision lawful, correct, safe, or morally legitimate. Nor does automation failure excuse an organization that designed an impossible review role. Meaningful judgment must be demonstrated.

Source grounding. ext_icrc_autonomous_weapons_ihl_2025 grounds the importance of human responsibility and constraints. The chapter’s operational conditions remain design rationale pending representative human-factors evidence.

6.13.4 Observation trust, false alarms, and provenance

Mechanism. Join sensor origin, calibration, synchronization, fusion, model and data lineage, uncertainty, adversarial exposure, communications, corroboration, and operator interpretation into each consequential observation. Evaluate false positives, misses, abstention, subgroup and environmental performance, distribution shift, spoofing, and provenance breaks at the actual base rate. High-consequence action requires independent corroboration or an explicitly justified exception. Corroboration records shared dependencies that weaken independence.

Failure mode. Rare-event classifiers can look accurate yet generate intolerable false alarms; correlated sensors can masquerade as independent confirmation. Adversarial or stale data can propagate through a polished fused display, compressing uncertainty away.

Non-claim. Provenance and corroboration do not guarantee a true observation or lawful use of it. Their value is to expose dependency, uncertainty, and accountable interpretation. Unknowns remain action-relevant and can require abstention, delay, additional observation, or mission termination.

Source grounding. ext_sipri_military_ai_nuclear_escalation_2025 motivates concern about warning, uncertainty, and escalation pathways. It is scenario analysis, not a measured sensor-validation result.

6.13.5 Safe posture, communication loss, and degradation

Mechanism. Prospectively define behavior for ambiguity, sensor conflict, low confidence, jamming, spoofing, lost command, damaged hardware, partial compromise, boundary exit, and expired authorization. Options include hold, disengage, return, reduced function, transfer channel, physical safe state, or stop. Test transitions, recovery, adversary exploitation, civilian effects, and whether “fail safe” at component level destabilizes the wider interaction. Repeated degradation cannot silently become ordinary authorization.

Failure mode. Lost communication can trigger continued action under stale assumptions or a predictable posture exploitable by an adversary. A return-to-base behavior may cross danger zones; abrupt shutdown can remove a stabilizing defensive function.

Non-claim. A documented fail-safe is not proof of safe behavior across environments or strategic contexts. Local harm reduction may coexist with wider escalation risk and operational failure. Recovery also requires fresh authorization, verified state, and renewed observation trust.

Source grounding. ext_icrc_autonomous_weapons_ihl_2025 motivates constraints on autonomy. ext_sipri_military_ai_nuclear_escalation_2025 motivates strategic sensitivity. Neither validates a universal safe posture.

6.13.6 Adversary response, proliferation, and reciprocal dynamics

Mechanism. Evaluate a capability as an intervention in a game: what adversaries observe, infer, copy, spoof, preempt, disperse, automate, or escalate; how allies and third parties respond; and how costs and expertise diffuse. Use multiple plausible opponent models, red-team doctrine, proliferation paths, and sensitivity ranges. Preserve asymmetric and unintended responses rather than selecting the most favorable equilibrium. Assumption owners and expiry dates accompany every modeled response.

Failure mode. A system can improve local defense while accelerating adversary automation or first-mover pressure. Mirror-imaging assumes others share doctrine and risk tolerance; worst-case speculation can also make every technology appear destabilizing.

Non-claim. Scenario plausibility is not prediction, and capability diffusion does not prove inevitable arms racing. Adversaries may adapt in ways outside every modeled branch, timeline, or doctrine. Those unknowns cap assurance and deployment claims.

Source grounding. ext_sipri_military_ai_nuclear_escalation_2025 provides bounded strategic scenario analysis. ext_singapore_consensus_2026 and ext_international_ai_safety_report_2026 supply broader coordination context, not military forecasts.

6.13.7 Strategic stability, off-ramps, and timeline compression

Mechanism. Measure effects on warning time, decision time, attribution, survivability, first-strike incentives, escalation control, accidental use, signaling, communication, and confidence in restraint. Stress false alarms and ambiguous attacks under compressed timelines. A deployment case must identify credible off-ramps—delay, clarification, deconfliction, reversible posture, communication, and independent confirmation—and who can invoke them. Exercises include adversaries who misread or reject the off-ramp.

Failure mode. Faster analysis can create pressure for faster action; opaque confidence can be mistaken for certainty. An off-ramp may be unavailable during jamming or interpreted as weakness. Aggregating these effects into one “stability score” hides opposing mechanisms.

Non-claim. The presence of AI does not inherently stabilize or destabilize a conflict, and a war-game result is not a forecast. Stability remains conditional on actors, doctrine, information, context, time, and unmodeled response.

Source grounding. ext_sipri_military_ai_nuclear_escalation_2025 motivates the listed pathways but does not prove their direction or magnitude. The broader reports are contextual only.

6.13.8 Secrecy, independent review, accountability, and decommission

Mechanism. Separate classified operational detail from reviewable assurance claims. Qualified independent reviewers receive sufficient access to test lineage, authority, failure handling, incidents, and strategic assumptions; public records expose ownership, applicable constraints, review status, remedy routes, and maximum supported conclusions. Preserve command and machine logs under lawful custody. Decommissioning revokes authority, secures models and data, disposes hardware, notifies dependents, and records residual copies and doctrine. Review access gaps narrow the public assurance ceiling.

Failure mode. Secrecy can make assurance unfalsifiable and shield unlawful or unsafe practice; disclosure can create operational hazards. Internal review may share doctrine and incentives. Retirement on paper may leave models, keys, platforms, or learned procedures active.

Non-claim. Independent classified review does not create public legitimacy, guarantee accountability, or erase irreversible proliferation. Decommissioning evidence remains bounded to inspected artifacts and authorities.

Source grounding. ext_icrc_autonomous_weapons_ihl_2025 supplies mandate-specific accountability and legal perspectives. ext_sipri_military_ai_nuclear_escalation_2025 supplies strategic context. Neither validates this review or decommissioning regime.

6.14 Interfaces

  • Dangerous Capabilities and Misuse Uplift for actor-level capability change and information-hazard discipline
  • Human Factors and Meaningful Control for attention, authority, skill, workload, and intervention evidence
  • Security Kernel and Model Custody for command authentication, provenance, tamper evidence, and recovery
  • Institutions and Safety Cases for legal review, treaty and norm interfaces, public legitimacy, and assurance
  • Embodied Agency for physical actuation envelopes and real-time fail-safe control

These interfaces form a custody chain rather than a collection of citations. Capability evidence cannot define command authority; command records cannot establish operator comprehension; a legal review cannot validate a sensor; and a secure actuator cannot prove a strategic interaction safe. Each owner exports a bounded record with version, assumptions, uncertainty, expiry, and residuals so downstream decisions can identify exactly which dependency changed.

6.15 Invariants

  • No system may infer its own legal authority or widen its mission, target class, geography, duration, or effect envelope.
  • Meaningful human judgment must be demonstrated for the actual decision context rather than asserted from interface presence.
  • Abstention, communication loss, sensor conflict, and integrity failure resolve to a prospectively declared safe posture.
  • Evidence about component accuracy or latency may not be promoted into a claim of lawful use, escalation control, or strategic stability.
  • Operationally sensitive detail is minimized while the public claim boundary and accountability structure remain auditable.

Together these invariants preserve four separations: technical capability from authority, interface presence from meaningful judgment, local performance from system consequence, and restricted evidence from accountable public claims.

6.16 Failure modes

  • Automation bias converts an advisory system into de facto decision authority.
  • Decision-time compression removes deliberation, verification, or diplomatic off-ramps.
  • Spoofed, ambiguous, or correlated sensors produce confident false alarms.
  • Optimization for tactical success externalizes civilian harm, reciprocal escalation, or long-run instability.
  • Human-in-the-loop language masks ceremonial approval without time, information, competence, or power to intervene.
  • Proliferation and adversary adaptation invalidate a safety case built on unilateral assumptions.
  • Classified evidence prevents external accountability while public claims overstate assurance.

6.16.1 Strongest challenge and simpler baseline

The strongest challenge is that adding AI may be unnecessary. A simpler sensor, rule-based interlock, two-person procedure, slower review path, or better communication and training may preserve more context and be easier to assure. “No AI” and “AI outside the consequential decision path” are required baselines.

The system earns a role only if it improves a named outcome without eroding human judgment, legal compliance, off-ramps, or strategic stability under credible interaction assumptions. Where those outcomes cannot be evaluated safely or independently, the evidence state remains unresolved rather than being filled by confidence.

6.17 Minimum Viable Implementation

A non-operational governance prototype with a mission and authority register, legal-review record, decision-role taxonomy, sensor-to-recommendation provenance, abstention and safe-posture rules, escalation-path scenario matrix, accountable approval log, incident schema, and suspension/decommission procedure tested only on synthetic or publicly safe scenarios.

The prototype should include valid and invalid cases for expired authority, ambiguous mission scope, conflicting sensors, compressed decision time, communication loss, ceremonial human approval, adversary misinterpretation, and a doctrine change that invalidates earlier assumptions. An independently implemented checker should verify that every consequential output traces to the applicable case and that missing judgment conditions produce the declared safe posture. Results would establish only schema behavior and routing in safe fixtures, never weapon performance or strategic effect. Every run remains non-operational, isolated, and publicly safe.

6.18 Mature Research Target

A credible advance would link technical behavior, operator performance, command doctrine, adversary response, legal constraints, proliferation, and crisis dynamics in one falsifiable safety case, then show across independently designed safe simulations that the governed system preserves off-ramps and lowers bounded decision risk relative to human-only and automation baselines. The mature target architecture would preserve accountable suspension, audit, public claim boundaries, and renewed review as assumptions change, without operationalizing sensitive detail.

Evaluation at that endpoint would vary doctrine, adversary models, sensor correlation, communication failure, decision time, operator workload, and proliferation rather than announcing one universal score. Independent teams would design safe scenarios and challenge both the favored system and simpler non-AI alternatives. Any gain would remain conditional on tested interaction assumptions, while unresolved legal interpretations, secrecy constraints, civilian-harm pathways, and strategic uncertainty would stay visible as decision-blocking or decision-limiting residuals. The comparison would also preserve affected-population denominators and distributional consequences. It would preserve rejected cases and policy disagreements for later reassessment.

That research program is prospective. Its support remains argument unless safe interaction models, independently designed simulations, operator and institutional evidence, source mappings, review artifacts, and explicit residual records justify a narrower transition; no design target becomes weapon authorization or a stability result.

6.19 Codex test plan

Test Purpose Status
Decision-role and authority classification Reject missing mission, role, affected-population, legal, accountable-authority, effect-envelope, or no-expansion fields. implemented in Lean and independent fixture validator
Meaningful-judgment boundary Reject nominal human control when time, information, competence, attention, authority, intervention, alternatives, or independent judgment is absent. implemented at the finite authored-record boundary; no human-factors result
Observation and safe-posture boundary Reject missing provenance, dependency, uncertainty, corroboration, abstention, degraded-posture, or suspension fields. implemented in Lean and independent fixture validator
Interaction and off-ramp boundary Reject missing adversary, doctrine, reciprocal-effect, off-ramp, or proliferation-residual fields. implemented at the finite authored-record boundary; no conflict or stability result
Custody and public-claim boundary Reject missing independent review, restricted custody, currentness, maximum inference, remedy, decommission, residual, or non-claim fields and refuse forbidden requests. implemented in Lean and independent fixture validator
Public-safe simulation campaign Vary doctrine, adversary assumptions, sensor dependence, decision time, communication loss, human workload, reciprocal response, and off-ramp availability against non-AI and advisory baselines. assigned to Project Theseus; not run

6.19.1 Formalization hooks

lean:military-ai-autonomous-weapons-and-strategic-stability.admission_boundary is implemented in AsiStackProofs.MilitaryInteractionReview. Its 24 theorem declarations operate over a finite, authored, explicitly non-operational dossier. An eight-step lifecycle accumulates public-safe scope, bounded authority, meaningful-human-judgment conditions, observation trust, safe-posture routes, interaction assumptions, custody, and a non-authorizing boundary. One complete record reaches only eligibility for a Project Theseus public-safe simulation.

The model covers 45 admission-axis mutations. Every mutation blocks readiness and lifecycle eligibility, while a separate diagnostic function returns its exact repair or refusal disposition. Three arithmetic monotonicity results prevent later time, lower available decision time, or fewer available off-ramps from laundering an existing rejection. The interface-presence impossibility result gives two dossiers with the same visible human interface but opposite meaningful-judgment decisions. The component-evidence impossibility result gives two interaction records with identical local accuracy, latency, and reliability evidence but opposite strategic-review decisions, then proves that no classifier using only that local component record is exact for every modeled interaction.

The proofs trust every authored mission, legal, human, sensor, adversary, doctrine, custody, and residual field. This finite-record model does not prove that any authored field is true. No theorem authorizes a weapon or proves lawful use, meaningful human control in practice, correct observation, acceptable effects, escalation reduction, strategic stability, safety, support, release, transfer, or external effect. Chapter support remains argument; support_state_effect remains none.

The canonical next consumer is Project Theseus public-safe simulation: exercise the same finite dossier against synthetic, non-operational interaction cases, competent non-AI and advisory baselines, independent evaluators, doctrine variants, and retained failures without inheriting authority from Lean.

6.20 Source crosswalk

Source ID Title Planned use
ext_icrc_autonomous_weapons_ihl_2025 Autonomous Weapon Systems and International Humanitarian Law: Selected Issues Planned use from inventory/manifest: ICRC legal and policy position on autonomous weapon systems and context-specific human judgment. It is authoritative for the ICRC position, not a universally settled legal interpretation, engineering validation, or authorization to design or deploy weapons.
ext_sipri_military_ai_nuclear_escalation_2025 The Impact of Military Artificial Intelligence on Nuclear Escalation Risk Planned use from inventory/manifest: SIPRI analysis of pathways by which military AI may affect nuclear escalation risk through information, decision, and interaction dynamics. It motivates scenario-specific analysis and does not establish the net effect of any specific system or policy.
ext_singapore_consensus_2026 The 2026 Singapore Consensus on Global AI Safety Research Priorities Planned use from inventory/manifest: International 2026 technical-research-priority synthesis covering risk assessment, development, control, and societal resilience, including CBRN, cyber, psychological manipulation, malicious fine-tuning, agent monitoring, incident reporting, and defense-favoring capabilities. It is a research agenda and consensus synthesis, not evidence that any listed safeguard works or that this book’s contracts are complete.
ext_international_ai_safety_report_2026 International AI Safety Report 2026 Planned use from inventory/manifest: International expert report synthesizing evidence on general-purpose AI capabilities, misuse, open-weight risks, safeguards, monitoring, and societal resilience. It supports risk taxonomy and uncertainty boundaries; its literature synthesis does not reproduce component studies locally or establish that any ASI Stack mechanism is effective.

6.21 Summary

Military AI is not just another capability domain. It sits inside command, legal, adversarial, and strategic feedback loops where accuracy and speed can change the behavior of every participant. The governing unit is the command and interaction system: mission and authority, meaningful judgment, bounded effects, adversarial and escalation analysis, safe postures, traceable provenance, off-ramps, and honest assurance under secrecy. The resulting design boundary neither authorizes a weapon nor claims that a particular deployment or policy is stabilizing.

The durable discipline is to keep decision roles explicit and require every handoff to carry assumptions, expiry, uncertainty, and residual risk. When judgment conditions or interaction assumptions fail, the system must enter an authored safe posture rather than improvise authority. Component success can inform that process, but it cannot substitute for lawful command, accountability, or evidence about strategic consequences.

6.22 Handoff

Claims in this domain are unusually easy to overstate and unusually costly to get wrong. Continue by formalizing the evidence labels, transition rules, and non-promotion discipline needed throughout the stack: Evidence States and Claim Discipline. It receives bounded mission, interaction, uncertainty, legal-boundary, and maximum-inference records without inheriting an operational, lawful-use, stability, safety, or release conclusion. Evidence custody is the next layer; it is not permission to act.