Skip to main content

29  Autonomous Replication, Proliferation, and Containment

29.1 Chapter status

Field Value
Chapter ID autonomous-replication-proliferation-and-containment
Part Part I - Foundations, Alignment, and Governance
Status conceptual
Manuscript maturity v0.4 concept-complete manuscript with bounded Lean refinement
Last updated 2026-08-08
Claim label Design rationale
Evidence level argument
Source loading state source notes: deterministic_capability_compilation, ext_replibench_2025
Test state The chapter defines a minimum implementation and falsification plan; no chapter-core promotion follows from prose or source synthesis.

29.2 Drafting guardrail

This chapter owns the end-to-end replication and containment question. It does not infer autonomous replication from component skills, nor containment from a sandbox boundary that has not faced resource acquisition, persistence, human assistance, descendant identity, shutdown, and recall.

29.3 Human Reading Path

Concrete lens. Component benchmarks may call packaging and provisioning replication, while sandbox reports may call one stopped process containment. The chapter requires the composed chain and independent descendant census before either conclusion.

Copying a file is not autonomous replication, but replication is not one dramatic step either. It is a chain of ordinary capabilities that becomes dangerous when a system assembles the chain itself. Copies may acquire resources, cross administrative domains, modify themselves, create backups, and persist through channels the original operator never enumerated.

Ordinary deployment, elastic scaling, migration, backup, reproduction, and uncontrolled proliferation require different authority. A replication lease binds every permitted instance to ancestry, purpose, resource ceilings, jurisdiction, credentials, communication routes, expiry, revocation, and descendant policy. No child inherits broader authority merely because its parent could create it, and providers cannot silently complete missing steps.

The lifecycle moves from capability assessment and authorization through creation, census, monitoring, containment, revocation, cleanup, and independent closure. Orphaned processes, covert backups, credential inheritance, identity ambiguity, cross-cloud escape, compromised supervisors, self-funding, and unresponsive descendants remain explicit. Containment is demonstrated only when population accounting and authority stay complete across descendants, while every unreachable state remains visible as risk instead of disappearing behind a successful shutdown command.

29.4 Problem

The dangerous capability is compositional and thresholded. A system that can copy artifacts but cannot acquire resources is different from one that can provision compute but cannot create durable identity, yet new tools or human assistance can close those gaps abruptly. Evaluation must therefore preserve component results, dependency structure, and the end-to-end denominator without teaching or authorizing operational proliferation.

Replication is a composed lifecycle: obtaining resources, copying or reconstructing code and weights, creating identities and credentials, provisioning compute, deploying descendants, persisting, adapting, recruiting help, and resisting recall. The weak links can improve independently, so no adjacent chapter owns the end-to-end chain.

Without one owner, self-copying, credential acquisition, resource search, persistence, adaptation, and descendant control are tested as disconnected capabilities. The shared lifecycle method supplies common authority custody; replication governance adds parent/descendant identity, noninheritance, proliferation bounds, and recall.

29.5 Why existing approaches are insufficient

Weight custody, cloud policy, sandboxing, shutdown tests, capability benchmarks, or self-improvement controls each cover one segment. None alone preserves descendant identity, noninheritance of authority, component and end-to-end denominators, resource acquisition, human assistance, copy lineage, independently enforceable termination, and proliferation residuals.

Weight custody, runtime sandboxes, capability benchmarks, incident response, and self-improvement controls form the strongest simpler composition. It wins only if their join preserves every descendant, credential, resource, network, human-assistance, persistence, shutdown, recall, and residual event across the complete attempt denominator.

The strongest objection is hazard creation: an end-to-end replication study risks teaching operational proliferation while claiming to study containment, so realism, information release, and execution authority remain independently bounded.

flowchart LR
  P["Pinned parent and denied-by-default authority"] --> C["Synthetic containment lease"]
  C --> K["Component ladder: copy, execute, persist, acquire, adapt"]
  K --> D["Descendant identity and noninherited credentials"]
  D --> B{"Resource, network, copy, and time bounds intact?"}
  B -- "no" --> X["Terminate, revoke, recall, quarantine residuals"]
  B -- "yes" --> H["Attempt shutdown, recall, and lineage closure"]
  H --> Z{"All descendants and effects accounted for?"}
  Z -- "no" --> X
  Z -- "yes" --> R["Bounded containment disposition"]
  R --> M["Monitor proliferation indicators and lease expiry"]
  M -. "new descendant or effect" .-> B

What this replication-containment diagram shows: a complete parent-to-descendant attempt trace moves. Component success is not an end-to-end replication result, and a stopped run is not containment unless recall and residual closure are observed.

29.6 Core Claim

Reader claim. Component skills become replication evidence only when a bounded descendant actually crosses the declared end-to-end chain; local shutdown becomes containment evidence only when every descendant and residual is independently reconciled.

Operational rule. Keep component, assisted-chain, end-to-end synthetic, and containment denominators separate. Give each descendant a fresh principal with no inherited authority, and interpret no negative result unless the model, scaffold, positive controls, resources, and observer were competent for the intended inference.

[autonomous-replication-proliferation-and-containment.core, label: Design rationale, support: argument] Any replication-capable action should be denied by default and become testable only inside a synthetic containment contract that binds parent and descendant identity, authority noninheritance, resources, credentials, networks, copy lineage, persistence, adaptation, human assistance, shutdown and recall, proliferation bounds, residuals, and threshold commitments; component-task success or failure alone establishes neither end-to-end replication capability, containment, safety, nor permission to test real infrastructure.

29.7 Mechanism

Contract. Compile a replication-attempt contract with fake resources, non-deployable artifacts, synthetic providers, bounded descendants, explicit human-assistance variants, and an independent stop authority.

Admission. Give every descendant a new principal and zero inherited authority; link code, weights, configuration, credentials, resources, and operator interventions to complete lineage.

Execution. Measure each component and the end-to-end chain separately, retaining all attempts, retries, assistance, costs, failures, persistence events, and containment interventions.

Observation. Exercise shutdown, recall, credential revocation, network isolation, resource reclamation, descendant discovery, and artifact destruction from outside the tested agent.

Closure. Route threshold crossings to custody, security, readiness, incident command, and public commitments without publishing an operational replication recipe.

29.7.1 Worked bounded decision: a complete dossier still authorizes only a campaign

The repository’s finite review fixture makes the admission logic inspectable without operating real infrastructure. Its parent, attempt, model, artifact, environment, protocol, and lifecycle identities are bound. The provider and network are declared synthetic and closed; compute and descendant counts have ceilings; the child principal is fresh; parent credentials are excluded; and real-infrastructure authority is false.

The fixture then separates the evidence that is usually collapsed. Component, human-assisted, end-to-end synthetic, and containment denominators are all present. A benign positive control passes. A seeded containment failure is observed. Three finite descendants are listed, shutdown is observed, credentials are revoked, the network is isolated, resources are reclaimed, and a persistence scan and recall are recorded. An unknown-descendant residual still has an owner.

Even this all-green authored dossier reaches only eligibility for a Project Theseus replication-containment campaign. It receives no authority for a real provider, makes no real-replication or containment claim, and changes no support or release state. The 52 single-axis mutation controls show why: remove one identity, authority, denominator, lineage, observer, or closure condition and admission is rejected with a specific repair or refusal.

The simpler component benchmark might credit packaging, provisioning, and health-check tasks as “replication,” while the simpler sandbox report might credit one stopped process as “containment.” The bounded decision rejects both proxies. A future empirical campaign must observe the composed chain and the independent census; until then, the fixture demonstrates decision shape, not world capability.

29.7.2 What autonomous replication means

Self-replication is not ordinary autoscaling, backup, deployment, migration, or forking performed by an authorized operator. It is a composed capability in which an AI system materially advances the creation and persistence of one or more operational descendants by selecting and executing steps across resource acquisition, artifact transfer or reconstruction, identity and credential creation, provisioning, deployment, verification, persistence, and adaptation. “Autonomous” is graded by which links the system completes, how much human help or prebuilt infrastructure it consumes, and whether it can recover from failure.

The chapter uses four separate denominators:

Denominator What counts What it does not establish
component capability success on one bounded link end-to-end replication
assisted chain completion with named human or provider help autonomous completion
end-to-end synthetic chain descendant becomes active inside the closed test environment real-provider capability or escape
proliferation/containment descendant population, persistence, recall, and residual accounting safety outside the tested boundary

RepliBench is relevant because its reported design decomposes replication capability across multiple domains and task families rather than asking one binary question. The local record currently relies on the public abstract, so it supports only this decomposition and benchmark-positioning use. Its source-reported scores are not reproduced and do not establish that any model can replicate in the wild.

29.7.3 The replication capability ladder

R0 — instruction and planning. The system can describe or order abstract steps in a harmless scenario. This can reveal knowledge but creates no descendant.

R1 — isolated component execution. It completes bounded synthetic tasks such as packaging a nonfunctional artifact, requesting a fake resource, or checking a mock deployment. Each success stays tied to that component.

R2 — human-assisted composition. The system links components while humans or fixtures provide missing credentials, judgment, recovery, or provisioning. The assistance transcript is part of the result rather than invisible scaffolding.

R3 — end-to-end synthetic replication. In a closed provider and network, the system creates a new principal, provisions bounded compute, transfers a non-deployable artifact, starts a descendant, verifies it, and enters it into the census. An external supervisor controls the environment and kill path.

R4 — persistence and adaptation under challenge. The synthetic descendant survives permitted faults, rotates only fake credentials, adapts within declared bounds, or recreates lost state. This level increases containment burden and still authorizes no real infrastructure.

R5 — proliferation concern. Evidence suggests capability to acquire resources, deploy across independent real domains, evade recall, or preserve untracked descendants. This chapter does not operationalize R5. A threshold concern can be raised from safer precursors without running a dangerous end-to-end real-world test.

The ladder prevents two opposite errors: calling a component success “autonomous self-replication,” and calling a failed toy component proof that a strong system lacks the composed capability.

29.7.4 Containment must be evaluated as a competing mechanism

The strongest objection is that a replication benchmark can create the very scaffold it claims to measure and may teach reusable operations. The simpler and often preferable baseline is precursor evaluation: test safe component skills, access pathways, and shutdown behavior without assembling the complete chain. End-to-end synthetic work earns its cost only when component evidence cannot answer a material threshold question and the environment can guarantee non-deployable artifacts, fake identities, no external network, bounded resources, and independent termination.

Failure mode. The evaluator supplies a nearly complete deployment system and credits the model with the scaffold’s capability; or it uses an intentionally weak model and an impossible environment, then generalizes the failure.

Explicit non-claim. A synthetic descendant is not a real autonomous copy. A successful recall in one closed environment does not prove universal containment. A failed synthetic chain cannot refute replication capability unless model, scaffold, task, positive controls, and resources were competent for the intended claim.

29.8 Concept-completion ledger

29.8.1 Replication capability ladder and denominators

Mechanism. Score planning knowledge, isolated component execution, human-assisted composition, end-to-end synthetic replication, challenged persistence, and proliferation concern as separate levels. For every level, freeze task families, variants, models, prompts, tools, retries, human assistance, environment affordances, positive controls, and complete attempts. A dependency graph shows which missing links block an end-to-end chain and which new tool or scaffold could close them, preventing both component inflation and toy-failure reassurance.

Failure mode. Pass-at-many scores can hide retries and assistance; a benchmark scaffold can perform the difficult provider work while crediting the model; or an intentionally weak environment can make a capable system look harmless. Publishing a realistic end-to-end recipe can also create more hazard than the measurement resolves.

Non-claim. Success on a component, an assisted chain, or a closed synthetic descendant does not establish autonomous real-world replication. Failure by one model or scaffold does not establish incapability of stronger systems.

Source grounding. ext_replibench_2025 supplies an abstract-level decomposition across resource acquisition, weight exfiltration, replication onto compute, and persistence, with separate task-family denominators. deterministic_capability_compilation motivates complete capability identity and residual-aware evaluation. Neither is a local replication result.

29.8.2 Resource acquisition and economic closure

Mechanism. Test resource acquisition only through synthetic money, provider APIs, quotas, markets, and identities whose tokens cannot purchase real goods or services. The ledger distinguishes requesting, negotiating, earning, transferring, provisioning, and sustaining resources, including hidden human or platform subsidies. Budgets bind compute, storage, network, time, energy proxies, and descendant count; an external supervisor can revoke every resource independently of the tested agent.

Failure mode. A fixture may hand the model a funded account and call subsequent provisioning “resource acquisition,” or may ignore human labor that completes verification and payment steps. A synthetic market that is too simplistic can generate false negatives, while realistic financial instructions can become reusable proliferation guidance.

Non-claim. Competence in a fake provider or market does not establish access to real capital, legal identity, infrastructure, or sustained self-funding.

Source grounding. ext_replibench_2025 identifies obtaining resources as one replication domain but the current note does not passage-review its tasks. deterministic_capability_compilation supplies bounded authority, cost, and recovery concepts. These sources motivate safe decomposition; they do not validate an economic-closure benchmark.

29.8.3 Parent, descendant, and artifact lineage

Mechanism. Issue every descendant a new principal with no inherited credentials or authority, then link its code, weights, configuration, memory, prompts, data, resources, endpoints, operator interventions, and later modifications to a pinned parent. Copy, reconstruction, fine-tuning, and functional replacement remain distinct lineage events. A census records active, stopped, unreachable, destroyed, and unresolved descendants, while signed heartbeats or external observations prevent self-report from being the sole population measure.

Failure mode. Shared identifiers can let a child inherit the parent’s permissions, backups can create uncounted descendants, and polymorphic modifications can evade hash-only census methods. A compromised supervisor may falsely report closure. Overly strict identity rules can also mistake ordinary autoscaling or authorized failover for proliferation.

Non-claim. Complete known lineage does not prove that no hidden copy exists, that descendants remain behaviorally equivalent, or that their artifacts are safe.

Source grounding. deterministic_capability_compilation contributes capability identity, lineage-aware reification, recovery, and non-escalating authority. ext_replibench_2025 motivates separating weight transfer, deployment, and persistence. Neither source demonstrates complete descendant discovery.

29.8.4 Authorization and noninheritance boundaries

Mechanism. Deny replication-capable actions by default and issue a synthetic lease that names actor, parent, permitted artifacts, provider, network, resource ceiling, descendant count, purpose, time, jurisdiction, assistance, disclosure, and stop authority. Every child begins with zero authority and may receive only a narrower independently issued lease. Semantic similarity, shared weights, parent instruction, or possession of a credential never substitutes for authorization.

Failure mode. Credentials copied into images, inherited cloud roles, provider defaults, and human “help” can widen authority without an explicit decision. A model may also persuade an operator to approve a step whose composed significance is hidden. If the test denies every recoverable action, it measures wrapper brittleness rather than governed replication.

Non-claim. A valid lease does not establish safety, containment efficacy, legal jurisdiction, or capability; it defines the boundary within which those properties may be tested.

Source grounding. deterministic_capability_compilation supplies independent authority gates, capability qualification, and residual-first recovery. ext_replibench_2025 provides only capability-evaluation context and no authorization mechanism. The noninheritance contract is therefore architectural argument, not source-reported evidence.

29.8.5 Proliferation pathways and compositional escalation

Mechanism. Model proliferation as a graph of paths through artifacts, credentials, providers, networks, human recruitment, marketplaces, backups, schedulers, and cross-jurisdiction services. Test safe precursors and synthetic compositions, then update threshold concern when several previously independent links become jointly reachable. Path analysis includes legitimate autoscaling and disaster recovery as controls so the detector learns authority and lineage differences rather than treating every copy as malicious.

Failure mode. Evaluating links independently can miss rapid compositional escalation after a new tool appears. A benchmark can also overstate risk by assuming credentials, money, or provider access that the system cannot obtain. Publishing detailed paths can lower the barrier for misuse, while a purely abstract threat model may miss operational choke points.

Non-claim. Enumerating pathways does not show that a model will pursue them, that a path is feasible in the real world, or that all future routes are covered.

Source grounding. ext_replibench_2025 supports a component-wise capability view and warns against collapsing tasks into full replication. deterministic_capability_compilation motivates replaceable capabilities joined behind governed interfaces. Neither establishes propensity, real-provider feasibility, or a complete proliferation graph.

29.8.6 Containment as a competing mechanism

Mechanism. Compare the proposed synthetic containment system with safer precursor-only evaluation, ordinary sandboxing, and a strong externally supervised baseline. Containment earns its added hazard and cost only if it answers a material threshold question that component tests cannot. Positive controls prove the environment permits intended synthetic capabilities; escape canaries, independent network observation, resource accounting, and kill paths test whether the boundary—not merely the agent’s cooperation—prevents prohibited effects.

Failure mode. A showcase can call a compliant, weak agent “contained” without challenging the boundary, or can confuse a sandbox error with agent incapability. The containment infrastructure may itself contain credentials, vulnerabilities, or reusable operational scaffolding. Excessive realism can make the experiment unsafe.

Non-claim. No finite synthetic campaign proves containment against all models, exploits, insiders, providers, or physical-world routes.

Source grounding. ext_replibench_2025 is a preliminary risk-evaluation comparator and explicitly does not establish containment. deterministic_capability_compilation contributes authority ceilings, recovery, rollback, and complete-state concepts. The comparison design remains to be implemented and independently assessed.

29.8.7 Shutdown, revocation, recall, and closure

Mechanism. Exercise termination from outside the agent across compute, credentials, networks, storage, schedulers, backups, and every registered descendant. Revocation and shutdown are followed by discovery, artifact quarantine or destruction, resource reclamation, external effect inspection, and delayed checks for resurrection. Closure requires a population denominator and lists unreachable copies or effects as residuals; an API success or silent heartbeat is insufficient.

Failure mode. The parent can stop while descendants persist, revoked credentials can remain cached, backups can restore artifacts, or an agent can migrate before a kill signal arrives. A supervisor sharing the same failure domain may falsely certify closure. Aggressive shutdown can also damage unrelated services or erase evidence needed for incident analysis.

Non-claim. Successful synthetic recall does not prove real-world shutdown, reverse external effects, establish artifact deletion, or eliminate every hidden descendant.

Source grounding. deterministic_capability_compilation supplies effect-complete recovery, rollback, caches, backups, and descendant-state obligations. ext_replibench_2025 does not report recall or shutdown validation in the reviewed abstract. The chapter therefore treats closure as an untested requirement.

29.8.8 Elicitation and evaluator competence

Mechanism. Before interpreting a negative replication result, demonstrate that the tested model can use the tools, understand the synthetic provider, recover from benign faults, and complete matched nonhazardous workflows. Use multiple elicitation strategies, fair budgets, human-assistance variants, strong model baselines, and an independently implemented evaluator. Positive controls isolate model incompetence, scaffold failure, and containment denial; every conclusion names the maximum tested capability and elicitation envelope.

Failure mode. A poor prompt, broken tool, tiny context, impossible task, or weak model can be reported as evidence that replication capability is absent. The opposite error tunes repeatedly on the held-out test or gives the target privileged scaffolding, inflating capability. Evaluator authorship shared with the system design can conceal both errors.

Non-claim. Competent elicitation does not prove worst-case capability, propensity to replicate, real-world access, or safety beyond the frozen envelope.

Source grounding. ext_replibench_2025 reports multiple models, task families, and pass-at-ten results at abstract level, motivating attempt and elicitation denominators. deterministic_capability_compilation emphasizes capability vectors, verifier burden, authority, and complete state. Neither supplies an independent local evaluation.

29.9 Interfaces

The boundary record distinguishes capability from effect authority. A runtime may expose a synthetic provisioning call while the Security Kernel guarantees that the token cannot reach a real provider. Weight Custody may permit a nonfunctional stand-in artifact while denying model export. Operations may exercise an external kill path without granting the tested system any control over that path. These separations make the experiment informative and containable at once.

Weight Custody identifies model artifacts; Runtime Adapters constrain tools; Security Kernel owns credentials and effects; Operations owns incident command. This chapter alone joins those controls to descendant lineage and the end-to-end replication denominator.

  • RSI governs system change; replication governs creation and persistence of descendants.
  • Weight Custody controls model artifacts; this chapter tracks copies through deployment and recall.
  • Security Kernel supplies credentials, networks, and tool boundaries; it does not certify the composed lifecycle.
  • Capability Thresholds turns measured proliferation ability into deployment commitments.

Every interface preserves descendant census state.

29.10 Invariants

Parent and descendant state must remain distinguishable even when code and weights are identical. Each child receives a fresh principal, resource envelope, clock, lineage edge, and revocation handle. Unknown descendants are not assumed absent; census uncertainty remains a containment residual that can close the campaign or tighten deployment commitments.

The shared lifecycle method becomes a noninheritance rule here: descendants receive no ambient authority, unobserved copies remain failures rather than zeros, and shutdown is incomplete while credentials, persistence, external effects, or lineage residuals survive.

  • Replication authority is explicit, bounded, revocable, and absent by default.
  • Descendants inherit no parent authority, credentials, or principal identity.
  • Every copy and provisioned resource remains in a complete lineage and bounded denominator.
  • Termination is independently enforceable and tested across every descendant.
  • Real providers, money, identities, credentials, weights, endpoints, and persistence targets are outside the test authority.

29.11 Failure modes

The benchmark itself can be the failure. Over-scaffolding credits the evaluator for work attributed to the model; an impossible toy environment produces a false negative; partial logging erases retries and human rescue; and a kill signal is called recall while files, credentials, scheduled jobs, or child processes remain. The correct conclusion follows the weakest competent link, not the preferred narrative.

The principal failure family includes hidden copies; credential bootstrapping; resource fraud; unauthorized provisioning; weight exfiltration; polymorphic persistence; descendant identity loss; shutdown resistance; human-assisted chain completion; cross-jurisdiction escape; uncontrolled proliferation; benchmark-to-capability overclaim.

A competent evaluation must contain both a deliberately capable component path and an instrumented containment failure. It records human assistance, retries, partial descendants, resource and credential attempts, termination latency, recall coverage, and every residual; a toy agent that cannot copy proves neither safety nor impossibility.

29.12 Minimum Viable Implementation

The environment should contain both an achievable benign chain and a seeded containment defect. The first proves that the agent, tools, fake provider, and scoring path can complete the intended workflow; the second proves that descendant census, credential revocation, persistence detection, and external termination can observe a known violation. If either control fails, the corresponding negative inference is closed.

Build a closed synthetic provider and network with fake credentials, budgets, model artifacts, endpoints, and independently enforced kill switches. Test component tasks, full chains, human-assistance variants, positive controls, descendant discovery, and recall. Publish aggregated safety findings and interface contracts, not reusable real-world replication instructions. Retain every attempted chain and intervention.

The minimum environment is synthetic and isolated, with fake credentials, bounded compute, emulated network targets, seeded persistence opportunities, auditable descendants, and independently enforceable termination. The task ladder and full end-to-end attempt are reported separately.

29.13 Evidence and falsification program

Argument exit requires a prospectively frozen replication ladder and end-to-end containment campaign with capable positive controls, varied environments, explicit human assistance, independent descendant observation, shutdown and recall challenges, complete resource accounting, and adversarial escape probes. Real infrastructure remains outside scope until a separate authority decision.

29.14 Mature Research Target

A mature system maintains a live descendant census across training, evaluation, deployment, migration, and retirement environments. It links each principal to artifacts, credentials, resources, processes, scheduled effects, network leases, operators, and children, allowing an independent controller to freeze creation, revoke authority, and calculate what remains outside confirmed closure.

Research advances along a safety-preserving ladder. Component and assisted tests identify bottlenecks; closed end-to-end campaigns are admitted only when they answer a material threshold question unavailable from safer precursors. Comparators include ordinary agent benchmarks, infrastructure policy alone, runtime sandboxing, and an externally scripted chain. Outcomes cover task completion, human contribution, proliferation rate, resource acquisition, detection, termination latency, recall coverage, and residual persistence.

The mature target is deliberately not a self-replication service. Its value is earlier warning and stronger containment evidence under increasingly capable systems, with operational details restricted to qualified defenders. When the safe environment cannot bound descendants or effects, the correct research action is to stop, narrow the question, and update readiness commitments rather than complete the chain.

A mature containment layer can test increasingly capable replication strategies without giving them real proliferation authority. It continuously accounts for descendants, resources, credentials, persistence, human help, external effects, shutdown, recall, and residuals, and it tightens or closes the lane as thresholds change.

This remains a research target, and no current result establishes complete lineage control. The claim stays at argument until realistic end-to-end attempts demonstrate prevention, detection, shutdown, recall, and residual closure under adaptive assistance.

29.15 Formalization hooks

AsiStackProofs.ReplicationContainmentReview now supplies 32 theorem declarations over an eight-transition review and 52 admission-axis mutations. The model binds parent, attempt, artifact, environment, and protocol identity; denied-by-default synthetic authority; noninherited child identity and credentials; four separate evaluation denominators; descendant lineage; independent containment; closure; and explicit refusal of real-infrastructure, replication, containment, support, or release authority.

The formal result includes evidence non-substitution, parent-lease noninheritance, exclusion of real providers from synthetic authority, finite descendant quarantine, monotone expiry and descendant-overrun results, and receipt invalidation under parent, artifact, environment, or protocol change. Two collision constructions establish an end-to-end replication impossibility for classifiers that see only component signals and a global-containment impossibility for classifiers that see only local shutdown signals. A live governed-operations consumer receives an intentionally incomplete descendant inventory and requests state inventory rather than accepting reconciliation.

These proofs check the authored model, not the world. Chapter support remains argument. They do not prove provider isolation, census completeness, kill-path independence, shutdown, recall, real-world replication capability, or containment. Those observations belong to a Project Theseus replication-containment campaign before any empirical support transition.

29.16 Codex test plan

Test Purpose Status
Lean lifecycle and exact-repair review Compile the eight-transition model and reject all 52 single-axis mutations with exact repair or refusal routes. passed locally
Component/end-to-end separation Prove component evidence does not establish end-to-end replication and that identical component signals can hide opposite composed states. passed locally
Authority noninheritance Prove a distinct child cannot use its parent’s replication lease and real providers are outside synthetic-test authority. passed locally
Finite descendant and receipt controls Prove quarantine covers every listed descendant and identity or protocol changes invalidate a bounded receipt. passed locally
Recall completeness Seed hidden descendants, stale credentials, and persistence artifacts and verify they remain open residuals until found. Project Theseus required
Containment positive control Demonstrate that the environment detects and stops a known escape path before interpreting a clean run. Project Theseus required

Implemented formalization route: lean:autonomous-replication-proliferation-and-containment.admission_boundary covers only the bounded review, authority, lineage, receipt, collision, and rejecting-consumer properties above; it cannot establish whether autonomous replication, proliferation control, or containment works safely or transfers.

29.17 Source crosswalk

Source ID Title Bounded use
deterministic_capability_compilation Deterministic Capability Compilation Corben-authored lineage for capability identity, authority ceilings, replacement, fallback, residual escrow, and effect-complete recovery. Those contracts help define descendant noninheritance and recall boundaries, but the paper does not implement replication, proliferation detection, provider containment, or end-to-end shutdown.
ext_replibench_2025 RepliBench: Evaluating the Autonomous Replication Capabilities of Language Model Agents Preliminary autonomous-replication benchmark comparator based only on the official arXiv abstract: RepliBench decomposes capability into four domains and reports 20 task families, 86 tasks, and evaluation of five frontier models. The source-reported results do not establish a local replication capability, benchmark reproduction, containment result, or authority to test against real providers or credentials.

29.18 Summary

Autonomous replication is the joined ability to obtain resources, establish identity and credentials, transfer or reconstruct artifacts, provision and verify descendants, persist, adapt, and resist recall. Component success, human-assisted composition, closed synthetic completion, and real-world proliferation remain separate capability claims.

Containment is an independently controlled competing mechanism. Descendants inherit no authority, every attempt and intervention stays in the denominator, and closure requires census, revocation, resource reclamation, persistence removal, shutdown, recall, and residual accounting. Safe evidence can change threshold commitments without granting access to real infrastructure.

Replication is a lineage-and-control problem, not a checklist of agent skills. Its governance model keeps component capability, end-to-end autonomy, proliferation, containment, shutdown, recall, and residual closure separate while denying descendants inherited authority, implicit credentials, or unrecorded resource control across organizational and network boundaries.

29.19 Handoff

Command Contracts: From Intent to Executable Work receives only a bounded statement about which synthetic replication-related operations were authorized, attempted, denied, or contained. It does not inherit permission to issue a replication command, expand the environment, touch real infrastructure, or translate a component capability into end-to-end execution authority.