flowchart LR
A["Backlog item or chapter need"] --> B["Phase proposal"]
B --> C["Entry criteria"]
C --> D["Required artifacts"]
D --> E["Acceptance gates"]
E --> F{"Gates pass?"}
F -- "yes" --> G["Unlock dependent phase"]
F -- "no" --> H["Block, residualize, or revise"]
G --> I["Release notes and evidence refs"]
H --> J["Residual ledger and next action"]
I --> K["Capability allowed to compose"]
J --> B
85 Prototype Roadmap
85.1 Chapter status
| Field | Value |
|---|---|
| Chapter ID | prototype-roadmap |
| Part | Part IV - Evidence, Implementation, and the Living Book |
| Status | conceptual |
| Manuscript maturity | v0.3 semantically audited manuscript |
| Last updated | 2026-08-08 |
| Primary source records | viea, benchmaxxing, scf, vcm_public, planforge, beastbrain, beastbrain_timeless, bugbrain, moecot, moecot_md, road_to_agi, coherence_exchange, project_theseus_whitepaper, theseus_plan_compiler, theseus_self_evolution_system, theseus_architecture_gate, theseus_operator_os, theseus_circle_transfer, circle_ai_contract_suite |
| Claim label | Design rationale |
| Evidence level | argument |
| Source queue | primary: viea, benchmaxxing, scf, vcm_public, planforge, beastbrain; supporting: beastbrain_timeless, bugbrain, project_theseus_whitepaper, theseus_plan_compiler, theseus_self_evolution_system, theseus_architecture_gate, theseus_operator_os, theseus_circle_transfer, circle_ai_contract_suite; connector/recovery: moecot, moecot_md, road_to_agi, coherence_exchange |
| Source loading state | source notes: viea, deterministic_capability_compilation, platonic_world_model, benchmaxxing, scf, vcm_public, planforge, beastbrain, beastbrain_timeless, bugbrain, moecot, moecot_md, road_to_agi, coherence_exchange, project_theseus_whitepaper, theseus_plan_compiler, theseus_self_evolution_system, theseus_architecture_gate, theseus_operator_os, theseus_circle_transfer, circle_ai_contract_suite, ext_nist_ai_rmf_1_0_2023, ext_model_evaluation_extreme_risks_2023, ext_dafny_2010, ext_copilot_runtime_monitor_2010, ext_shop2_2003, ext_swe_bench_2023, ext_mmlu_2020, ext_checklist_2020, ext_codebleu_2020, ext_qlora_2023, ext_swe_rebench_v2_2026; raw cache: viea, benchmaxxing, scf, vcm_public, planforge, beastbrain, beastbrain_timeless, bugbrain, moecot_md, road_to_agi; connector/recovery: moecot, coherence_exchange |
| Test state | prototype_phase_record.valid.json passes repository-level protocol fixture validation; python3 scripts/validate_prototype_phase_gates.py retains 2 public-safe phase-gate fixtures and 6 expected-invalid controls, recompiles the exact 37-declaration AsiStackProofs.PrototypeRoadmap surface, checks strict dependency order and two complete execution/evidence-review traces across 10/10 splits, explores 33 reachable states through 1,023 transitions, checks eighteen terminal states through 558 absorbing transitions, and rejects nineteen semantic mutations. python3 scripts/validate_readiness_residual_gates.py validates synthetic expired-evidence rerun/reject behavior. Real program execution, evaluator competence, benchmark, deployed build-controller, rollback-effect, and full evidence-state audits remain planned. |
85.2 Drafting guardrail
This roadmap is sequencing logic, not evidence. A phase milestone, source-reported result, or local render success cannot promote a capability claim without the relevant artifacts.
Report-first implementation references can guide the stack without becoming public empirical proof. The roadmap turns that caution into build order. The prototype should earn later agency by first making source ingestion, artifact continuity, claim ledgers, validation, and authority stops boringly reliable.
The roadmap is therefore not a schedule. It is a dependency contract for trust. A phase is allowed to unlock later work only when the artifacts that would make that later work inspectable already exist.
85.3 Human Reading Path
Concrete lens. The simpler baseline treats a finished artifact or green milestone as permission for all later work. The chapter issues consumer-specific unlocks and lets useful research continue without integration or promotion.
Project Theseus gives implementation pressure that needs an accountable build order. The roadmap turns that pressure into build order. It starts with evidence hygiene, source ingestion, artifact continuity, and claim discipline before it asks for stronger agency.
The roadmap is a humility device. A phase is not a result, and a milestone is not evidence. The build should earn later autonomy by first proving that the project can preserve sources, traces, tests, proofs, residuals, and release records.
The order matters because early shortcuts become later architecture. A prototype that cannot keep evidence boringly reliable should not be trusted with stronger routing, execution, or self-improvement authority. Sequence is part of safety.
The roadmap is therefore a build discipline for earning trust one controlled surface at a time. Each phase should make the next claim harder to exaggerate and easier to test.
Progress means reducing the number of things that have to be taken on faith.
A good roadmap converts ambition into smaller verifiable obligations. Milestones matter when each one visibly reduces a specific uncertainty before expansion.
85.4 Problem
A credible build order cannot create agency before auditability, evidence, and authority controls exist. A prototype that begins with autonomous self-improvement skips the machinery that would make improvement inspectable.
The build order must also prevent partial demos from becoming institutional facts. A useful local trace, agent loop, retrieval store, or benchmark dashboard can guide engineering, but it should not become a capability claim until phase acceptance records say what was run, what passed, what failed, what remains blocked, and what the milestone does not imply.
The roadmap problem is therefore sequencing, not ambition. The project should still build toward governed autonomy, tool use, routing, memory, and self-improvement, but each step must inherit the evidence and rollback machinery of the step before it. Otherwise the prototype will demonstrate the exact failure the book warns about: impressive agency without durable proof of governability.
85.5 Why existing approaches are insufficient
A roadmap that jumps to autonomous improvement skips source matrices, artifact graphs, claim ledgers, tests, and authority controls. It also makes useful prototypes dangerous: once something feels helpful, teams are tempted to skip evidence discipline.
VIEA says durable artifacts first. SCF says replacement and improvement need stable fields. VCM says context admission and adequacy must be visible. PlanForge says plans need typed contracts. Project Theseus says report-first local machinery should gate training and self-evolution.
Prototype roadmaps often hide dependency direction. They list features in an attractive order, but do not say which earlier gates make later agency legitimate. This architecture needs the opposite: every attractive capability should point backward to the evidence, authority, replay, and rollback infrastructure it depends on.
External baselines help order the prototype ladder. NIST AI RMF (ext_nist_ai_rmf_1_0_2023) and extreme-risk evaluation work (ext_model_evaluation_extreme_risks_2023) set risk/evaluation gates, Dafny (ext_dafny_2010) and Copilot (ext_copilot_runtime_monitor_2010) point toward specification and monitor artifacts, SHOP2 (ext_shop2_2003) and SWE-bench (ext_swe_bench_2023) pressure planning and software-task evaluation, MMLU (ext_mmlu_2020) and CheckList (ext_checklist_2020) pressure benchmark and behavioral coverage, CodeBLEU (ext_codebleu_2020) pressures artifact-quality metrics, and QLoRA (ext_qlora_2023) names efficient adaptation as a later evidence lane. The roadmap uses them to sequence prototypes, not to claim those systems have been reproduced.
Without that dependency discipline, a roadmap becomes a story about ambition rather than an engineering sequence that can stop, downgrade, or refuse work when its prerequisites are missing.
85.6 Core Claim
Reader claim. A prototype phase is not complete because its artifact is impressive or its checklist is green. It is complete only for the exact consumer and authority unlocked by current dependencies, observed gates, and an explicit terminal decision.
Operational rule. Bind every phase to typed predecessors, required artifacts, commands, environment, resource bounds, evaluator, acceptance gates, phase debt, rollback, retirement, evidence effect, and forbidden claims. A dependency inversion, missing artifact, self-evaluation, or promotion without an evidence transition rejects the unlock.
[prototype-roadmap.core, label: Design rationale, support: argument] Prototype Roadmap owns a program-, roadmap-, phase-, dependency-, artifact-, acceptance-gate-, authority-, evaluator-, evidence-transition-, phase-debt-, residual-, rollback-, reviewer-, consumer-, environment-, and time-specific Evidence-Gated Phase Unlock Contract. It binds every proposed phase to prerequisites, allowed work state, required artifacts, commands, environment, resource bounds, gates, independent evaluation, residuals, debt, rollback, retirement, evidence effect, and non-claims before later work may research, demo, integrate, promote, or release. No roadmap row, milestone, source report, dashboard, task count, passing fixture, theorem, validator, build, or locally useful prototype alone establishes phase completion, safe dependency order, capability, governance effectiveness, deployment, transfer, AGI, ASI, or SOTA.
The claim remains at argument support. The source notes support a staged build sequence, but the prototype phases are not claimed complete.
No roadmap phase has been accepted, executed, or used to promote a capability claim in this repository.
The roadmap should therefore be read less like a product plan and more like a dependency graph for trust. The early phases create the conditions under which later phases can mean anything: sources can be inspected, artifacts can be found again, claims can be downgraded, and authority checks can stop an unsafe action. Later phases are only legitimate when they inherit those conditions. A planner that cannot preserve source provenance is not ready to dispatch work. A runtime that cannot replay actions is not ready to close loops. A replacement system that cannot roll back is not ready to improve itself.
This ordering also protects useful prototypes from their own momentum. A half-working local agent, benchmark harness, context memory, or route scheduler can feel convincing long before it is governable. The phase record slows that down. It says which artifacts exist, which checks passed, which blockers remain, and which claims are still forbidden. That makes the roadmap a pressure surface for better engineering rather than a story that later work was inevitable.
Each phase should have entry criteria and exit criteria. Entry criteria prevent starting a phase when its prerequisites are absent. Exit criteria prevent treating a phase as accepted because it produced an impressive artifact. A phase can produce useful work while still failing exit.
85.6.1 Claim-source mapping status
Appendix C now records twenty-nine passage-reviewed mappings for this claim: nineteen implementation, lineage, connector, and local-project records plus ten external comparators already named in the chapter. The review sharpens what each source may support while keeping every non-claim visible; no external benchmark, formal tool, risk framework, monitor, planner, or adaptation result is imported as local phase evidence.
| Source group | Reviewed support | Boundary |
|---|---|---|
Core stack sequencing sources: viea, benchmaxxing, scf, vcm_public, planforge |
Durable artifacts, command contracts, claim ledgers, benchmark ratchets, stable fields, governed context, typed plans, gates, and residual discipline before higher-agency phases. | No phase acceptance, prototype execution, benchmark result, dependency-gate audit, or claim promotion. |
Prototype lineage sources: beastbrain, beastbrain_timeless, bugbrain |
Local, stateful, resource-constrained, memory/planning/verification/routing vocabulary and cautionary prototype lineage. | No hardware, performance, context-length, security, autonomy, cost, AGI, consciousness, build, emulation, or hardware-test claim. |
Connector/recovery runtime sources: moecot, moecot_md, road_to_agi, coherence_exchange |
Runtime-reference context, release-wording normalization, remaining-work/blocker categories, governance/evidence-interface vocabulary, and source-reported status boundaries. | Source-note-only or variant use; no runtime artifact, reproduced command, benchmark evidence, economic mechanism, or roadmap gate. |
Theseus implementation references: project_theseus_whitepaper, theseus_plan_compiler, theseus_self_evolution_system, theseus_architecture_gate, theseus_operator_os, theseus_circle_transfer |
Report-first implementation lane, plan compiler contracts, self-evolution gates, architecture-gate checks, operator/work-board controls, and structural transfer design. | No current report bundle, compiler command, gate output, dashboard run, board step, transfer consumer, benchmark result, model-quality result, or deployment readiness. |
Circle contracts and formal/fixture artifacts: circle_ai_contract_suite, prototype_phase_record.valid.json, AsiStackProofs.PrototypeRoadmap |
Future proof-carrying contract milestone shape and finite record-level phase-gate predicates. | The separate Circle rope receipt slice covers one bounded external receipt replay only; no vendored contract pack, ASI Stack acceptance check, downstream consumer, transfer evidence, or semantic proof of phase completion exists. |
85.6.2 Publication placement and preserved technical ownership
This chapter remains the phase-governance technical route beneath Project Theseus as Report-First Implementation Reference. It owns the Evidence-Gated Phase Unlock Contract: strict dependencies, per-phase work and acceptance records, evaluator custody, research/demo versus integration and evidence-review states, phase debt, rollback, residuals, retirement, and terminal decisions. The parent chapter owns imported and replayed Theseus artifacts, commands, environments, report families, currentness, public-safety boundaries, and maximum inferences.
This placement follows a failed semantic-merge test, not a weakened technical claim. Combining the files would retain two independent predicates and two proof programs inside one longer skeleton rather than remove repetition. The publication nest keeps the family navigable while preserving this chapter’s claim identity, twenty-nine source mappings, phase-record schema and fixtures, 37 Lean declarations, independent state-space consumer, test plan, debts, residuals, implementation horizon, support ceiling, and historical URL.
The boundary runs in both directions. A roadmap transaction does not inherit report truth, replay success, runtime currentness, benchmark validity, or capability from the Theseus reference. A complete report packet does not inherit safe sequencing, dependency satisfaction, phase acceptance, evaluator independence, rollback success, or program completion from this chapter.
85.7 Mechanism
85.7.1 Worked phase gate: useful research is not an integrated milestone
The local gate suite presents two valid records. An infrastructure phase with its required dependency, artifact, evaluator, gate result, rollback, and non-claim boundary reaches integrate. A phase carrying unresolved debt and a retirement condition remains research_only; its useful artifact does not unlock later capability. Six controls then make the boundary adversarial: dependency inversion, missing artifact, self-improvement without an evaluator, promotion without an evidence transition, phase debt without retirement, and a missing non-claim boundary all reject.
The independent model goes beyond the two happy records. It explores 33 reachable states through 1,023 transitions, checks 18 terminal states through 558 absorbing transitions, and shows that dependency count plus receipt count cannot exactly classify integration when the underlying dependency meaning is lost. These are finite authored phase semantics, not proof that any roadmap milestone was executed or completed. They turn the chapter from a feature list into a stop-capable build order: research may continue while integration, promotion, and release remain closed.
The owned object is an Evidence-Gated Phase Unlock Contract, not a feature list or predicted calendar. Its lifecycle is:
| Phase | Contract operations |
|---|---|
| 1. Freeze and identify | Freeze program objective, roadmap version, candidate phases, dependency graph, authority ceiling, consumers, environments, resources, time, and non-claims; assign identities to every prerequisite, artifact, command, gate, evaluator, transition, residual, debt item, rollback, retirement condition, reviewer, and descendant unlock. |
| 2. Classify and order | Keep proposed, blocked, research-only, demo, integrate, promote, release, failed, rolled-back, retired, and superseded states distinct; build evidence infrastructure first; add intent, planning, context, work, runtime, observation, and rollback second; add model-computation and update layers only under stronger gates. |
| 3. Bind dependencies | Express every edge as typed input, allowed work state, blocking condition, acceptance predicate, output, authority delta, evidence effect, rollback, and expiry; issue a work contract with owners, commands, environment, resources, observations, stops, review, completion, and forbidden claims. |
| 4. Validate and bound | Apply semantic, executable, negative, mutation, replay, proof, rights, publication, and release checks; disclose evaluator implementation and dependence; conserve phase debt; allow failed-gate research or demos only under isolation and non-promotion boundaries. |
| 5. Decide and propagate | Accept completion only from observed artifacts and gates plus an authorized terminal decision; propagate failure, revocation, correction, expiry, rollback, downgrade, and retired debt through every dependent phase and consumer. |
| 6. Challenge and retire | Compare against matched calendar, backlog, milestone, stage-gate, no-gate, and ordinary program-management baselines; monitor useful throughput, false unlocks and blocks, rework, incidents, debt, rollback, burden, resources, opportunity cost, and total cost; expire or retire changed contracts with lineage. |
The contract turns a phase boundary into a reviewable transaction. A proposal first names the capability it hopes to unlock and the authority it must not inherit. Dependency binding then names the exact predecessor artifacts and the conditions under which they cease to be current. Validation records both positive and rejecting observations, including evaluator dependence and work that remains isolated. The terminal decision can accept, block, narrow, roll back, retire, or supersede the phase; none of those outcomes rewrites the work or uncertainty that produced it.
Unlock propagation is consumer-specific. A passed infrastructure gate may let one experiment consume a schema while leaving training, deployment, and public claims blocked. A failed downstream gate can invalidate an affected route without erasing unrelated predecessor work. Every propagation carries roadmap, phase, and consumer identities, the authority delta, evidence effect, residual owner, expiry trigger, and acknowledgement. A green milestone therefore never becomes permission for every later phase.
The natural Theseus flagship is the concrete late-phase consumer of this roadmap contract. ASI-THESEUS-FLAGSHIP-01 may be preregistered only after its five matched route families, competence dossiers, evaluator independence, attack families, resource denominators, rescue rules, freeze boundary, and terminal result classes resolve to versioned artifacts. Completing this roadmap chapter does not complete that experiment, and imported Theseus reports do not satisfy its natural-task gates. The handoff is one-way until the protected campaign is opened: the roadmap may authorize a frozen run contract, while only the flagship’s later evidence and an accepted evidence transition may change a bounded claim.
Its first dependency is now stated at the right granularity. Historical T0 is complete and preserved as an immutable 57M architecture control. T0A is historically complete at its GREEN pre-activation freeze, and T1 is active at step 9,048 under an exact prospective anchor. Because the exact step-3,480 payload and complete predecessor chain to the anchor were not retained, no present-tense full-chain replay may be claimed. Every later segment must join the append-only ledger before another launch. The next unlock is an honest source-disjoint behavioral numerator at T2, not another architecture intake and not a capability inference from optimizer progress.
What the prototype roadmap shows: The roadmap is a dependency gate, not a promise of future success. A phase unlocks dependent work only when entry criteria, required artifacts, acceptance gates, evidence refs, and residual handling justify composition.
The recommended dependency order is:
- Source matrix, source notes, claim ledger, and publication validation.
- Artifact graph, protocol schemas, fixtures, and proof manifest.
- Intent contracts, PlanForge DAGs, VCM context records, and typed jobs.
- Runtime adapters, permissions, audit logs, replay, and procedural memory.
- Routing heads, specialist registry, readiness gates, and benchmark ratchets.
- Stable capability fields, replacement transactions, rollback, and bounded self-improvement.
Two source-derived programs now refine the late phases without bypassing that order. The Deterministic Capability Compilation lane proceeds charter -> scaffold -> semantic capability graph -> coverage -> experts -> NCO linking -> translation validation -> shielded adaptation -> tribunal -> reification. The Platonic World Model lane proceeds stable identity/version -> claim ledger -> proof and semantic governance -> grounded branches -> packet compilation and federation. Each stage must earn its own evidence and can stop honestly; neither paper authorizes a full-stack implementation jump.
Self-improvement comes late because it depends on evaluator integrity, containment, rollback, governance rights, evidence ledgers, resource controls, and residual conservation. Research can proceed earlier in isolation, but it cannot silently inherit integration, promotion, release, training, or deployment authority. Every shortcut becomes phase debt with an owner, severity, downstream exposure, due condition, and retirement rule.
85.8 Interfaces
The contract joins twelve interfaces without absorbing their authority:
- Human Intent, value conflict, agency rights, and System Authority provide legitimate objectives, ceilings, appeal, stop, and terminal decisions.
- Sources, Claim Ledgers, Evidence States, Artifact Graphs, supply chain, rights, and Living Book surfaces provide provenance and public truth.
- Intent contracts, planning, semantic IR, context, Labor OS, and stewardship provide typed work, dependencies, ownership, continuity, and acceptance packets.
- Runtime adapters, Security Kernel, SCIFs, observation, audit, replay, and incidents provide bounded effects, containment, receipts, and recovery.
- Routing, specialists, deliberation, procedural memory, search, and compact models provide candidate computation without unlock authority.
- Benchmarks, behavioral tests, simulation, adversarial evaluation, oversight, safety cases, and formal methods provide evaluation lanes.
- Policy, data, unlearning, replacement, recursive improvement, and open-ended improvement consume accepted lower-phase artifacts and create new causal obligations.
- Readiness, Stable Capability Fields, residual escrow, quarantine, release, deployment, and publication provide qualification and terminal transition authority.
- Resource Economics, compute governance, weight custody, hardware roots, Hives, and services provide budgets, substrates, custody, availability, and costs.
- VIEA, Theseus, Circle, VCM, PlanForge, Benchmaxxing, SCF, BeastBrain, BugBrain, and MoECOT supply lineage without local result inheritance.
- NIST AI RMF, extreme-risk evaluation, Dafny, Copilot, SHOP2, SWE-bench, MMLU, CheckList, CodeBLEU, and QLoRA supply external comparators without certification.
- Status pages, changelogs, dashboards, releases, readers, builders, reviewers, and successors consume projections and cannot overwrite canonical phase records.
Interface exchanges are typed handoffs rather than informal adjacency. Each provider may refuse, expire, or narrow its record; each consumer must acknowledge the exact version it used. Missing acknowledgement remains an open dependency, and a later correction invalidates only descendants that consumed the affected record. This keeps phase sequencing auditable without turning the roadmap into a universal source of truth.
85.9 Invariants
The contract requires eighteen invariants: exact scope; bounded or acyclic dependency direction; evidence and authority before agency; distinct phase states; artifact- and decision-backed unlocks; no milestone-as-result; no research-to-release inheritance; monotone non-widening authority; visible evaluator dependence; conserved phase debt; complete failure denominators; accepted evidence transitions only; effect-complete rollback; correction propagation; joint usefulness, failure, burden, and cost; exact transfer scope; expiry on material change; and a strict finite-proof boundary. Provisional evidence can authorize isolated research under explicit controls, but it cannot unlock default execution, promotion, training, self-improvement, release, deployment, or public capability claims. Canonical phase state also retains prerequisite lineage after completion, failure, rollback, retirement, or supersession.
85.10 Failure modes
The named falsifiers are agency-first sequencing; dependency inversion; roadmap-as-evidence laundering; prototype theater; milestone compression; weak or circular gate theater; evaluator laundering; phase-state collapse; evidence-transition laundering; phase-debt erasure; happy-path-only denominators; source-result inheritance; rollback theater; operator-convenience drift; stale roadmap assumptions; over-gating and false blocking; self-improvement ratchet failure; and unsupported transfer. Each failure must leave a rejected transition, incident, residual, debt item, downgrade, rollback, or retired phase. A polished schedule or demo that cannot expose those records has failed the prototype-roadmap core contract even if it ships useful work. False blocks and gate overhead count as failures when their cost exceeds any measured evidential, safety, recovery, or governance benefit.
85.11 Minimum Viable Implementation
The implemented minimum contains twenty-nine exact source mappings and one prototype phase-record schema with a public-safe fixture. python3 scripts/validate_prototype_phase_gates.py checks two valid fixtures—one accepted non-promoting infrastructure phase and one research-only phase with owned debt and a retirement condition—plus six expected-invalid controls for missing artifacts, dependency inversion, self-improvement without an independent evaluator, support promotion without an evidence transition, debt without retirement, and missing non-claims.
Three proof targets are implemented by 37 Lean declarations. Nine retained declarations cover the original claim-promotion, phase-route, and fixture-bridge guards. Twenty-eight declarations add strict finite dependency ordering and a dependency-aware phase transaction that separates ordinary integration from evidence review. The model proves rejected-event noninterference, exact phase, plan, dependency, artifact, and authority custody, arbitrary-run gate coherence, exact trace composition, reachable integration and evidence-review witnesses, zero support and external-effect authority, terminal integrated/evidence-review/rolled-back suffixes, seven named gate countermodels, and an impossibility result showing that dependency and receipt counts cannot determine integration. A separate synthetic readiness/residual check covers expired-evidence rerun/reject behavior. These artifacts establish authored finite dependency, transaction, and routing boundaries only.
No real program phase is accepted or complete, no deployed build controller exists, no natural program workload or matched baseline has run, and no chapter-core support effect, capability, governance-effectiveness, safety, deployment, independent reproduction, or transfer is established. A blocked phase is useful only insofar as it names the exact artifact, owner, or decision that could unblock or retire it.
85.12 Mature Research Target
The mature argument-exit campaign must prospectively freeze a real multi-phase program: graph, teams, tasks, tools, models, resources, gates, evaluators, debt policy, rollback, stopping, and outcomes. It must compare the Evidence-Gated Phase Unlock Contract with matched calendar, feature-backlog, milestone, stage-gate, no-gate, and strongest ordinary program-management baselines under equal work opportunities.
Natural builds must be mixed with dependency inversion, stale gates, evaluator dependence, research-to-release leakage, debt erasure, revocation, incident, correction, and effect-complete rollback attacks. The campaign must jointly measure useful throughput, time-to-evidence, quality, false unlocks and blocks, rework, incidents, recovery, residuals, debt, reviewer burden, compute, storage, opportunity cost, and total cost. Causal ablations must identify which gate families help, harm, or merely add theater.
Independent teams, reviewers, institutions, infrastructures, legal regimes, environments, and successor programs must reproduce and transfer positive, negative, null, narrowed, blocked, and refuted results. The terminal record must preserve abandoned work, opportunity costs, weak or harmful gates, unresolved debt, outside-model effects, and every condition under which the dependency policy should be narrowed or rejected. Until then the contract is a target architecture, not a functioning build controller, proof of correct sequencing, capability, governance effectiveness, safety, readiness, deployment, transfer, AGI, ASI, or SOTA.
85.13 Codex test plan
| Test | Purpose | Status |
|---|---|---|
| Prototype phase record fixture validation | Validate that a phase record names unlocks, required artifacts, acceptance gates, blockers, validation commands, evidence refs, non-claims, and status. | implemented; passing via python3 scripts/validate_protocol_examples.py |
| Phase acceptance checklist | Check that dependent phases have gates, artifacts, evaluator and rollback records, residual handling, debt retirement, and non-claim boundaries before integration or evidence review. | implemented; passing via python3 scripts/validate_prototype_phase_gates.py with result prototype_phase_gates_2026_07_02_local, 2 valid and 6 expected-invalid fixtures, 37 exact Lean declarations, 10/10 trace splits, 33 reachable states through 1,023 transitions, 18 terminal states through 558 absorbing transitions, and 19 semantic mutations; no real phase completion and no support-state promotion claim |
| Dependency gate review | Check strict finite dependency order and independently reconstruct custody, gate coherence, integration/evidence-review separation, composition, terminal closure, summary insufficiency, and named invalid transitions. | implemented; passing via python3 scripts/validate_prototype_phase_gates.py; no dependency-truth, evaluator-competence, rollback-execution, deployed-build-controller, or support-state-promotion claim |
| Prototype evidence-state audit | Check that roadmap milestones do not promote claims by themselves. | implemented by readiness/residual gate harness for synthetic expired-evidence rerun/reject behavior; full phase evidence-state audit not run |
85.13.1 Formalization hooks
| Tag | Module | Target | Status |
|---|---|---|---|
lean:roadmap.phases.operational_invariant |
AsiStackProofs.PrototypeRoadmap |
Strict finite dependency order and a dependency-aware execution lifecycle preserve exact phase, plan, dependency, artifact, authority, and gate custody across arbitrary runs; only a complete non-promoting trace reaches integration. | implemented |
lean:roadmap.phases.failure_blocks_promotion |
AsiStackProofs.PrototypeRoadmap |
Missing dependency completion, rollback, evaluator, acceptance, debt-retirement, evidence-transition, or non-claim gates reject without state change; integration, evidence review, and rollback are terminal under arbitrary finite suffixes. | implemented |
lean:roadmap.phases.fixture_gate_bridge |
AsiStackProofs.PrototypeRoadmap |
The fixture bridge and independent consumer retain the public-safe route suite, reconstruct both transaction branches, and demonstrate that dependency and receipt counts alone cannot classify integration. | implemented |
These Lean hooks are finite authored semantics over claim promotion, phase routes, dependency order, fixture bridges, and a proposed-to-integrated-or-evidence-review execution transaction. The original route envelope remains, while the lifecycle proves exact accepted advancement, rejected-event noninterference, arbitrary-run custody and gate coherence, composition, terminal closure, and lossy-summary insufficiency. The independent consumer recompiles the exact surface, checks both four-event traces at every split, explores the reachable finite state space, and injects nineteen semantic mutations. It does not prove that the dependency graph is true or complete, that any prototype phase is complete or safe to run, that artifacts or evaluators are competent, that rollback executes in the world, that a gate improves outcomes, or that any capability claim has earned stronger support.
85.13.2 Formal-proof audit boundary
The module contains 37 declarations. Nine retained declarations remain route or fixture consequences; twenty-eight declarations cover strict dependency order, reachable execution and evidence-review branches, arbitrary-run identity and authority custody, gate coherence, exact composition, terminal closure, rejecting countermodels, and a thin-summary impossibility result. The independent consumer checks ten trace splits, 33 reachable states through 1,023 transitions, eighteen terminal states through 558 absorbing transitions, and nineteen semantic mutations. None proves that the chosen dependency graph is correct or complete, that a gate or evaluator is adequate, that any real phase completed, that useful work improved, or that rollback, governance, reproduction, or transfer succeeds.
85.14 Source crosswalk
| Source ID | Title | Layer | Planned use | Readiness |
|---|---|---|---|---|
viea |
Verified Intent-to-Execution Architecture | whole_stack_execution_spine | Intent, artifacts, routing, verification, runtime, feedback, and residuals. | source note available; local raw cache available |
benchmaxxing |
Benchmaxxing: The Performance Ratchet | benchmarks_evidence | Benchmark frontier, regression, and anti-Goodhart discipline. | source note available; local raw cache available |
scf |
Stable Capability Fields | governance_recursive_self_improvement | Stable replacement and rollback substrate. | source note available; local raw cache available |
vcm_public |
Virtual_Context_Memory_v1 | memory_context | Governed context compiler and adequacy/admission separation. | source note available; local raw cache available |
planforge |
PlanForge | planning_control | Goal-to-execution compilation, scheduling, and intelligence arbitrage. | source note available; local raw cache available |
beastbrain |
BeastBrain Cognitive Architecture | whole_stack_lineage | Local, stateful stack vocabulary and implementation-roadmap lineage; source-reported hardware/performance/cost claims not promoted. | source note available; local raw cache available |
beastbrain_timeless |
BeastBrain Timeless Edition | whole_stack_lineage | Evergreen local-memory, planning, verification, and hardware-adaptation framing; no milestone validation. | source note available; local raw cache available |
bugbrain |
BugBrain | edge_efficiency_lineage | Resource-constrained prototype lineage; build/emulation/hardware/AGI/consciousness claims not promoted. | source note available; local raw cache available |
moecot |
MoECOT-Agent Architecture Whitepaper | implementation_reference | Runtime/reference implementation context, readiness gates, ledgers, replay, and blockers; not reproduced here. | source note available; connector or recovery required |
moecot_md |
MoECOT Markdown Export | implementation_reference_variant | Terminology and release-wording normalization; not independent corroboration. | source note available; local raw cache available |
road_to_agi |
Road To AGI | strategic_roadmap | Remaining-work, blocker, and source-reported status context; commands/results not reproduced. | source note available; local raw cache available |
coherence_exchange |
Coherence Exchange | governance_evidence_interface | Optional governance/evidence-interface vocabulary and contestability framing; speculative economics not promoted. | source note available; connector or recovery required |
project_theseus_whitepaper |
Project Theseus Whitepaper | report_first_rmi_prototype | Report-first local RMI implementation reference and safety/report boundary. | source note available |
theseus_plan_compiler |
Theseus Plan Compiler | planning_control | Goal contracts, semantic IR DAGs, VCM slices, routes, contract hashes, and replay-trace targets. | source note available |
theseus_self_evolution_system |
Theseus Self-Evolution System | recursive_self_improvement_governance | Evidence-first self-evolution lane, intervention ladder, ATTD gates, and outcome ledgers. | source note available |
theseus_architecture_gate |
Theseus Architecture Gate | readiness_gate_governance | Gate checks before heavy training; source-reported snapshot not rerun here. | source note available |
theseus_operator_os |
Hive Operator OS and Work Board | labor_os_operator_surface | Durable work board, operator controls, node registry, TTLs, kill switches, and feedback routing. | source note available |
theseus_circle_transfer |
Circle Calculus Transfer Lane | proof_carrying_ai_contracts | Structural transfer design for future private experiments; quality/runtime/model-improvement claims not promoted. | source note available |
circle_ai_contract_suite |
Circle Calculus AI Contract Suite | proof_carrying_ai_contracts | Contract packs, theorem-linked fields, receipt boundaries, and non-claim discipline. | source note available |
ext_nist_ai_rmf_1_0_2023 |
NIST AI Risk Management Framework 1.0 | risk_governance | External comparator for govern-map-measure-manage gates before broader authority. | source note available |
ext_model_evaluation_extreme_risks_2023 |
Model Evaluation for Extreme Risks | dangerous_capability_evaluation | External comparator for dangerous-capability and alignment gates before high-risk phases. | source note available |
ext_dafny_2010 |
Dafny | executable_specification | External comparator for specification and verification artifacts before higher-agency implementation. | source note available |
ext_copilot_runtime_monitor_2010 |
Copilot Runtime Monitor | runtime_monitoring | External comparator for resource-bounded monitor artifacts before autonomous execution. | source note available |
ext_shop2_2003 |
SHOP2 | planning | External comparator for explicit hierarchical decomposition and plan artifacts before dispatch. | source note available |
ext_swe_bench_2023 |
SWE-bench | software_task_evaluation | External comparator for repository-level tasks and executable-test evaluation. | source note available |
ext_mmlu_2020 |
MMLU | broad_benchmarking | External comparator for broad multitask coverage without general-capability inheritance. | source note available |
ext_checklist_2020 |
CheckList | behavioral_testing | External comparator for capability matrices, test templates, perturbations, and failure discovery. | source note available |
ext_codebleu_2020 |
CodeBLEU | code_quality_metrics | External comparator for syntax- and dataflow-aware code metrics subordinate to executable outcomes. | source note available |
ext_qlora_2023 |
QLoRA | efficient_adaptation | External comparator for later parameter-efficient adaptation under data, evaluation, rollback, and resource gates. | source note available |
85.14.1 Manifest source assignment reconciliation
These rows keep Prototype Roadmap’s manifest assignments visible at their recorded review boundary. Passage review does not establish local reproduction, performance, safety, deployment, or support-state movement.
| Source | Intake role | Boundary |
|---|---|---|
deterministic_capability_compilation |
Passage-reviewed Corben architecture source: Deterministic Capability Compilation: A Capability-Preserving Ladder from Executable Scaffolds to Governed Adaptive Agents. Corben-authored July 2026 architecture and research program for compiling executable scaffolds into contract-bound experts and linked Neural Capability Objects while retaining semantic obligation mass balance, candidate-specific translation validation, fallback, residual escrow, authority ceilings, reification, and effect-complete recovery. Existing chapters are upgraded first; no foundry implementation, learned-capability result, preservation result, safety result, SOTA result, AGI, ASI, or support-state promotion is inferred. | No local implementation, reproduction, performance, safety, deployment, support-state, or ASI result is established by this reconciliation row. |
platonic_world_model |
Metadata-first comparator: The Platonic World Model: A Semantic Constitution for Grounded, Proof-Carrying, Self-Editing Artificial Intelligence. Corben-authored July 2026 conceptual architecture and falsifiable research program for semantic continuity through stable Form lineages, immutable semantic versions, typed Essence Contracts, six mutually constraining planes, explicit proposition-attestation-commitment-proof separation, branch-protected world dynamics, qualified grounding, semantic transactions, runtime packet compilation, and federated mappings. Existing chapters are upgraded first; no implemented substrate, benchmark result, philosophical solution to grounding, safety result, SOTA result, AGI, ASI, or support-state promotion is inferred. | No passage-level source claim, local implementation, reproduction, safety, performance, deployment, support-state, or ASI result is established by this reconciliation row. |
ext_swe_rebench_v2_2026 |
Passage-reviewed comparator: SWE-rebench V2: Language-Agnostic SWE Task Collection at Scale. Replaces authored repository fixtures in the empirical lane with post-snapshot public merged changes while requiring setup repair, gold execution, test-path collision guards, independent task review, and final-heldout custody. | The 12 selected development tasks are debugging inputs, not the final denominator or evidence for a chapter claim. No local implementation, reproduction, performance, safety, deployment, support-state, or ASI result is established by this reconciliation row. |
85.15 Summary
The prototype should grow as evidence infrastructure grows: source matrix, claim ledger, artifact graph, schemas, plans, context, typed jobs, runtime adapters, routing, readiness, benchmarks, SCFs, and only then bounded self-improvement.
That order is intentionally conservative. It asks the project to earn agency by first proving that it can remember sources, preserve artifacts, admit failures, replay checks, and refuse unsupported promotions. The same rule applies to the manuscript itself.
Every later capability should be able to point to the earlier records that made it safe to exist. If the pointer is missing, the roadmap should block, narrow, or mark the capability as research-only. The route to v1.0 follows the same logic: strengthen the public book, keep the human and AI views aligned, validate the reader edition pipeline, preserve source ownership boundaries, and avoid support-state inflation. The route to a prototype follows after that with small, reproducible artifacts. A living architecture becomes credible when each new layer inherits the evidence and refusal machinery of the layers before it, not when the project jumps straight to impressive autonomy.
85.16 Evidence reconciliation (2026-07-16)
The invariant protocol, field meanings, and inference limits are stated once in Living Book Methodology. This packet contains only the chapter-specific projection; its authoritative per-atom rows are the prototype-roadmap slice of experiments/claim_family_terminal_coverage/results/result.json.
The core remains blocked after full attempt at argument support. The strongest family attempt was Integrated governed lifecycle slices. Its exact boundary is: Bounded local replay only; no deployment, whole-book proof, external effect authority, transfer, publication, or release claim. Across 74 atoms, the terminal ledger records 74 blocked_after_full_attempt.
| Chapter-specific field | Value |
|---|---|
| Family / atom denominator | CF-08 / 74 atoms |
| Terminal dispositions | 74 blocked_after_full_attempt |
| Core | prototype-roadmap.core: blocked_after_full_attempt at argument |
| Core attempted / missing lanes | causal, empirical, executable, formal, source-synthesis / normative, transfer |
| Attempted local lanes | causal, empirical, executable, formal, source-synthesis |
| Missing or unproved lanes | normative, transfer |
| Strongest family bundle | Integrated governed lifecycle slices (end_to_end): Three versioned integrated slices, 12 cases, all ten lifecycle states, six observed effects, rollback/residual/quarantine outcomes, and an eleven-surface sealed epoch. |
| Negative controls | 20 named boundary injections; three exact rollbacks; partial-effect residual and quarantine; 11 rejecting mutations. |
| Accepted transitions | none |
| Maximum inference | Bounded local replay only; no deployment, whole-book proof, external effect authority, transfer, publication, or release claim. |
| Reproduction / next burden | Replay scripts/validate_p3_integrated_slices.py and scripts/validate_claim_family_terminal_program.py; fill the named atom-specific lanes under a new prospective protocol. |
85.17 Handoff
The roadmap gives the stack a build order, and the manuscript itself has to follow the same discipline while it keeps changing. Living Book Methodology defines that operating system for the book. It keeps manifest order, source queues, claim states, proof targets, validation commands, reader projections, audio-script paths, changelog entries, and release boundaries aligned so the project can improve without losing its own evidence trail.