Skip to main content

GenesisCode Roadmap: Agent-Native, Universal, Efficient, Self-Hosted, and Trustworthy

Last audited: 2026-08-11

Status: canonical strategic plan for the path from v0.2 to v1.0 and the post-v1 frontier. This file is not release evidence. Repository history, clean-clone verification, and remote CI evidence are mandatory under R0.1.

Program objective: fully complete this roadmap end-to-end in a beyond state-of-the-art manner: AI-first, universal across declared product and hardware profiles, super efficient, fully self-hosted above a minimal auditable host, reproducible, formally hardened, validated by executable evidence, and eventually able to conduct independently verified proof-carrying algorithm discovery through Genesis Foundry without enlarging production trust. This describes the durable program, not the state of any Codex goal or work session.

GenesisCode is not trying to be another general-purpose language with an AI plugin. It is a deterministic software substrate designed for AI-authored and AI-iterated systems. Its differentiator is the combination of compact machine-facing semantics, content-addressed identity, sealed effects, deny-by-default capabilities, deterministic replay, semantic patches, evidence-carrying packages, and a self-hosted toolchain whose claims can be independently checked.

Its first delivery wedge is a fully functional, agent-usable, self-hosted GenesisCode Core whose production semantics are owned by GenesisCode above a minimal stage0 and whose bootstrap and release claims can be checked independently. Its second wedge is a finite Genesis Foundry Foundation that calibrates proof-carrying search against that frozen Core and publishes the first independently governed genesis/algorithms package release. Its third wedge is a finite GenesisExamples/GenesisApps seed: executable teaching projects and maintained application systems that consume the released algorithm package and prove useful end-to-end construction, beginning with a library/CLI, MCP server, web server, and a separate full-stack site/application served by that stack. Its fourth adoption wedge is GenesisBench: expose the frozen language, disclosed algorithm package, released examples, and bounded documentation to a model, require it to learn an unfamiliar executable formal system and complete real work, and score artifacts mechanically rather than by style or judge preference. This ordering prevents benchmark scaffolding, mutable language behavior, target-model failures, or unverified search output from steering an unfinished language, library, or application contract. Retained internal conformance fixtures may improve GenesisCode before v1, but executable Foundry work waits for exact reviewed R9.4.f; the curated example/app seed waits for F2.r; and new model campaigns, held-out commissioning, public ranking, and GenesisBench release work wait for that finite seed. High-quality outputs may become packages, examples, applications, or training material only through separate provenance, license, safety, verification, deduplication, destination-review, and held-out-firewall processes. Search, demo, or benchmark success never makes generated output authoritative by itself.

The roadmap is intentionally ambitious, but it is not a list of unrelated ambitions. It is a dependency-ordered release program. The first priority is making GenesisCode safe and effective for agents now. Performance tiers, broad domain support, and recursive self-improvement follow only after the semantic and evidence foundations they depend on are real.


1. Product thesis and boundaries

1.1 North-star outcome

A local or remote AI agent should be able to learn the stable GenesisCode profile from a small, generated context bundle; create or repair a complete product; receive structured diagnostics; run it under explicit capabilities and resource bounds; inspect or replay every effect; package it with provenance and obligation evidence; and deploy it without trusting raw model output. Within a declared target profile, the application author writes GenesisCode plus GenesisCode-owned manifests, policies, schemas, tests, and assets. They do not hand-maintain HTML, CSS, JavaScript, TypeScript, Python, Rust, C/C++, Swift, Kotlin, Java, shader source, linker scripts, or deployment YAML to finish the product.

This is a one-language authoring contract, not a claim that browsers, operating systems, silicon SDKs, databases, GPU drivers, or every third-party library cease to exist. GenesisCode may compile to or generate standards-compliant foreign artifacts and may consume audited components behind the capability boundary. Those implementation details are content-addressed, reproducible, inspectable, replaceable, and never become an undocumented second source language for the application.

The v1 toolchain should be self-hosted above a deliberately small stage0 host. “Fully self-hosted” means the language-level frontend, typechecker, formatter, optimizer, patch engine, obligation logic, package/registry logic, documentation generators, and build orchestration are authored in GenesisCode and bootstrap to a reproducible fixpoint. It does not mean pretending that hardware, an operating system, a WebAssembly engine, or the minimal evaluator/effect host does not exist.

The program delivers four dependency-ordered outcomes with independent version, release, compatibility, and evidence authorities:

  1. GenesisCode: the language, runtime, package/build/deployment system, documentation, agent interface, and the governed GenesisExamples/GenesisApps distribution surfaces.
  2. Genesis Foundry + Genesis Algorithms: an isolated proof-carrying search system plus a destination-governed first-party .gc algorithm package collection. The engine proposes; the package maintainers independently verify, accept, version, and support ordinary library releases.
  3. GenesisBench: the self-hostable evaluation program for whether a model or agent can acquire an unfamiliar executable formal system from bounded evidence and construct trustworthy GenesisCode software using the same disclosed language and package surface available to users.
  4. Genesis Model: a profile-bound, self-hostable specialist trained only from governed, lineage-preserving trajectories after the training/evaluation firewall is proven.

Their flywheel is deliberate but not circular authority. GenesisCode defines executable semantics and constraints. Foundry searches frozen profiles and retains candidates, proofs, costs, failures, and counterexamples; independently accepted results are copied into ordinary versioned genesis/algorithms releases that contain no mutable research dependency. GenesisExamples teaches one concept or workflow with the smallest complete executable project; GenesisApps composes those capabilities and released algorithms into maintained, deployable systems. GenesisBench then measures acquisition and trustworthy use of that disclosed language/library/example environment and produces independently checkable failures and successes. Those failures improve the language, diagnostics, cards, tools, libraries, examples, applications, and tasks. Verified artifacts may enter packages, examples, applications, or a training corpus only through explicit destination-owned promotion. The Genesis Model proposes more artifacts but never scores, verifies, authorizes, or promotes itself. Outcome versions advance independently, and no search, package, example, benchmark, or model release silently changes language semantics.

After the Foundry Foundation release, that ecosystem may also operate GenesisChallenge, a separately governed, non-ranking contribution contest for solving real GenesisCode, Genesis Algorithms, Foundry, and, once independently released, GenesisBench problems. GenesisChallenge is not a fifth product, a benchmark track, a difficulty band, or a source of release authority. Its accepted work enters the destination repository only through ordinary independent review, verification, provenance, and promotion controls; open or prize-bearing work can never be mixed into a benchmark score.

Each product therefore has its own release train and acceptance authority:

Product Canonical release identity May consume Must never control Independent release condition
GenesisCode language/runtime/profile/source/artifact manifest plus separately versioned example/app catalog manifests named immutable Genesis Algorithms releases, typed benchmark failures, and reviewed promoted artifacts mutable Foundry state, benchmark scores, held-out custody, corpus admission, or model weights language/runtime/package/toolchain claims meet their own compatibility, product, assurance, and E4 gates; every released example/app meets its declared target and maintenance contract
Genesis Foundry + Genesis Algorithms Foundry snapshot/grammar/catalog/search/verifier/receipt manifest plus a separate signed package-release manifest frozen GenesisCode releases, prior immutable algorithm-library releases as baselines, and independently authored tasks/checkers production semantics, destination review, package acceptance, benchmark scoring, or same-run library mutation calibration is independently reproduced; each library export has destination-owned contract/provenance/proof-test/cost/compatibility evidence; production remains valid with foundry/ absent
GenesisBench protocol/epoch/scaffold/adapter/scorer/registry manifest frozen GenesisCode and genesis/algorithms releases plus disclosed model interfaces language or library semantics, model training, solver output, Foundry search state, or historical result rewriting construct validity, challenge scale, custody, reproducibility, governance, and public-methods gates pass for the named cohort
Genesis Model weights/tokenizer/runtime/training-lineage manifest admitted corpus artifacts and frozen language/benchmark profiles evaluator, held-out tasks, promotion rules, capability grants, or release signing the named model package meets its declared quality, safety, license, provenance, local-resource, and reproducibility profile

No outcome waits for an unrelated outcome to reach the same major version. GenesisCode v1 Core ships without requiring Foundry, Genesis Algorithms, the post-Core example/app catalog, GenesisBench, GenesisChallenge, or a Genesis Model package. The first Foundry Foundation begins only against immutable GenesisCode v1 Core. The first curated GenesisExamples/GenesisApps seed begins only against the first independently promoted genesis/algorithms release and remains runnable with foundry/ absent. The first public GenesisBench and GenesisChallenge releases begin only against that Core identity, algorithm release, and finite public example/app seed; later releases name explicit compatible ranges and do not wait for unrelated Foundry or universal-showcase expansion. A Genesis Model package supports explicit language, algorithm-library, example/app, and benchmark profiles and fails closed outside them. Portfolio milestones may require evidence from multiple outcomes, but a release is blocked only by its declared prerequisites and compatibility tests.

The benchmark is an instrument for independently measuring and continuing to harden a usable language, not a substitute for finishing the language or a post-language marketing exercise. Before v1, only model-free benchmark protocol conformance, shared regression fixtures, security maintenance, and preservation of already completed evidence may run; they cannot publish a new model comparison or divert the self-host critical path. Candidate generation and research may run only within the activation boundaries in section 3.18, while promotion of a task, result, package, language change, corpus artifact, model, or release is discrete, independently verified, append-only where historical evidence is involved, reversible where state changes, and incapable of rewriting prior results.

After the v1 trust baseline, Genesis Foundry becomes the first post-Core program: an independently bounded system for discovering and comparing algorithms from typed computational mechanisms. It consumes frozen GenesisCode semantics, CoreForm, obligations, artifacts, replay, and at most a prior immutable genesis/algorithms release; generates candidates under quarantine; verifies them through independent checkers; retains failures and counterexamples; and emits reviewed promotion proposals. It cannot install a rewrite, expand the kernel, alter a verifier, admit training data, or authorize a package, model, benchmark, or language release. A Foundry snapshot has its own schema/catalog/task/search/verifier/receipt identities and an explicit compatible GenesisCode range.

Genesis Algorithms is the supplied library output, not the search engine under another name. It is an ordinary, signed, semantically versioned GenesisCode package collection containing independently accepted reference algorithms, specialized variants, exact applicability contracts, deterministic fallbacks, provenance/licenses, complexity/cost evidence, examples, and conformance tests. It contains no Foundry runtime, mutable registry, candidate history, hidden oracle, or same-release circular dependency. Foundry may search against release N as a baseline and propose changes for N+1; destination maintainers reverify those proposals without the search engine before release. GenesisCode, Bench, applications, and models may depend on a named library release but never on foundry/ or mutable research state.

GenesisExamples and GenesisApps are governed GenesisCode distribution surfaces, not additional semantic authorities or independent language implementations. GenesisExamples contains compact, executable, concept- and workflow-focused projects; GenesisApps contains maintained production-style systems assembled from the same public packages and capabilities. Existing parser, conformance, benchmark-reproduction, and host-workflow fixtures remain truthfully classified as fixtures rather than being relabeled as applications. Every catalog entry binds exact GenesisCode, target-profile, dependency, and Genesis Algorithms identities; capabilities and budgets; build/run/deploy/replay commands; deterministic acceptance; source/generated/dependency inventories; provenance and license; maintenance owner; and evidence level. Entries use a named released algorithm where it materially serves the workload, never as dependency theater, and build without Foundry code, candidate history, or mutable research state. Public examples may define disclosed benchmark capability families, but a public solution is never copied into held-out scoring material.

1.2 What “amazing” means here

GenesisCode v1 must be strong in all of these dimensions at once:

  1. Agent legibility: the useful language surface fits within explicit context budgets and exposes machine-readable schemas, diagnostics, repair hints, examples, and capability contracts.
  2. Mechanical trust: generated code is never accepted because it looks plausible. Semantics, effects, policy decisions, patches, packages, and releases carry independently checkable evidence.
  3. Determinism: pure evaluation, canonical identity, effect logging, replay, concurrency, builds, and bootstrap outputs have specified cross-machine behavior.
  4. Efficiency: the common agent edit-check-run loop is interactive, the reference interpreter is credible, compiled tiers are validated, and long-lived services have bounded memory and resource behavior.
  5. Self-host authority: self-host claims describe semantic ownership and bootstrap closure, not merely a GenesisCode wrapper that routes to Rust implementations.
  6. Self-hostable infrastructure: no required cloud control plane, hosted registry, proprietary model, or external telemetry service is needed to build, verify, package, or deploy software.
  7. Practical reach: real programs across the declared v1 domain matrix work end-to-end, while unsupported domains are named rather than implied by broad marketing language.
  8. Replaceability: an independent implementation can pass the kernel and artifact conformance suites from the published specifications.
  9. Benchmark validity: model comparisons use precommitted tasks, deterministic scoring, reproducible run records, contamination labels, temporal challenge sets, and independently replayable evidence rather than opaque leaderboards.
  10. Learning flywheel: benchmark failures improve GenesisCode, verified successes seed useful examples and packages, and only governed artifacts enter a lineage-preserving corpus for a future GenesisCode-native model.
  11. One-language product construction: declared web, service, desktop, mobile, game, data/AI, embedded-Linux, and microcontroller profiles can be authored, tested, built, deployed, operated, and maintained without handwritten foreign-language glue.
  12. Authentic targets: a filename, archive, descriptor, headless plan, empty Wasm export, or shell smoke token never counts as an application. Target support means the artifact executes the declared workload in the real browser, OS, simulator/device, container/runtime, or board and exposes exact evidence of that execution.
  13. Proof-carrying discovery: post-v1 algorithm research preserves the exact task, grammar, semantic profile, search lineage, counterexamples, proof/test scope, cost samples, and promotion decision. A new hash is not novelty, bounded verification is not a general proof, and search output never becomes production authority by proximity.
  14. A useful algorithm commons: the supplied genesis/algorithms collection offers stable, documented, portable .gc APIs with exact behavioral and complexity contracts. Foundry discoveries can improve it, but baseline implementations, known algorithms, negative results, and independently reviewed fallbacks are first-class; novelty is never an admission requirement.

1.3 One-language authoring contract

Layer Allowed application-owned inputs Toolchain responsibility Acceptance rule
Language and product logic .gc, Genesis package/workspace manifests, capability policy, typed schemas, tests, assets, and generated-lock inputs Parse, type/effect check, optimize, compile, link, package, and diagnose No handwritten foreign source is required in a canonical starter or flagship product.
Web presentation Genesis document/component/style/route/accessibility IR Emit HTML/CSS/JavaScript/Wasm/service-worker artifacts and source maps as deterministic build outputs Browser-matrix behavior, accessibility, hydration/SSR identity, and network policy pass from the .gc source alone.
Native/mobile presentation Genesis app, lifecycle, UI, permission, asset, and platform-service declarations Generate platform project metadata, bindings, resources, signing plans, and executable packages Real install, launch, lifecycle, accessibility, and device/simulator tests pass; generated Swift/Kotlin/XML/project files are not edited.
Games, graphics, and media Genesis scene/ECS/render/shader/audio/asset descriptions and logic Generate validated shaders, pipelines, asset packs, and target adapters A real interactive frame/audio/input loop meets target budgets; a frame-plan hash alone is insufficient.
Services and operations Genesis routes, schemas, migrations, jobs, deployment and observability declarations Generate portable components, service/container/unit manifests, SBOM/provenance, and reviewable deploy plans The service handles real protocol traffic, persistence, upgrade/rollback, and fault injection without handwritten shell/YAML.
Embedded and hardware Genesis embedded-profile source, board manifest, pin/peripheral policy, static resource budget, and tests Select BSP/HAL components, AOT compile/link, generate startup/link layout, flash/debug/OTA plans, and hardware evidence The exact firmware boots and performs the workload on a named board; descriptor-only or host-bridge simulation does not count.
Extension packages Genesis bindings and capability declarations Import WIT/schema/header metadata and isolate audited foreign internals behind components Application authors remain in GenesisCode; foreign internals are explicit supply-chain dependencies, not ambient escape hatches.

The contract is profile-scoped. A target that cannot meet it is reported as unsupported or experimental. GenesisCode does not silently emit a misleading file extension, invoke an undeclared cloud compiler, or ask the agent to finish generated glue by hand.

1.4 Non-goals before v1

  • Syntax breadth is not a goal. New syntax must reduce total agent/human complexity or unlock a required semantic capability.
  • A production JIT is not automatically a v1 requirement. It becomes critical-path work only if the interpreter, bytecode VM, prelude snapshot, and warm daemon cannot meet the measured product budgets.
  • “Supports anything” is not an acceptable release claim. The north star is universal product construction through a finite, versioned profile and extension matrix; every release names exactly which products, hosts, boards, services, and lifecycle operations have authentic L4/L5 proof.
  • Raw LLM inference never enters the pure kernel. Model access is an explicit host effect with capability, provenance, budget, redaction, and replay rules.
  • One aggregate leaderboard never mixes raw-model acquisition, open agent scaffolds, GenesisCode-adapted systems, and embedded local systems. They answer different questions and remain separate tracks.
  • GenesisBench never requires one vendor’s API, chat template, special tokens, native tool-call encoding, reasoning-effort control, streaming protocol, tokenizer, hidden chain of thought, CLI, or hosted service. Provider-specific integrations are adapters and named campaign scaffolds, never the semantic benchmark contract.
  • Open-ended repository contributions, bounties, and improvement contests are not benchmark items. GenesisChallenge records them on a separate evidence and credit surface, and no contest outcome changes a GenesisBench cohort, rank, difficulty calibration, or historical score.
  • GenesisCode novelty is not evidence of training-data absence. Public fixtures establish conformance and practice, not temporal cleanliness.
  • Training a language model from scratch is not an early milestone. Begin from a capable, legally usable open base unless measured evidence shows that specialization cannot meet the declared local envelope.
  • “English and GenesisCode” describes the intended human/model interface, not the full source of domain knowledge. Changeable knowledge should move into versioned packages and capability schemas; stable general knowledge may remain in model weights; tools and user context remain explicit inputs.
  • A continuously running proposal loop is not continuous self-promotion. No model may grant itself authority, alter its evaluator, admit its own training data, or approve a language/model release.
  • GenesisCode does not require agents to modify generated Rust or shell glue. Any unavoidable host surface must be small, generated where possible, versioned, and independently tested.
  • Passing one local script or the existence of a JSON report is not release proof.
  • Genesis Foundry is not a GenesisCode v1 prerequisite, an autonomous optimizer, a general algorithm-invention claim, or permission to enlarge the kernel. Before v1, exploratory boundary/schema notes may be retained as non-normative design inputs when they do not preempt the active release path; F2 execution and completion begin from a separate bounded workspace only after the production semantics and evidence interfaces it consumes are frozen. GenesisExamples/GenesisApps wait for finite F2.r, and GenesisBench/GenesisChallenge wait for finite R8.3.b, not open-ended Foundry or universal-showcase expansion.

2. Truth model

The project currently has broad functionality and many gates, but a single green checkmark hides important differences. From R0 onward, every capability and public claim uses this maturity ladder:

Level Name Meaning
L0 Specified Normative behavior, failure semantics, resource model, and compatibility rules exist.
L1 Implemented A reachable implementation exists; experimental wiring and wrappers qualify only for L1.
L2 Verified Positive, negative, differential, property, and boundary tests pass on the reference host.
L3 Reproducible Hermetic checks pass from a clean checkout on the supported host matrix with pinned inputs.
L4 Product-proven The feature succeeds in representative agent and domain workloads under declared SLOs.
L5 Release-attested A signed, immutable release evidence bundle binds source, toolchain, tests, artifacts, and claims.

Rules:

  • feature_matrix.md must report the level, not a binary checkmark.
  • A checked roadmap task cites durable evidence and its exact input identity. Mutable .genesis/perf/ files are local observations, not L5 evidence.
  • A check_* command is read-only. It must never repair state, regenerate reports, download undeclared inputs, or turn a missing artifact into a pass.
  • An update_* command may regenerate derived material, but its output must be reviewed and then checked separately.
  • Release evidence is content-addressed, schema-versioned, reproducible offline from declared inputs, and independently verifiable without trusting the producing CLI.
  • Routing a command through a .gc entrypoint establishes reachability, not self-host semantic ownership.
  • Public documentation states the lowest level that all supported platforms satisfy.

2.1 Evidence classes

Class Purpose Durability
E0 local observation Development timing, debug output, exploratory report Ignored and disposable
E1 checked test Named test/golden/property result tied to a source revision Reproducible in-tree
E2 gate manifest Machine-readable inputs, commands, versions, outputs, durations, and resource use Versioned schema; generated artifact may be ignored
E3 release candidate bundle Full host-matrix evidence, negative controls, SBOM, provenance, benchmark samples Immutable candidate artifact
E4 release attestation Signed hash tree over source, binaries, evidence, compatibility profile, and claims Published and independently mirrored

2.2 Definition of done for a task

A task may be checked only when its line contains done YYYY-MM-DD; evidence: <stable path or test>; input: <revision/hash>. The cited check must include at least one negative control when the task protects a trust boundary. Performance tasks include raw samples, machine profile, variance, warm/cold distinction, and a reproducible command. Documentation alone cannot close implementation work.

2.3 Open-standards posture

GenesisCode should define new formats only for genuinely GenesisCode-specific semantics. It should pin and profile mature open standards for envelopes, transport, components, updates, and inventories rather than create gratuitously incompatible substitutes. Versions below are the audited 2026-07-10 baseline; R0 must pin exact schemas and provide an explicit upgrade process rather than following latest implicitly.

Area Adopt/profile GenesisCode-specific extension
Agent transport MCP 2025-11-25 over stdio first, with version/capability negotiation; experimental Tasks are negotiated and never required by the core profile Semantic workspace snapshots, capability minimization, evidence IDs, and deterministic tool-result contracts
Attestation framework in-toto Attestation Framework v1.2, Statement v1, and an explicitly pinned authenticated envelope/bundle profile Predicates for semantic patches, obligations, replay verification, bootstrap witnesses, and agent provenance
Build/source provenance SLSA 1.2 provenance and verification expectations Hermetic/reproducible-build claims, Genesis profiles, semantic inputs, and independent verifier results
Portable extension ABI WebAssembly Component Model and WIT with WASI 0.3 as the async-capable candidate profile and a declared WASI 0.2 compatibility strategy Mapping Genesis values/effects/capabilities/resource charges to component worlds
Package/update trust TUF 1.0.34 threshold roles, delegation, freshness, rollback/freeze protection, and recovery Content-addressed package graph, evidence policy, transparency proofs, and offline promotion
Software inventory SPDX 3.0.1 or a later explicitly profiled compatible version Genesis package/profile/capability/evidence relationships not represented by the base SBOM

The project may use narrower subsets or stricter rules. Every deviation records why the standard is insufficient, how interoperability is retained, and who independently verifies the extension.

2.4 Foundry discovery maturity

Foundry results use a separate ladder so research artifacts cannot borrow GenesisCode’s L-levels or a product release’s authority:

Level Meaning Explicit non-claim
FD0 Generated The canonical candidate graph parses and its lineage is retained. Not known to execute or satisfy a task.
FD1 Executable It executes inside the declared sandbox on at least one admitted instance. Not correct, general, fast, or safe to promote.
FD2 Bounded verified Independent checks pass for the exact finite domain, properties, and semantic profile in the receipt. Not a proof outside those bounds.
FD3 Generalized Precommitted unseen sizes/distributions, metamorphic properties, adversaries, and ablations pass. Not historical novelty or production readiness.
FD4 Independently reproduced A separately provisioned verifier/operator reproduces correctness and cost evidence from the closed bundle. Not authorized for a production destination.
FD5 Promoted The destination’s independent authority accepts a semantic patch/package/rewrite under named rollback and compatibility rules. Not valid for any unlisted profile, target, or future version.

Proof status remains orthogonal: unproved, tested, exhaustive-bounded, solver-certified, proof-assistant-checked, and independently-replayed are never collapsed into one verified flag. Semantic quotienting likewise means quotienting by a named certified equivalence relation, not deciding arbitrary program equivalence.


3. Historical audit log and current correction

Sections 3.1-3.23 preserve the dated decisions and counterexamples that produced the current plan. They are historical records, not an executable queue. When an older entry conflicts with a later entry, the latest dated correction controls; section 3.24 is the current correction. Current task selection comes only from section 10 and the machine-generated execution slice.

3.1 What exists

  • An 18-crate Rust workspace implements the kernel, effects, CLI, types, optimizer, obligations, patches, packages, registry, VCS, WASI, WebAssembly, graphics, and runtime benchmarks.
  • The pure-kernel, seal, no-user-panic, deterministic effect log/replay, deny-by-default capability, and self-host-boundary gates exist and currently pass when their prerequisites are built.
  • Production command routing is artifact-only/self-host-first, strict fallback guards exist, and stage2 WebAssembly execution uses wasmi.
  • Effect rows and deterministic concurrency specifications/implementations already exist. Their roadmap work is audit, completion, and product proof, not invention from scratch.
  • The warm command and JSON request protocol exist. A generated MCP server does not yet appear to be the default agent interface.
  • The authoring skill, pointer guide, prompt pack, recipes, agent index, generative workloads, and capability gauntlets exist and pass static structure checks.
  • The project has substantial domain wiring across services, browser, data, process/network, GPU, hardware, mobile deployment, collaboration, UI, graphics, and XR.
  • Recent interpreter work added inline small integers, hybrid vectors, compiled parse fast paths, fast primitive wrappers, n-ary application support, and inline environment slots.
  • A local strict workload measured approximately: fib(25)=163ms, vector build =6ms, map build =402ms, string concatenate =15ms, self-host parse =3656ms, dispatch =1211ms. These are E0 observations, not release claims.

3.2 Material gaps found by this audit

  1. ROADMAP.md is untracked, so the nominal canonical plan is not yet repository history.
  2. feature_matrix.md reports a nearly all-green binary view and “Known gaps: none” despite open semantic, performance, self-host, agent-training, evidence, and release work.
  3. Several evidence checks treat absent mutable reports as a reason to regenerate them. Audit and mutation are therefore conflated.
  4. Resolved 2026-07-10: scripts/check_root_lock_policy.sh no longer assumes Python 3.11 tomllib; its required TOML subset is checked by a POSIX awk parser with embedded adversarial controls.
  5. Cold hardening checks can rebuild much of the workspace. One no-panic/TCB run took about 182 seconds and regenerated roughly 4.8 GiB under .genesis/build/cargo, making the default agent loop too expensive.
  6. The two largest Rust crates, gc_effects and gc_cli_driver, are roughly 40k and 32k Rust lines respectively. This concentration increases review, compile, and semantic-ownership risk.
  7. Self-host .gc sources are much smaller than the remaining Rust semantic surface. Current command routing does not prove that typechecking, obligation decisions, packaging, canonicalization, or patch semantics are GenesisCode-authoritative.
  8. The current authoring bundle is broad but not yet a compact, versioned, training-ready SDK. It lacks a stable diagnostic catalog, held-out task corpus, leakage policy, and model-agnostic scorecard.
  9. Interpreter closure capture, linked lexical frames, let allocation, long tail-loop behavior, and map performance remain open.
  10. Heap ownership, cycles, allocation accounting, deterministic out-of-memory behavior, and long-lived daemon leak limits are not yet closed as language/runtime contracts.
  11. The existing roadmap places bytecode/JIT work before the agent loop and treats a heavyweight JIT as mandatory without a measured decision gate.
  12. Version/profile compatibility, migration tooling, release evidence authority, gate cost, and disk lifecycle need first-class ownership.

3.3 Immediate interpretation

GenesisCode is ready for controlled experimentation and agent-assisted development inside the repository. It is not yet ready to be presented as a stable training target or trusted production language. The shortest path to that status is R0 plus R1, not a JIT. The shortest path to a defensible “fully self-hosted” claim is the ownership ledger and bootstrap program in R4, not more routing coverage.

3.4 Independent review delta: 2026-07-14

An independent dirty-worktree review exercised the newly built trust and agent surfaces rather than accepting their generated reports. It confirmed that the MCP server, warm protocol, structured diagnostics, compact cards, independent evidence verifier, signed failing baselines, and most of the gate architecture are substantive. Its full local run reported 1,565 passes, five failures, and 15 ignored tests across 269 binaries. Those counts are E0 observations, not clean-clone or release evidence.

The review also established the following release-blocking or roadmap-relevant deltas:

  1. Denied-effect replay diverges between Rust and self-host engines on both native and WASI mirrors. This is one deterministic trust-path defect, not two unrelated test failures; R0.2.f owns its closure.
  2. Spawn and persistent host-bridge kill/reap stress tests race under load, cli_agent_plan is load-sensitive, and a default regression test recursively launches a multi-minute changed-fast pipeline. R0.2.g must make the default suite hermetic and repeatably load-stable without hiding failures behind ignores.
  3. The new MCP/warm code has 14 Clippy warnings. R0.2.h makes warning-free all-target/all-feature builds part of the ordinary checkpoint, including newly generated modules.
  4. The current work accumulated hundreds of dirty/untracked files without a remote checkpoint, so neither bisectability nor the rewritten GitHub CI has been proven. R0.1.a and R0.1.g are the first priority.
  5. Generated state reached about 59 GiB and deterministic cleanup classified about 51 GiB as rebuildable. Cleanup safety is strong, but automatic quota/lease enforcement is still absent; R0.4.g owns bounded steady-state caches.
  6. Startup and fib observations drifted backward while PB-5 map construction and PB-7 self-host parsing remain honestly below budget. R2.1 and R2.3 prioritize general runtime fixes and regressions over microbenchmark-specific shortcuts.
  7. The review found workload-specific recognizers in crates/gc_kernel/src/compiled_runtime/patterns.rs; R2.1.g has since deleted that production path, installed a governed tombstone, and added fail-closed scans against equivalent benchmark dispatch. Its removal exposed an honest PB-4 regression that R2.1.h must close through semantically general optimization.
  8. MCP can suppress accepted in-flight responses when stdin closes, and concise human diagnostics are sometimes too terse. R1.3.e and R1.2.f close those lifecycle and usability gaps without creating a second semantic API.
  9. The trust-infrastructure investment was necessary, but the project now has enough governance machinery. R0.4.h imposes consolidation, and milestone effort shifts to transactional agent sessions, corpus/held-out evaluation, the generated authoring skill, instruction/content separation, PB-4/PB-5/PB-7, and product proofs.

3.5 Three-product direction review: 2026-07-15

The benchmark-first product review is accepted with stricter normalization. GenesisCode, GenesisBench, and Genesis Model are coupled programs, but their versions and authorities remain separate. GenesisBench measures cold acquisition of an unfamiliar executable formal system plus trustworthy software engineering, not syntax imitation. The current 27 public cases are nine task lineages under three context conditions and cannot support 27 independent observations. A credible Preview therefore requires lineage-scale private challenges, four isolated tracks, a fixed reference scaffold, canonical adapters, predeclared statistics, useful post-release overlays, and maintenance tasks. The English-intent surface begins before the specialist model as a quarantined data-generation instrument; the specialist begins from an audited open base only after the corpus firewall passes. Model inference remains an explicit effect, general-domain knowledge remains available through weights/packages/tools/context, and continuous proposal generation never acquires continuous promotion authority.

3.6 Independent implementation review delta: 2026-07-16

A second independent review exercised GenesisBench end to end through real CLI run, validation, replay without model access, independent rescoring, deterministic bundling, and registry surfaces. It confirmed the three-product split, docs publication, replay parity, host-bridge hardening, warning-free build, and substantive benchmark trust core. Its reported 1,340 passing, two failing, and 19 ignored tests across 227 binaries are E0 observations from that checkout, not release evidence. That checkout exposed generated-view and ignored-output hygiene defects. A separate hosted Linux run then exposed process-boundary defects that the local summary did not classify: aggregate resident-memory policy had been implemented with a per-process virtual-address-space limit; successful children that closed standard input early were treated as FFI failures; WebSocket operations held a global network lock across blocking I/O; and persisted bridge counters allocated duplicate IDs under concurrent processes. These distinct observations remain distinct in the evidence record.

The review changes this roadmap in seven ways:

  1. Generated views become a transitive closure, not a final manual regeneration step. A change to any canonical input must make focused local checks and remote CI fail before publication, and one explicit updater must regenerate the complete declared dependency closure without checks mutating state.
  2. Test-lane membership remains auditable. Moving tests from the default lane to stress, performance, platform, or release lanes requires a manifest diff, rationale, owner, trigger, and proof that required CI still executes them; a lower default test count is never accepted on count alone.
  3. Repository hygiene distinguishes tracked/source roots from classified generated roots. Ignored Quarto, cache, build, simulator, and benchmark outputs are excluded from source-topology scans only through the generated-state policy, remain bounded/cleanable, and cannot hide untracked source or evidence.
  4. R1.4.o begins with a staged real-model reality gate: one immutable frontier general model attempts every public task class before broad baseline publication. The result may be poor or invalid; preserving the complete failure record is the value. A deterministic mock proves conformance, not model usefulness.
  5. GenesisBench may remain a project-controlled Preview, but a public benchmark/leaderboard and E4 claims require an independent non-project reviewer/operator, conflict disclosure, and separately controlled signing or mirroring authority. Agent-family diversity helps challenge concentration but does not impersonate human or organizational independence.
  6. Genesis Model receives a formal research-transfer lane. Project Theseus is treated as a report-first implementation reference and The ASI Stack as a living synthesis, claim/evidence ledger, and hypothesis source. Neither project, prose corpus, dashboard, or latest report becomes model architecture or training authority by proximity. R8.5.i and MQ-1 through MQ-8 require clean pinned releases, permission/license classification, artifact-level lineage, replay where possible, Genesis-specific replication, simpler baselines, ablations, negative/null retention, privacy/firewall review, and independent promotion decisions before a learning affects tokenizer, architecture, data, training, routing, memory, inference, or deployment.
  7. Cross-platform boundary claims name the exact quantity, scope, enforcement mechanism, sampling/error bound, and fallback semantics. Process-tree memory cannot be substituted with per-process address space; successful early pipe closure is not equivalent to child failure; blocking stream I/O cannot monopolize unrelated bridge state; and persistent IDs require process-safe allocation. Native single-run success never substitutes for supported-host CI plus repeated high-contention and lifecycle stress.

Point 1 had no implementation owner despite being stated as an invariant. This review therefore reopens the current M0 acceptance state through R0.4.i; prior M0 evidence remains immutable historical evidence, but the current tree is not allowed to claim complete generated-authority closure until that task passes. R1.4.o and broad baseline publication depend on the reopened task.

The current local Theseus and ASI Stack worktrees are active and dirty, so their moving state is intentionally not pinned here. R8 freezes clean content-addressed source editions when the model lane actually begins; importing today’s mutable state would create false currentness and contaminate later experimental attribution.

3.7 Universal product and hardware audit delta: 2026-07-16

This audit tested the repository against the stronger one-language authoring objective rather than inferring product support from operation names, archive suffixes, or deterministic mock workflows. The existing foundation is useful: GenesisCode already has capability-shaped filesystem, process, network, database, crypto, media, plugin, browser, graphics, GPU, XR, package, and deployment surfaces; native/WASI/Wasm hosts; deterministic frame planning; mobile/edge target labels; and agent workflow fixtures. However, the generated capability ledger correctly rates the broad surfaces only L1. Reachability is not a product stack.

The audit establishes these gaps and owners:

  1. gcpm build currently accepts web, desktop, service, ios, android, edge, and service-runtime, but several outputs are envelopes rather than executable products. The iOS and Android archives contain metadata, Genesis source, and a runtime descriptor without an executable app runtime; edge/service Wasm exports an empty function; web/desktop/service packages are CoreForm maps; launch scripts verify bytes and print synthetic boot/smoke tokens. R6.3 and R6.5-R6.7 must replace suffix-level conformance with authentic compilation, installation, launch, protocol, and device evidence while preserving these formats only as explicitly labeled descriptors if still useful.
  2. Browser support exposes a small first-party window/input/audio/storage profile and deterministic graphics plans, not a complete website/application authoring system. There is no normative Genesis document/component/style/layout/router/forms/accessibility/SSR/hydration/PWA contract or compiler that removes handwritten HTML/CSS/JavaScript. R5.6 and R6.5 own that product layer.
  3. gc_gfx provides valuable headless resource/frame codecs and scene/UI planning, but a frame hash is not a game or production UI toolkit. The repository lacks a complete retained/reactive app model, accessible widgets, text/layout/input composition, ECS, deterministic simulation clock, physics/collision, animation, asset pipeline, Genesis-authored shaders, game networking, save state, editor/inspector, and real multi-target render/audio loops. R5.6, R5.8, R6.6, and R8.2 own closure.
  4. Mobile and desktop workflows do not yet prove installable applications, native lifecycle, permissions, accessibility, input, sensors, store packaging, signing, crash reporting, update/rollback, or real device behavior. R6.6 turns generated project glue into a toolchain-owned detail and requires simulator plus physical-device evidence without application authors editing it.
  5. Service/data primitives exist but not a cohesive batteries-included product platform. Typed routing/middleware, authentication/session policy, schema/query/migration tooling, queues/jobs/cache/object storage/email, observability, configuration/secrets, and local production orchestration need one compatible package and capability model. R5.7 and R6.4 own it.
  6. The hardware workflow is a shell plugin fixture, not a GPIO or firmware implementation. There is no bare-metal/no-OS Genesis profile, static-memory verifier, interrupt semantics, AOT MCU backend, linker/startup generation, BSP/HAL registry, GPIO/ADC/PWM/I2C/SPI/UART/CAN/USB/BLE/Wi-Fi surface, flasher/debugger/OTA flow, board simulator, or hardware-in-the-loop gate. R5.9 and R6.7 add Embedded Linux first, then representative RP2040, ESP32, and Arduino-compatible MCU families without silently changing core semantics.
  7. Raspberry Pi-class Linux devices are best treated first as reproducible linux-arm64 hosts with explicit GPIO/I2C/SPI/serial capabilities, not conflated with bare metal. Tiny MCUs require a validated constrained Genesis profile with closed-world AOT, statically bounded memory/stack/effects, fixed-width numeric types, and explicit unsupported features; they cannot inherit arbitrary-precision/dynamic host behavior by accident.
  8. Foreign ecosystems remain necessary under the hood, but the application authoring boundary is undefined. R5.5 expands schema/WIT/header import, generated bindings, component sandboxing, and audited package adapters so a missing domain library does not force the user or model into a second source language.
  9. The current ten-archetype gauntlet and five flagship systems under-sample the north star. R8.2 expands to eighteen authentic product archetypes and R8.3 requires maintained cross-target products, real hardware, upgrades, and zero handwritten foreign application source. GenesisBench later measures these tasks by independent product lineage rather than multiplying target conditions.
  10. Genesis Model cannot compensate for a missing platform contract. It may retrieve domain packages and propose GenesisCode, but compiler/runtime/package/deployment evidence remains authoritative. The model succeeds only when a general model and the eventual specialist can build and maintain products through public Genesis interfaces without private glue or privileged target knowledge.

The practical conclusion is not to put every framework into the kernel. Keep the language and TCB small; build breadth as versioned Genesis libraries, capability profiles, target adapters, component packages, generated bindings, and evidence-backed product kits. Common product paths are first-party and coherent. Long-tail domains use the same extension boundary without weakening the one-language authoring contract.

3.8 Independent runtime and benchmark-readiness review delta: 2026-07-17

The fourth independent review confirms that the runtime sprint materially improved the project rather than merely increasing its governance surface. The workload recognizers are gone and permanently guarded against reintroduction; PB-4 and PB-5 recovered through general execution improvements; cycle collection, deterministic logical allocation/live-heap metering, fallible bulk allocation, hostile-blob bounds, and native worker-process containment now have differential and adversarial coverage. These are implementation results, not evidence that every roadmap product is complete. PB-1/startup drift, PB-7, persistent-sharing stress, host-handle teardown, authentic model runs, and universal product execution remain open under their existing owners.

The review changes execution priority and acceptance in six ways:

  1. R1.4.o is now the immediate product-proof lane. Its first named campaign uses Codex CLI with the provider-catalog slug gpt-5.6-luna at xhigh reasoning effort, followed by resource-fit local MLX models. Codex CLI is an Open Agent scaffold, not a raw-model Cold Acquisition result; the same underlying model may enter Cold Acquisition only through the fixed reference scaffold and an independently pinned immutable provider revision.
  2. A model alias, CLI display name, or local directory name is not an immutable model identity. Every publishable attempt binds the exact Codex/adapter/runtime executable digest and version, requested model/effort, provider-returned revision/fingerprint when available, weights and tokenizer digests for local models, context/scaffold identity, repository snapshot, capability envelope, complete event transcript, and all missing identity facts. Missing immutable identity forces an explicit unranked/unknown label; it never blocks preserving the attempt.
  3. The reality gate fires before further benchmark armor. One public case from each of the nine task classes runs under one frozen predeclaration and no between-class prompt, scaffold, tool, or budget tuning. Only after all nine attempts, including failures and invalids, validate and replay may the campaign expand to all 27 public conditions. Results feed typed language/documentation/tooling defects, but benchmark outcomes never weaken deterministic acceptance.
  4. The existing command adapter is insufficient for claiming an agentic Codex CLI run because it has an empty working directory and no workspace tool contract. R1.4.o must version a capability-minimal Open Agent workspace-harness revision that provisions a copy-on-write benchmark checkout, makes frozen source/runtime/docs readable, confines edits to declared candidate paths, mediates tools through the disclosed scaffold, blocks private held-out and ambient user state, records every event, and destroys no failed attempt. It extends the front door rather than silently broadening the fixed Cold Acquisition adapter.
  5. R1.4.o owns a campaign-execution regression beyond R0.2.g’s retained closure because a scoring integration test exceeded two minutes and failed only under the full parallel suite, while a scheduled hosted run remained queued long enough to lose operational value. Test-lane ownership must be based on measured resource behavior: long benchmark/scoring workflows move to a required serial or stress lane with an explicit trigger, while their small contract core remains in the default lane. Scheduled, push, and pull-request concurrency groups must neither starve cron evidence nor permit stale runs to consume capacity indefinitely.
  6. Roadmap growth is now scope-neutral before each milestone. Every new task must name its milestone, critical-path dependency, evidence predicate, and an existing task it refines, defers, consolidates, or makes unnecessary. A milestone may grow only with an explicit capacity/sequence decision and updated exit criteria; frontier breadth cannot displace the next falsifiable product proof.
  7. A 2026-07-18 local release-full audit passed the common gates, generated-authority freshness, warning-denied compilation, no-panic compilation, kernel tail stress, native/WASI workflow parity, real Apple M1 GPU parity, and source-decomposition parity before reaching the existing strict GCPM target-runtime gate at 45 minutes. That gate correctly rejected all four targets because only synthetic adapters were available and no real iOS, Android, edge, or service runtime commands were declared. This is a truthful product-readiness blocker, not permission to weaken CI=true, forge non-synthetic evidence, or raise GB-4 around duplicate work. R0.4.j refines the already-complete gate-budget work: reuse identity-matched private evidence across nested checks, provision a named reference runner with authentic target integrations, distinguish infrastructure failure from expected unsupported-product status, and re-establish the complete 45-minute/20-GiB release contract before claiming a release candidate.

3.9 Independent first-contact and benchmark-signal review delta: 2026-07-18

The fifth independent review exercised the retained Open Agent campaigns, local workspace suite, Clippy, and runtime timing. Its reported 1,397 passes, zero failures, 23 ignored tests across 231 binaries, warning-free Clippy, and approximately 0.34-second fib(25) observation are E0 checkout observations rather than release evidence. More importantly, it confirmed that the third Codex CLI/gpt-5.6-luna xhigh commissioning campaign preserved eight verified, independently rescored attempts and one genuine model/noneditable-input-drift deployment invalid after earlier harness faults were typed and removed. The observed seven perfect small-class scores and one 9,797-basis-point repair score show that this public small-context commissioning set has little remaining frontier discrimination. This is a prioritization signal, not the BQ-5 saturation trigger: one mutable-provider Open Agent family, one epoch, and one invalid cell cannot establish formal multi-family saturation or a ranked model claim.

The review adds four bounded corrections to the active sequence:

  1. R0.4.k reopens generated-authority completeness without rewriting R0.4.i’s historical evidence. Changing cli_genesisbench_front_door.rs left the GenesisBench protocol closure and check_agent_authoring_bundle.sh gate input stale while the transactional graph still reported itself fresh. The immediate generated outputs must be repaired, and the graph must then discover transitive gate-input dependencies so the same class of partial freshness cannot recur.
  2. R1.4.o keeps the failed expansion predicate immutable. The eight-verified/one-invalid commissioning campaign cannot be retroactively expanded under its old predeclaration. A complete 27-condition Luna follow-up requires a new frozen campaign/cohort after the deployment failure is dispositioned through general diagnostics/cards/tooling rather than task-specific prompting; all cells, including invalids, remain publishable. Medium and large contexts carry priority because the small public tier is already weakly discriminative.
  3. R1.4.q distinguishes a valid held-out protocol and commitment ledger from a useful private benchmark. The current 90-lineage commitment structure reports private_verified=false; M1 therefore requires at least 45 real, independently authored or independently governed private payloads with complete custody, oracle, alternative-solution, leakage, and blind-execution evidence. Placeholder commitments, generator labels, or balanced metadata cannot satisfy this requirement.
  4. R1.4.o’s publication step becomes an immediate commissioning deliverable: publish the benchmark card, methods, exact campaign chronology, all invalids and harness corrections, the deployment failure disposition, limitations/non-claims, and reproducible first-results analysis before presenting a leaderboard. Then close R0.4.j and begin R2.3 startup/snapshot work as the standing runtime priority. This refinement consumes existing M0/M1 hardening and benchmark-publication capacity; it adds no benchmark track, task class, runtime tier, or product target.

3.10 Release-graph and active-delivery audit delta: 2026-07-21

This audit compared the prose release contract, machine-readable roadmap graph, local branch, remote branches, and protected CI rather than treating a generated manifest as sufficient. It found six release and ordering defects that could make an otherwise rigorous plan execute incorrectly:

  1. The local v0.5 custody tranche is preserved at commit 84a09a4, but it is not on origin/main. PR #5 is mergeable but blocked because test_suite correctly rejects the undeclared GB-8 Python packaging module scripts/lib/genesisbench_mlx_custody.py; the aggregate test job consequently fails. This is current R0.4.k evidence, not an excuse to bypass CI. No benchmark inference or downstream publication may use that tranche as a released authority until the exact commit or a reviewed successor passes required checks, is promoted to main, and the temporary branch is deleted.
  2. The policy made R1.4.o start-ready even though the practical order required R0.4.k first. R1.4.o now depends on R0.4.k. Campaign predeclarations may be prepared offline, but no model inference begins against an unpublished or red authority.
  3. The prose promised independent GenesisCode, GenesisBench, and Genesis Model release trains, while the graph made R9.1 wait for all of R8.5, including specialist training and packaging. Milestones are now explicit portfolio checkpoints with product-scoped submilestones. GenesisCode v1 waits for GenesisCode product proof, not a GenesisBench public release or Genesis Model package; each optional linked product must still be independently releasable and truthfully marked compatible, preview, or absent.
  4. R8.5 sequenced corpus creation before the training/evaluation firewall and put independent benchmark adoption after the complete model-training chain. The corrected graph forks after verified ecosystem promotion: the GenesisBench adoption lane can close independently, while the Genesis Model lane establishes BQ-7/BQ-8 isolation before corpus admission and joins research-transfer, teacher, and base-model evidence before any training.
  5. The mandatory local upgrade-health profile reproduced a host-bridge lifecycle failure: spawn_per_op_timeout_kills_bridge_processes_and_recovers returned gpu/bridge-reap because the process group survived repeated termination sweeps. P1.5 records the active defect and R2.2.f remains the sole roadmap owner; M1-A daemon sign-off and every robust-service claim wait for cross-host bounded kill/reap and no-surviving-descendant evidence.
  6. The generated v1 ancestry could bypass R0 checkpoint discipline, R1.7 typed intent, R2.4 runtime observability, R3.3 optimizer/tier-controller closure, and R3.4’s explicit JIT-or-no-JIT decision; R1.5 also consumed only terminal R1.4.n rather than the complete public R1.4.a-n authority. The corrected graph explicitly fans R1.4.a-n into R1.5.a, orders R2.4 before R3.3 and R3.3 before self-host compiler migration, makes R8.1 wait for R1.7, makes R9.1 wait for the R3.4 decision, and fans every non-sequential R0.1 leaf into release assembly. Real-model campaigns and private/public GenesisBench governance in R1.4.o/q/p, the R8.5 benchmark/model lanes, and optional JIT implementation remain outside GenesisCode’s unconditional release dependency.

Capacity does not increase. Until the protected v0.5 tranche is green on main, R0.4.k is the only repository-changing publication lane. Private challenge commissioning may proceed under role-separated custody without target-model access, and analysis/specification work may continue without claiming the unpublished runtime. After promotion, the current executable tranche is: finish R0.4.j; complete the agent-authoring lane and R2.2.f daemon/handle closure; run the separately frozen Qwen3 4B, Qwen3 8B, and Luna cohorts; complete R1.4.q/p; close R2.1.h; then make R2.3 startup/snapshot work the standing runtime priority. R4.1 may proceed only as non-preemptive specification/ownership work until M1-A and M1-B are secure.

3.11 Genesis Foundry research-direction delta: 2026-08-03

The Foundry proposal belongs in this ecosystem because GenesisCode already has most of the trust substrate algorithm discovery needs: canonical CoreForm identity, pure deterministic evaluation, explicit semantic profiles, e-graph infrastructure, obligations, replay, semantic patches, content-addressed artifacts, and separate proposal/promotion authority. It does not belong in the production optimizer or the v1 critical path. The audited direction makes six corrections to the initial concept handoff:

  1. Foundry is a separately versioned research program that depends one-way on frozen GenesisCode interfaces. Production crates and the default genesis CLI never depend on Foundry code; accepted results cross only through reviewed versioned artifacts and semantic patches.
  2. Arbitrary semantic equivalence, mechanistic duplication, and historical novelty are not decidable by declaration. Every quotient, law, dominance result, and novelty label names its finite grammar, semantic profile, proof/test relation, literature-search limits, and defeaters. Unknown remains unknown.
  3. The first slice is pure, finite, and calibrating: fixed-size bounded sequence/comparator tasks, deterministic enumeration, symmetry reduction, independent exhaustive verification, mutation controls, and known-result rediscovery. General recursion, effects, mutable heap graphs, LLM search, Circle Calculus, and Theseus scheduling enter only after that slice proves the artifact and authority model.
  4. The proposed nine-crate decomposition is a target boundary map, not permission for premature fragmentation. Begin with one separately locked foundry/ workspace and a small number of modules/crates; split only when dependency, compile-cost, ownership, or independent-verifier evidence justifies it.
  5. Candidate execution is hostile-input execution. It receives no ambient filesystem/network/process/model authority, has deterministic logical resource limits plus hard host containment, cannot observe hidden verifier state, and cannot alter task, reward, checker, registry history, or promotion policy.
  6. Foundry quality is measured against simpler baselines. Search-space compression, rediscovery rate, counterexample yield, verifier false-accept controls, Pareto stability, abstraction reuse, total compute/storage, and human review burden are retained with negative results. A more complex search engine, law, or abstraction is rejected unless a predeclared ablation shows net value.

F2 takes this direction to its logical conclusion after v1 while allowing only non-preemptive schema/ADR research earlier. The initial browser draft remains a design input, not normative authority; F2’s checked artifacts and independently replayable evidence become the source of truth.

3.12 Portfolio concentration and model-readiness audit delta: 2026-08-03

The latest project-wide review is accepted where repository evidence supports it. The project has enough L1 breadth; the dominant risk is concentration failure, not missing ambition. This revision makes five execution corrections:

  1. Independent utility is mandatory. GenesisCode and GenesisBench release on their own evidence and remain useful with no Foundry workspace or Genesis Model package. Foundry is research infrastructure, not a language or benchmark dependency. The Genesis Model is the final optional product and cannot become a hidden prerequisite for authoring, verification, deployment, scoring, or release.
  2. v1 is finite. GenesisCode v1 Core freezes the language, deterministic runtime, package/evidence system, agent interfaces, native/WASI path, self-host authority, and one authentic CLI/WASI plus service profile. Web/UI, desktop/mobile, game/media, Embedded Linux, MCU/boards, GPU/XR, and later targets are independently versioned platform packs. Every pack retains the one-language and authentic-target burden, but an unfinished pack cannot hold the core release hostage. The eighteen-archetype gauntlet and ten flagships become portfolio maturity evidence rather than v1 Core prerequisites.
  3. Foundry calibration precedes model implementation. This audit originally made F2.a-q the MF calibration boundary and all F2.r-z optional. Section 3.25 later strengthens the finite boundary to F2.r so the independently promoted Genesis Algorithms release also precedes Bench, Challenge, and Model; F2.s-z remains optional expansion.
  4. Genesis Model has one hard readiness gate. R8.5.t is a recorded independent decision, not a slogan. It cannot pass until GenesisCode v1 Core and Genesis Algorithms are stable, GenesisBench has an independently useful trust release and a proven training/evaluation firewall, MF Foundation is reproduced, the Theseus/ASI/Foundry transfer pack is reviewed, legal/privacy/security/compute/release feasibility is closed, and a simplest-adequate open baseline strategy is predeclared. Corpus construction, teacher selection, tokenizer/base selection, training, distillation, packaging, and integration remain non-start-ready before that decision. Generic model-provider support and typed intent may continue because they are language capabilities, not Genesis Model implementation.
  5. The current slice is bounded and generated. One repository-changing production task may be active. A separately custodied immutable benchmark campaign may coexist only with disjoint files, roles, and authorities. The existing roadmap policy orders the immediate slice as finish/publish R0.4.j, close P1.5 through R2.2.f, then resume the M1-A/M1-B and runtime sequence. Merely start-ready work is not permission to widen WIP. python3 scripts/lib/roadmap_execution_manifest.py --slice reports the selected task, blockers, owner paths, guards, rollback, and queued successors from existing authorities.

This order supersedes older prose that allowed pre-gate model architecture/tokenizer experiments or made every target pack a v1 Core blocker. It does not discard the universal one-language goal: it converts that goal into independently shippable, evidence-bearing packs so useful releases arrive before the complete portfolio does.

3.13 Executable release-lane isolation audit delta: 2026-08-03

A dependency audit of the generated execution manifest found that the prose-level Core/pack split was not yet mechanically complete. Historical edges through R8.4.e, R8.5.s, R8.5.r, and AB-7 could still make GenesisCode Core, GenesisBench Trust, Foundry calibration, or a profile-bound Genesis Model wait for the universal eighteen-archetype portfolio. This revision closes that gap:

  1. GenesisCode v1 Core depends on R8.2.a-b and the Core human/documentation surface, never R8.2.c-r or R8.3. Pack documentation, device soak, and flagship operation remain mandatory for the pack that claims them.
  2. GenesisBench Trust depends on a compatible frozen GenesisCode profile and its own independent operators, validity, custody, governance, and adoption evidence. It does not require universal product maturity or a Genesis Model.
  3. At this audit MF depended on signed v1 Core interfaces and F2.a-q only. Section 3.25 replaces that boundary with F2.a-r so the initial independently governed algorithm-library release is included; F2.s-z remains separately gated expansion and cannot be pulled into Model readiness by a scalar F2 completion claim.
  4. Genesis Model readiness depends on stable Core, GenesisBench Trust, MF, the firewall, and accepted transfer evidence. Model release evaluation covers its declared compatible profiles; a coupled language/model proposal gets the full four-cell design, while a model-only proposal uses an explicit frozen-language two-cell control rather than inventing a language change.
  5. The execution policy’s closed release_lane_contracts table names each release root, mandatory ancestors, forbidden workstreams/task families, and narrow exceptions. The execution-manifest self-test injects representative cross-lane leaks and treats acceptance as a release-policy violation. Future prose, milestone, or policy edits must preserve these negative dependencies, not merely avoid graph cycles.
  6. Shared service, build, test, and supply-chain authorities are profile-parametric rather than aggregate portfolio gates. Core requires the service/data/operations subset it actually demonstrates, but not the data/ML library. Product-generator adversarial testing and product/hardware supply-chain qualification provide reusable harnesses; each platform pack must bind and pass its own applicable cases, while Core, GenesisBench, Foundry, and Genesis Model cannot inherit unfinished pack cases through a scalar workstream dependency.

This correction removes accidental blockers only. It does not promote an unfinished capability, reduce any claimed profile’s evidence burden, or allow Core, Bench, Foundry, Model, or a platform pack to borrow another lane’s evidence.

3.14 CI control-plane and delivery-liveness audit delta: 2026-08-04

The seventh independent review identified a release-control failure that is accepted after direct verification against GitHub Actions history and repository runner inventory. At the audit boundary, the latest successful main CI run was 29664738972 on 2026-07-18. The subsequent history contained fifteen failed and sixteen cancelled default-branch runs, while the repository had zero registered self-hosted runners. PR #13 repaired the deterministic nested write-skill evidence-binding failure and passed all 97 protected aggregate steps before merge, but that does not by itself restore default-branch or scheduled-full observability. Local and PR green remain necessary observations; neither substitutes for a terminal run of the exact default-branch profile whose claims are being published.

This review makes five bounded corrections:

  1. CI liveness is part of R0.4.j, not external operations trivia. P1.6 records the active control-plane defect. It refines the existing M0/M1-A/M1-B release-profile task, consumes its operational-hardening capacity, and adds no product, platform, benchmark, gate family, or release lane.
  2. Concurrency policy must preserve authoritative work. Pull-request runs may cancel superseded runs for the same PR. Push, schedule, and operator-dispatched runs use event/profile-aware groups or an explicit bounded FIFO queue so a newer run cannot silently replace the sole pending evidence for an older immutable revision. Every cancellation has a typed supersession or operator disposition; scheduled full evidence cannot share a starvation domain with ordinary pushes.
  3. Unavailable self-hosted labels are discovered before dispatch. A GitHub-hosted preflight checks the exact required label conjunction and online/busy state, binds the inventory observation, and prevents an impossible job from entering the 24-hour self-hosted queue. Absence becomes a bounded infrastructure-failure or unsupported-profile result. It may be non-blocking only for a release profile that explicitly makes no claim for that optional platform pack; it is a hard failure for any profile that claims authentic device coverage. A deterministic surrogate can prove harness semantics but never satisfy authentic GPU evidence.
  4. An independent watchdog observes the observer. A disjoint workflow and concurrency authority checks that the latest pushed main revision receives its required aggregate disposition within two hours, every scheduled or dispatched full run terminates within sixty minutes, a successful full run is no more than forty-eight hours stale while the project is active, and no undispositioned failure/cancellation or impossible runner queue persists across the next observation. Static checks prove the watchdog cannot share the failing workflow’s concurrency group, and fixture tests reject stale-success, cancelled-only, missing-schedule, wrong-head, impossible-label, and self-reporting-success mutations.
  5. The remaining recommendations already have owners. R1.4.o already requires immutable Qwen3 4B/8B and Luna campaigns plus the benchmark card, methods, chronology, invalid dispositions, and first-results report. R1.4.q already requires at least 45 custody-verified real private lineages. Sections 3.10, 7, and 14 already freeze milestone growth and WIP. Those requirements remain next in sequence and are not duplicated or displaced by another planning layer.

R0.4.j closes only after the exact reviewed main revision passes the complete standard and release-full profiles under these controls, the watchdog independently recognizes the evidence, the runner-absent negative path terminates within its bound, and the historical cancelled/failed sequence has an append-only disposition. This is a delivery-integrity requirement, not permission to treat optional unavailable platforms as implemented.

3.15 Release-DAG and milestone-flow audit delta: 2026-08-05

The latest independent review verified that the prior CI corrections work and independently reproduced the remaining release-full failure. Both hosted v0.1 pair workers completed their cold child in approximately 42.7 minutes under the unchanged 45-minute worker bound, then gave the warm child only the remainder of one shared 50-minute pair envelope: 430674ms and 442492ms. This is the expected failure of an impossible serial contract, not evidence for raising GB-4 or the watchdog SLO. The in-flight v0.2 policy, schema, validator, and class-selective health runner correct the methodology by using independent matched cold and deterministically prewarmed cohorts rather than calling observations from unrelated hosted machines pairs. They are necessary but not sufficient: hosted CI still runs v0.1 until the worker reports, read-only aggregate, workflow fan-out, failure-tail retention, and watchdog recognition are promoted as one reviewed migration.

This review changes execution in four bounded ways:

  1. The next R0.4.j transaction is the hosted v0.2 cutover. Replace release_full_measurement_pair rather than tuning its constants. Run three independent cold and three independent warm cache-sensitive workers with fresh 45-minute bounds, execute invariant work once, execute three isolated stress/performance workers, derive warm precondition proof outside measured time, and give only the read-only aggregate authority to decide completeness. Permanent controls reject any surviving serial-pair route, fewer or even-sized cohorts, producer-authored verdicts, cross-run fan-out, or a workflow whose worst-case critical path exceeds sixty minutes.
  2. R0.4.j retains one exact exit. Policy or local runner completion cannot close it. The exact reviewed main revision must produce a complete v0.2 bundle, bounded diagnostic tails for every failed node, named target dispositions, append-only disposition of the v0.1 counterexample runs, a successful aggregate, and independent watchdog recognition within the existing freshness window.
  3. GenesisBench product proof moves directly behind R0.4.j. After the exact-main release cutover, the next repository-changing frontier is R1.4.o’s already-specified Qwen3 4B nine-class campaign: freeze and publish its predeclaration before inference, execute every cell once, preserve invalids and failures, validate/replay/rescore independently, and publish the result. R2.2.f and R1.5 remain start-ready M1-A work, but no longer serialize this dependency-free M1-B adoption proof.
  4. Long production work cannot silently starve a disjoint authorized lane. If one production task remains active for more than 72 hours or spans two independent review checkpoints, the execution slice must record a dated disposition for every start-ready parallel lane: active, externally blocked with the missing authority/operator/resource, or explicitly deferred with milestone impact. One immutable benchmark custody/execution campaign or private-lineage commissioning campaign may proceed only under the existing disjoint-file, disjoint-authority, frozen-predeclaration, no-target-leakage conditions. This rule exposes avoidable serialization; it does not increase repository-changing WIP, compel model access before custody freezes, or let parallel work delay the selected production frontier.

Parallel-lane disposition 2026-08-05: R1.4.q commissioning is externally blocked from activation because no disjoint role-separated independent author/admission/custody operator set and storage authority is recorded for this tranche. Repository-changing corpus/protocol work also overlaps the current generated-authority closure and cannot become a second production transaction. No target-model access is authorized. An out-of-repository no-target-access commissioning lane may activate immediately once those named authorities, conflict declarations, storage boundaries, and frozen admission contract are recorded; otherwise the blocker remains visible in the next execution review.

3.16 GenesisBench stratification and GenesisChallenge audit delta: 2026-08-06

The benchmark direction is accepted with one structural correction: calibrated evaluation and open-ended project improvement answer different questions and must not share a rank. This revision refines R1.4.o/q/p and R8.5.a-e,s, adds the post-M5-B R8.5.u-v contest work, and consumes their existing benchmark/adoption capacity. It adds no product, platform, Foundry, or model dependency to the release DAG: M1-B/M5-B gain explicit acceptance detail inside their existing benchmark tasks, while GenesisChallenge blocks no release. The frozen v0.1 public campaigns, task IDs, conditions, outcomes, and failed expansion predicates remain immutable commissioning anchors; no new label or profile is applied retroactively.

  1. Difficulty, context, and cost become orthogonal axes. Every future scored task declares exactly one empirically maintained easy, medium, or hard difficulty band; one existing small, medium, or large context condition where applicable; one run profile; one task class; one capability/product surface; one benchmark track; and the exact language, epoch, scaffold, adapter, and hardware cohort. Context length is never used as a proxy for difficulty, and no aggregate hides a missing required stratum.
  2. Difficulty is calibrated, not editorialized. Authors predeclare a structural rationale covering semantic depth, repository scope, ambiguity, authority minimization, maintenance horizon, target authenticity, and resource pressure. Independent non-target pilots across at least two materially different model families then calibrate the bands with a predeclared item-response model or a justified nonparametric equivalent. Labels and ranking weights are immutable within an epoch; low-discrimination, differential-item-functioning, leaked, or unsound items receive append-only dispositions rather than post-hoc relabeling.
  3. Run profiles answer different operational questions. smoke is a deterministic non-ranking conformance check. fast is a fixed, content-addressed, precommitted stratified subsample for inexpensive directional comparison. standard is the default statistically powered comparison over every required class and difficulty stratum. full is the complete declared compatible capability/product matrix, including target-authentic cases only where the frozen GenesisCode profile has the required L4 support. Fast results live in their own cohort, publish their sampling frame and wide uncertainty, and cannot satisfy Trust-release, universal-coverage, or strongest-model claims.
  4. Fast never means hand-picked. Each epoch publishes at least two deterministic fast forms derived before target-model access from lineage-level strata, with no condition siblings split across train/evaluation or counted as independent tasks. Selection optimizes declared coverage, not prior model scores; every omitted class or capability is visible. A fast result names the form identity and is comparable only with the same form, track, profile, scaffold, adapter, attempt policy, and compatible hardware class.
  5. Capability coverage follows the language, not marketing. The benchmark coverage ledger maps every released syntax/semantic form, effect/capability, package/build/deploy operation, agent workflow, maintenance operation, and L4 product/target pack to at least one eligible task or an explicit unbenchmarked/unsupported disposition. Easy cases isolate primitives and local reasoning; medium cases compose repository-scale libraries, CLI/service behavior, policy, packaging, and repair; hard cases require multi-subsystem construction, migration, optimization, authentic deployment, or maintenance without weakening obligations.
  6. Translation becomes behavioral reconstruction. Once a destination profile is L4, GenesisBench may commission license-clean source migration tasks that expose a foreign-language application and spec reconstruction tasks that expose behavior, assets, protocols, and acceptance tests without requiring source. Solvers author only GenesisCode-owned source and manifests. Acceptance uses behavioral, property, metamorphic, security, accessibility, resource, deployment, and seeded-maintenance tests rather than reference-patch equality; generated foreign target artifacts are disposable and non-editable. An unsupported target is a typed benchmark exclusion, not a model failure.
  7. Epochs rise without erasing the tide line. Public and private anchor lineages preserve longitudinal measurement; saturated items remain available as diagnostics and historical anchors while ranking weight rotates prospectively under BQ-5. New hard material is precommitted before target access, and difficulty advancement preserves overlapping anchors and publishes equating uncertainty. Tasks are retired for ranking only through governance; prior bundles, labels, weights, failures, and scores remain replayable.
  8. GenesisChallenge is a separate contribution contest. Its challenge record freezes the exact repository revision and reproduced defect or measured baseline; editable and protected scope; acceptance tests and hidden adversarial controls; capability, resource, security, and compatibility bounds; verifier identities; expiry; prize or credit terms; and destination-owned promotion path. Runtime optimization, correctness/security repair, diagnostics/agent tooling, library/platform capability, proof/fuzz/conformance, and benchmark-validity work are eligible. Challenge status is append-only (proposed, admitted, open, claimed, verified, accepted, merged, retired, invalidated, or disputed), and proposer, solver, model/provider, scorer, custodian, contest operator, and destination approver conflicts are explicit. A solver may receive signed attribution for a verified delta, but cannot admit, score, merge, promote, or rewrite its own challenge; evaluator changes are meta-challenges that must shadow-replay history and cannot verify themselves.
  9. The optimization target is a Pareto frontier, not one universal score. GenesisChallenge reports accepted correctness, safety, capability, latency, memory, energy, context, authority, proof burden, maintainability, and human-review deltas separately, with regressions and failed attempts retained. GenesisBench continues to compare acquisition and trustworthy construction under frozen contracts. Neither surface may claim that GenesisCode is universally best from a scalar composite or let benchmark-specific shortcuts enter production semantics.

3.17 Model-neutral benchmark portability audit delta: 2026-08-06

The concern is accepted. GenesisBench v0.1 already has model-agnostic semantic scoring, five closed adapter classes, fixed-scaffold cohorting, and typed harness invalids, but those facts do not yet prove that an arbitrary capable model can enter without imitating Codex CLI, OpenAI Responses, MLX, Qwen chat conventions, or a provider-specific reasoning control. R1.4.r now closes that portability proof before M1-B. It refines R1.4.e/f/g/k/l/n and the publication/governance work in R1.4.o/p; it does not rewrite completed evidence or retroactively reclassify any frozen campaign.

  1. One transport-neutral task authority feeds every model. The canonical benchmark input is ordered UTF-8 content blocks, declared roles/trust labels, immutable file/artifact attachments, generic tool schemas where the condition permits them, budgets, and an acceptance manifest. It contains no provider wire object, tokenizer ID, hidden system message, SDK type, special token, or model-family instruction. Adapters perform typed lossless transport mapping only and cannot summarize, improve, reorder, omit, or semantically rewrite task content.
  2. A minimum-common-denominator profile admits text generators. Every eligible non-interactive task has an artifact-text condition that requires only bounded text input and text output; the harness accepts versioned complete-file maps, unified patches, or declared final artifacts and performs parsing, workspace application, tests, and scoring outside the model. Bounded diagnostic-repair turns append canonical text observations without requiring native function calling. Models with chat, streaming, JSON mode, or tool calling can use those features only in separately keyed interaction conditions.
  3. Generic tools do not imply one native encoding. A canonical-tools condition exposes the same closed tool names, schemas, results, order, and authority through either provider-native tool calls or a canonical JSON-text emulation. Cross-encoding scripted controls must produce identical normalized tool requests, workspace transitions, candidate artifacts, resource charges, and scores. Provider-owned search, code execution, subagents, memory, retrieval, or undeclared tools are disabled in fixed-scaffold tracks and disclosed only in Open Agent cohorts.
  4. Capabilities are negotiated before task delivery. Each adapter declares support for plain completion, chat roles, multi-turn state, system messages, streaming, native tools, JSON constraints, seed, sampling controls, stop sequences, token accounting, context limit, image/audio inputs where applicable, and cancellation. Required-feature absence yields unsupported-interaction or unsupported-context, not a semantic model failure; it remains a visible missing cell and prevents complete-cohort eligibility rather than disappearing from the denominator. Optional unsupported controls are omitted and cohort-keyed rather than silently emulated; a provider rejection, timeout before valid delivery, malformed wire response, or adapter extraction defect remains an invalid infrastructure/harness outcome.
  5. Context is tokenizer-neutral at authority boundaries. Task and card budgets are canonical UTF-8 bytes plus structural node/file counts. Every adapter records the exact delivered bytes and, when available, provider/model token counts and tokenizer identity as non-authoritative resource facts. No canonical task is truncated, reordered, or selected using one model’s tokenizer. If exact content cannot fit a declared model context window, the complete condition is reported unsupported and a separately predeclared smaller condition may run; hidden per-model compression or prompt tuning creates a new scaffold cohort.
  6. Candidate extraction is plural and semantics-first. The closed extractor accepts the declared neutral artifact forms, preserves raw output, and never repairs generated code. Format variance that still yields the same declared workspace is normalization; absence of an extractable candidate produces a typed protocol outcome, not an invented semantic score. Once a candidate exists, all models face identical parser, type/effect, obligation, policy, test, resource, replay, and maintenance authorities; style, prose, markdown fences, identifier taste, and reference-patch resemblance do not affect quality.
  7. Failure attribution precedes model judgment. Every attempt terminates in exactly one independently rederivable class: verified, semantic-failure, model-protocol-failure, adapter-failure, scaffold-failure, provider-failure, infrastructure-failure, resource-limit, unsupported-interaction, unsupported-context, or abstained. Only a validly delivered request, conformant response capture, successful neutral extraction boundary, and intact scoring workspace can produce semantic-failure. Ambiguous ownership fails toward invalid/non-claiming status and retains all evidence.
  8. Portability is tested by metamorphic adapters, not brand count. One scripted semantic model runs through completion-only, chat, streaming/non-streaming, native-tool, JSON-text-tool, local direct-runtime, HTTP, and JSON-stdio fixtures; normalized requests, candidate workspaces, scores, and resource accounting must agree where capabilities overlap. Mutants that inject instructions, alter roles, drop content, use a different tokenizer to truncate, add retries, enable provider tools, repair output, leak secrets, or blame the model for harness failure must be rejected. At least three independently implemented real interface stacks and three materially different model families reproduce the common artifact-text matrix before GenesisBench Trust.
  9. Bring-your-own-model is a product surface. A self-hosting user supplies one declarative adapter manifest plus either a bounded JSON-stdio command or a loopback/remote endpoint. genesis bench adapter inspect, conformance, and dry-run validate capability negotiation, secret isolation, request/response closure, timeout/cancellation/reap, transcript retention, and model-free replay before any scored access. Public documentation provides completion-only, generic chat, native-tool, JSON-text-tool, local-runtime, and Open Agent templates, with no GenesisCode or provider SDK changes required to add a conforming model.

3.18 Self-host-first activation decision: 2026-08-06

Historical note: sections 3.24 and 3.25 supersede items 3, 4, and 6 after the unchanged R9.4.f Core handoff; they are retained here to preserve the decision trail.

The concern is accepted. GenesisBench protocol and fixture work has exposed valuable defects, but the program must not optimize an unstable language around benchmark campaigns or let benchmark publication compete with semantic ownership. The stable task and milestone IDs remain unchanged so completed history and evidence remain addressable; activation order, not historical identity, is changed here and in policies/roadmap_execution_v0.1.json. This decision supersedes every earlier execution-priority or parallel-campaign statement in sections 3.5 through 3.15 while preserving those sections as dated historical rationale and evidence.

  1. GenesisCode v1 Self-Hosted Core is the sole production critical path. The activation gate is exact reviewed R9.4.f evidence over the finite Core scope, including all R4 exit criteria: every production semantic decision above the specified stage0 reaches H2 GenesisCode authority; stage2/stage3 reach an H3 hermetic cross-host fixpoint; diverse double compilation and independent H4 conformance verification pass; legacy semantic authorities are unreachable; and the released language, runtime, compiler, formatter, type/effect checker, optimizer, patch/obligation/policy logic, package/registry/build orchestration, CLI/agent interface, documentation generation, native/WASI CLI, and service workflows are usable without GenesisBench, Foundry, a model, or mutable research state.
  2. Completed benchmark foundations are preserved, not expanded or erased. Existing public fixtures, protocol/scoring authorities, adapters, frozen attempts, invalids, and historical reports remain append-only. Before R9.4.f, benchmark code may receive only security/correctness maintenance required by GenesisCode gates, model-free conformance tests, regeneration needed to keep completed authorities reproducible, and the non-scoring interface canary defined in section 3.19. No benchmark-task target inference, private-lineage commissioning, campaign, rank, benchmark governance release, adoption claim, GenesisChallenge operation, or benchmark-driven language tuning begins.
  3. Superseded downstream activation. This decision originally activated GenesisBench directly after R9.4.f. Section 3.25 now requires F2.r/MF and the exact Genesis Algorithms release first; every later task, scaffold, adapter, context, scorer, and result binds both language and library identities.
  4. Superseded Foundry ordering. This decision originally placed Foundry after GenesisBench Trust. Sections 3.24-3.25 reverse that edge safely: independent F2.a-r Foundation precedes Bench, while only cross-ecosystem F2.y waits for compatible Bench Trust.
  5. Internal tests are not GenesisBench releases. Compiler differentials, conformance corpora, mutation controls, agent workflow tests, and authentic product acceptance remain mandatory throughout the self-host path. They establish language correctness and usability but carry no model-comparison, held-out, calibration, rank, adoption, or benchmark-release claim.
  6. The freeze remains fail-closed and machine-enforced. The unchanged pre-Core rule keeps Benchmark and Foundry implementation out before R9.4.f. Section 3.25 now makes F2.r the required ancestor of Preview, Trust, Challenge, and model-transfer work, while release-lane controls reject Bench/Model ancestry inside Foundry Foundation and reject F2.s-z expansion ancestry in downstream releases.

3.19 Round-10 delivery, portability, and self-host checkpoint correction: 2026-08-06

The review’s load-bearing portability concern is accepted, while its repository-delivery snapshot is superseded by direct evidence. Exact head 0914f08561b6fe904acdcd214f4ba18cd639b3a6 is remotely backed up on the existing review branch, the documentation workflow passed that head, and the standard workflow is still running. Those facts protect the work but do not close R0.4.j: only exact reviewed main release-full evidence and independent watchdog recognition can do that. No task is checked from this local or PR observation.

  1. Retire interface risk without reopening GenesisBench. One read-only, non-scoring portability canary may run for each transport-neutral authority revision after its model-free fixtures pass. It must use an already inventoried, digest-pinned local model from a materially different family and a direct-runtime or generic JSON-stdio stack that does not imitate Codex CLI or a hosted provider wire protocol; select the strongest predeclared host-fit model, currently Qwen3 8B, unless a score-blind resource preflight rejects it. The canary receives only fixed protocol content that tests lossless ordered blocks, candidate extraction, capability negotiation, timeout/cancel/reap, and transcript closure. It receives no GenesisBench task, private payload, scorer, answer, repair loop, tool authority, or language-change signal. Its only result is a typed pass, fail, unsupported, or infrastructure-failure interface-qualification record with exact model/runtime/adapter/input/output identities. It creates no attempt, score, cohort, rank, benchmark result, adoption claim, roadmap completion, or publication entitlement.
  2. Use four demonstrable self-host checkpoints inside existing tasks. SH-A closes R4.1 with the minimal stage0 contract, complete semantic-ownership ledger, structural boundary gate, and one selected migration slice. SH-B closes R4.2.a with the GenesisCode frontend/canonicalizer at H2 under strict no fallback on native and WASI, independently reproducing source-to-CoreForm bytes, hashes, spans, diagnostics, and malformed-input behavior. SH-C closes R4.2.b-e with all non-codegen semantic authorities at H2 and proves that deleting or disabling their legacy Rust routes cannot change the declared command matrix. SH-D closes R4.2.f-g plus R4.4.a-b with one hermetic reference-host stage2/stage3 bootstrap rehearsal before the cross-host claim; R4.4.c-e and R4.5 then establish release H3/H4. Each checkpoint publishes a runnable artifact, exact no-fallback reachability evidence, negative controls, residual H-levels, and the next smallest blocking edge. None is a release milestone or permission to weaken final cross-host, DDC, or independent-conformance acceptance.
  3. Migrate vertically instead of waiting for every execution tier. R4.2 frontend through package/policy migration begins after R4.1 and the relevant R5.1/R5.3 contracts freeze; only R4.2.f code generation waits for R3.1-R3.3. R4.4.a-b stage definitions and hermetic harness work may proceed after R4.1, while R4.4.c cross-host fixpoint remains blocked on the complete R4.2/R4.3 authority cutover. R4.5 conformance specifications may also begin after R4.1, but independent implementation acceptance remains blocked on the production authority switch. This exposes early end-to-end evidence without claiming that a routed wrapper or single-host rehearsal is H3.
  4. Freeze conceptual scope through Core release. Until R9.4.f, the named concepts remain GenesisCode, GenesisBench, GenesisChallenge, Genesis Foundry, and Genesis Model. Do not add a product, program, benchmark track, research lane, milestone family, task family, or governed gate family. A reproduced P0/P1 correction may refine an existing task or replace an equivalent concept one-for-one with a recorded deletion, dependency impact, and acceptance delta; it may not increase WIP or release ancestry. Ideas outside that bound go to non-authoritative notes without implementation or roadmap tasks. The execution policy exposes this freeze as machine-readable review authority.
  5. Finish the release transaction before opening another repository-changing lane. Hosted standard CI on the current head must terminate successfully, the coherent fixes must reach reviewed main, release-full must execute the complete v0.2 DAG on that exact revision within GB-4, and the disjoint watchdog must recognize fresh success. Infrastructure failures receive typed append-only dispositions and a retry of the unchanged revision; they are never converted into code passes. The interface canary is read-only and separately custodied, so it cannot delay, mutate, or authorize this transaction.

3.20 Round-11 closure and dependency-complete execution correction: 2026-08-08

The review correctly recognizes the first authentic release-full success and the remaining delivery concentration, but its claim that seventeen R1 tasks are immediately available is not supported by the generated graph. R1.6.e is load-bearing agent safety work, yet it depends on R1.3.f and R1.6.d; R1.3.f in turn depends on P1.5/R2.2.f lifecycle closure. Repository-changing WIP remains one. This correction closes the achieved release task, repairs the machine route that could otherwise skip unfinished sequential work, and records the exact order without adding a task, milestone, product, benchmark lane, or gate family.

  1. R0.4.j is closed only by exact evidence. Exact main revision da9359fd35221421563ea77b7dbe17e0a33d0e12 passed standard run 31249172621, release-full run 31249173667, aggregate 68fab0fa55bdf560a81f45b35168bb55123c9f1370b56c01c24ddb357fdd4345, and independent watchdog run 31250800213. The append-only chronology retains all failures and cancellations through that success, and the resulting closure remains productReleaseQualified=false with every deferred target explicitly unsupported-product.
  2. The execution frontier is dependency-complete rather than checkbox-shaped. A completed anchor in a sequential workstream now resolves to that workstream’s first unfinished task. After R2.2.f, the machine route advances through R1.3.f, the complete R1.5 authoring-skill sequence, the complete R1.6 collaboration sequence, the complete R1.7 typed-intent sequence, R2.1.h, the complete R2.3 startup/incremental sequence, and then the complete R4.1 SH-A sequence. A focused task may not be skipped merely because its first workstream task is already complete.
  3. Instruction/content separation is prioritized at its real dependency boundary. R1.6.e remains mandatory before typed English intent and product-agent acceptance, with provenance/trust labels and prompt-injection, confused-deputy, and indirect-tool-instruction controls. It does not preempt R1.6.a-d or the daemon lifecycle evidence it consumes, and it cannot be declared parallel-ready to manufacture velocity.
  4. SH-A starts after the finite agent/runtime frontier, not after benchmark work. R4.1 remains the first self-host authority checkpoint and GenesisCode v1 Self-Hosted Core remains the sole production release path. The benchmark freeze in section 3.18 is unchanged. The transport-neutral portability canary remains separately custodied, read-only, non-scoring, and non-blocking; its absence is visible interface risk but never authority to reopen a campaign or delay the selected production task.
  5. Every handoff must show progress without rewarding shallow closure. A work session either closes the selected task with its required dated evidence or records the exact unmet acceptance edge, reproduced counterexample or external authority needed, completed evidence, and next executable action. The generated focus cannot silently change to an easier task. This is transaction accountability, not permission to split semantic work into artificial checkboxes or treat local mutable output as completion.
  6. Working instructions must match current authority. Any AGENTS, skill, status, or handoff text that routes work into a pre-R9.4.f benchmark campaign or around the generated frontier is stale and fails review. GitHub availability affects publication evidence only; local implementation and verification continue on main whenever the selected task does not require the external service.

3.21 Round-12 lifecycle and evidence-control correction: 2026-08-08

The review accepts the lifecycle, watchdog, generated-authority, and timing criticisms that have executable counterexamples. It rejects the recommendation to manufacture velocity through seventeen parallel R1 tasks or to preempt the dependency-complete frontier with SH-A: only R1.5.c is independently start-ready, R1.6.e remains blocked by R1.3.f and R1.6.d, and repository-changing WIP remains one. This correction adds no product, task, milestone, benchmark lane, gate family, or release ancestor.

  1. Historical success is preserved but cannot satisfy a later regression. The exact R0.4.i and R0.4.k evidence remains the durable record of the atomicity and transitive-input contracts established on those reviewed revisions; reopening them would falsely invalidate completed dependents. P1.7 records a later regression introduced in the operational release path, and reopened R0.4.j owns its correction plus a complete rerun of the R0.4.i/R0.4.k guards. R0.4.j also owns P1.6 because its watchdog does not independently classify the promised standard dispatch. M0 and the R0 completion count remain open until both current defects pass on reviewed main; a generated count or prior green watchdog cannot overrule the open defect ledger and task text.
  2. Generated publication has one real resource owner. The canonical updater must run under a bounded aggregate observer that measures complete transaction wall time and generated-disk growth, owns the event channel, rejects network attempts, terminates and reaps every child process group, and emits a typed failure. Parallel nested checks retain their own hard duration and deny-network enforcement; only shared disk attribution may move to the aggregate, and only through an unforgeable supervisor path rather than a caller-selected environment switch. Node timeoutSeconds and diskMiB declarations become enforced limits, not schema-only metadata. Negative controls exceed each bound, omit the aggregate owner, inject a caller bypass, and prove no partial promotion.
  3. Standard means standard. The disjoint watchdog must classify the separately dispatched exact-main standard profile by event, title/profile, head revision, attempt, terminal state, and canonical-history membership. A push-fast success is not a standard disposition. Missing, overdue, cancelled, failed, stale, wrong-head, and superseded-only standard histories fail independently of full-run freshness; the two-hour standard bound and append-only incident disposition remain distinct from the sixty-minute full termination and forty-eight-hour successful-full freshness bounds.
  4. Cleanup errors cross the same boundary as operation errors. Persistent bridge malformed-response and disconnect paths, session eviction, owner shutdown, and daemon teardown may not discard a failed stop/reap result. Public operations return a typed cleanup/reap error that preserves the initiating failure as context; an explicit fallible close/drain path reports aggregate teardown failure, while Drop is only a bounded last-resort safety net. Mandatory fault injection reaches these public paths and proves bounded join/reap, no detached worker, no surviving descendant, deterministic diagnostics, and recovery of the next healthy request on every supported native host.
  5. Timing gates combine hard containment with statistical calibration. A semantic pass never converts a hard timing failure into success, but one noisy hosted sample cannot authorize a budget change. Local warm, local clean/fallback, and hosted cold/shared-runner classes are measured separately with exact toolchain, host class, cache precondition, workload identity, and competing-lane state. After five discarded warmups, at least thirty retained conformant samples establish median, nearest-rank p95, MAD, and a distribution-free confidence interval; the hard ceiling is derived from the declared host-class distribution and documented headroom, while rolling median/p95 trend alarms detect sustained regression. Budget changes preserve the failed samples, cannot relabel cold observations as warm, and ratchet downward when repeated conformant evidence permits.
  6. The corrected frontier remains sequential. First close P1.7 and P1.6 through reopened R0.4.j, ordering aggregate resource ownership before the hosted evidence rerun; then close P1.5 through R2.2.f. Only after exact reviewed evidence closes those boundaries does the existing R1.3, R1.5, R1.6, R1.7, R2.1, R2.3, and SH-A order resume. The portability canary remains optional, read-only, non-scoring, separately custodied, and non-authoritative; GenesisBench, Foundry, and Genesis Model activation remain frozen exactly as before.

P1.7 closed 2026-08-08 on exact main revision d07ff48e909fea43ec405b7160901b5bbe0b62ad. The generated-authority owner now bounds the complete transaction and each generation node, preserves nested duration/network enforcement, delegates only parallel disk attribution through an inherited append-only descriptor, and kills/reaps process groups on violation. The canonical update reached changed=0; the complete R0.4.i/R0.4.k and R0.4.j guard union passed; scripts/test_changed_fast.sh passed in 378,681 ms under its 720,000 ms/3-GiB envelope; and hosted push-fast run 31282158056 plus documentation run 31282158060 passed on that exact revision. Twenty telemetry controls and twenty-six generated-authority controls reject duration, disk, network, caller-bypass, missing-owner, concurrent-child, timeout, and partial-publication faults. This closes P1.7 only: P1.6, R0.4.j, M0, and release qualification remain open.

P1.6 closed 2026-08-08 on exact main revision 63db7e3f3a00da6b582aa0464ad1e12d9df2e85d. Push-fast run 31284538133, standard run 31284542152, release-full run 31284543591, and documentation run 31284538134 passed on that revision. The independently verified v0.2 aggregate has identity 6c5c4e3da4ce4ee04fd744caaa7d088746aa814463061021d702a8a4244733b9, authenticates all three cold, three prewarmed, one invariant, and three stress workers, derives its verdict read-only in 1,727 ms, and preserves productReleaseQualified=false; Android, edge, iOS, and service runtime remain explicit unsupported-product dispositions. Post-pair watchdog run 31287266902 independently observed exact-head standard 31284542152 and full 31284543591, reported exactHeadStandardFullPair=complete, emitted zero violations, and preserved releaseQualified=false. The append-only standard chronology retains the failed 31269647811 counterexample and all seven dispositions through the closing success. This closes P1.6 only: R0.4.j remains open because section 3.21 still requires statistically calibrated local-warm, local-clean/fallback, and hosted-cold/shared-runner timing classes; P1.5, M0, and product release qualification also remain open.

3.22 Round-13 readiness, self-host priority, and agent-operated calibration correction: 2026-08-09

The review’s corrected dependency analysis has teeth: only five of 300 open tasks were graph-ready on the reviewed snapshot, and R4.1.a was one of them. Sections 3.20 and 3.21 were right to reject seventeen parallel R1 tasks, but wrong to describe SH-A as dependency-deferred. Its deferral is solely the one-repository-changing-WIP rule. This correction makes that scheduling choice explicit, moves the complete R4.1 truth/boundary package directly behind lifecycle closure, and adds a global derived readiness view without adding a task, milestone, product, gate family, or authority.

  1. Graph readiness and selection are separate facts. python3 scripts/lib/roadmap_execution_manifest.py --ready reports every open task whose declared prerequisites are complete, then identifies which of those tasks the WIP-limited policy selected. The report is read-only, recomputed from the roadmap and policy, closed to unknown fields by implementation tests, and cannot authorize start, completion, promotion, capability, signing, or release. --slice remains the work selector; a ready-but-deprioritized task cannot widen WIP.
  2. SH-A is WIP-deferred, not dependency-blocked. Finish R0.4.j, close P1.5 through R2.2.f using the retained cross-host lifecycle evidence, then execute the complete R4.1 SH-A sequence before the longer R1.3, R1.5, R1.6, R1.7, R2.1, and R2.3 sequence. SH-A publishes only the minimal stage0 contract, semantic-ownership truth, structural boundary, and selected migration slice already declared in R4.1. It does not claim H2/H3/H4, broaden stage0, start R4.2 prematurely, or activate GenesisBench, Foundry, or Genesis Model work.
  3. The WIP limit remains one. The reordered frontier is a priority decision among graph-ready work, not a parallel lane. Every handoff continues to name the selected task, completed evidence, exact unmet edge, and next executable action. GitHub unavailability blocks only evidence that specifically requires it; it does not permit substitution of an easier ready task.
  4. Local timing represents the agent-operated product without inventing exclusivity. The original 5% local preflight repeatedly rejected an otherwise idle Apple M1 whose only application workload was the Codex operator itself: five observed one-minute load averages of 1.55-1.86 on eight logical CPUs corresponded to roughly 19-23%, with zero competing Cargo, Rust, Genesis, nextest, or Quarto process. A longer post-validation observation then stabilized near 40% with the same zero competing-build result while Codex, WindowServer, Spotlight, and macOS media analysis were the only material resident consumers, proving that a 30% replacement still encoded an unavailable idle-host assumption. After the complete changed-file profile passed, two further exact-clean-main admissions correctly found no competing build but rejected at the 50% boundary while contemporaneous one-minute load averages were 4.12 and 5.32 on eight logical CPUs, roughly 52-67%, again dominated by Codex and macOS indexing/media services. R0.4.j therefore keeps the general 5% unattended reference-lab control unchanged but gives its local agent-loop timing classes a separate policy-bound preflight: five one-second-spaced samples, maximum statistic, 75% logical-CPU ceiling, AC power, low-power mode disabled, nominal thermals, exact clean main, zero external competing build process, and hard process-group timeout/kill/reap. Exclusivity is not inferred from one process-name snapshot: argument-aware checks run before launch and once per second through termination, direct and wrapped Cargo/Rust/Genesis/nextest/Quarto commands are recognized, the complete measured-workload descendant closure is excluded even across nested process groups, and any matching external process becomes a retained competing-lane hard failure. Every observation records the exact applied ceiling and maximum observed competitor count; semantic passes require zero. This is a realistic class definition, not permission to relabel a contended run or weaken hosted/release containment.
  5. Calibration failures remain evidence. The exact-main hosted campaign retains the cancelled startup observation as an interrupted hard failure with reaped cleanup rather than excluding it from review. Five discarded semantic-pass warmups and thirty retained passes are still required independently for each class; failures do not consume either quota, authorize a ceiling change, or disappear from chronology.

3.23 Delivery and validation-economy correction: 2026-08-10

The prior plan correctly built fail-closed evidence machinery, but it made a release-qualification campaign the sole WIP item before the language subject was ready for release qualification. That was a sequencing defect. Repeating the identical approximately six-minute standard command across ninety successful class observations could refine a future SLO, but it could not change whether the operational release DAG, watchdog, host lifecycle, self-host contract, or language implementation worked. It therefore consumed the only implementation lane without advancing the product.

  1. R0.4.j closes on its actual operational acceptance. P1.7 established bounded aggregate publication ownership and passed the complete generated-authority guard union. P1.6 then produced a passing exact-main standard run, passing v0.2 release-full DAG and read-only aggregate, complete typed target dispositions, and independent watchdog recognition. Those were the task’s functional and delivery-liveness predicates. The later timing collector remains useful, but a calibration population cannot retroactively become a prerequisite for an already-proved release-control implementation.
  2. Statistical timing ratification moves intact to R9.1.c. The three classes, hard containment ceilings, five discarded warmups, thirty retained conformant samples, hard-failure retention, exact statistics, anti-relabeling rules, and trend policy are preserved. Existing identity-matched E0 observations remain reusable. Their completion first blocks a release-candidate decision at R9.1.c, not R0 implementation, M1 agent work, runtime work, or self-host migration. Until ratification, the existing 720,000ms local and 7,200,000ms hosted values remain conservative hard containment limits and no permanent SLO is claimed.
  3. Validation follows subject readiness. Contract and negative tests may precede implementation. Focused correctness and failure-path tests begin with the smallest runnable slice. One impacted integration profile runs after those pass. Cross-host matrices begin only when the claimed host implementation exists; statistical performance, fuzz, and soak campaigns begin only when their measured subject is correct, runnable, and stable enough that the result can change an engineering decision; release matrices and long soak begin only in R9. A synthetic adapter, stub, unsupported product, or known-broken functional path is not a useful subject for qualification.
  4. Identical success is not additional evidence during development. For one exact revision, inputs, host class, cache state, seed, and command, the first successful invocation closes the ordinary implementation check. One additional identical invocation is allowed only to reproduce or confirm a recorded flake or nondeterminism hypothesis. Further repetition is valid only when a declared independent variable changes or when one bounded autonomous harness is itself the task subject, such as a microbenchmark distribution, fuzz campaign, lifecycle stress run, or release soak. Whole-repository profiles are not rerun once per inner sample.
  5. Every campaign must be decision-bearing and autonomous. Before launch it names the decision the result can change, subject-readiness evidence, independent variable, observation-reuse rule, resource budget, stopping rule, and terminal artifact. It runs without model-turn polling and reports state transitions or terminal outcomes, not unchanged snapshots. If the result cannot alter implementation, containment, promotion, or release, the campaign is deferred or removed.
  6. The critical path is implementation-first and self-host-first. Close P1.5/R2.2.f, publish SH-A, freeze R5.1/R5.3 semantics, migrate R4.2.a-e production authority, finish only the runtime/tiering prerequisites required for compiler authority, then close R4.2.f-g and R4.3-R4.5. Remaining agent-product polish follows the self-host authority path; GenesisBench inference, Foundry implementation, and Genesis Model work stay frozen by their existing gates. A blocked frontier anchor no longer prevents selection of its next ready prerequisite.
  7. No safety evidence is discarded or weakened. All interrupted, failed, superseded, and successful timing records remain append-only observations. Hard timeout, disk, process-group kill/reap, capability, replay, panic, and semantic-parity checks remain mandatory at the phase where their subject exists. This correction removes redundant qualification work, not adversarial coverage.

3.24 GenesisCode completion handoff and post-Core fan-out correction: 2026-08-10

Historical note: section 3.25 supersedes this section’s immediate Bench/Challenge fan-out with the finite F2.r Foundry Foundation and Genesis Algorithms gate; the Core boundary and lane-isolation reasoning remain in force.

At this historical audit, the program still required a trustworthy language before spending its only implementation lane on products that evaluate, research, or improve that language, but it removed the then-current serialization of GenesisBench, Genesis Foundry, and GenesisChallenge after the language handoff. Section 3.25 subsequently found that immediate fan-out too weak: the current rule is exact reviewed Core, then finite F2.a-r Foundry Foundation and Genesis Algorithms, then downstream fan-out. The Core boundary, product independence, and lane-isolation conclusions below remain authoritative; the old activation order does not.

  1. Exact reviewed R9.4.f is the GenesisCode handoff. It closes M6 GenesisCode v1 Core, not every optional task in the durable program. On this audit’s derived graph, the R9.4.f ancestor closure contains 255 tasks, of which 80 are already complete and 175 remain open; the other 124 roadmap tasks cannot block Core. Counts are observational and must be recomputed from the execution manifest, while the task identity and release contract are normative.
  2. The implementation lane is GenesisCode-only until the handoff. Before R9.4.f, one repository-changing transaction follows the generated Core frontier. No target-model run, benchmark custody or commissioning, GenesisChallenge operation, Foundry implementation, Genesis Model work, or optional platform-pack breadth may consume that slot. The only concurrent activity is read-only assurance that directly satisfies the selected GenesisCode task, runs on separately provisioned resources, touches no repository or release authority, and requires no attention from the active implementation agent.
  3. Core is small enough to finish and complete enough to build on. The handoff includes the frozen language and effect semantics, bounded runtime and validated execution tiers, H2 self-hosted semantic authority, H3/H4 bootstrap and independent conformance evidence, package/build/registry/deployment Core, generated human and agent documentation, safe agent interfaces, native/WASI CLI, authentic service operation, security/compatibility/support policy, signed artifacts, and independently reproduced release evidence. It explicitly excludes universal web/mobile/game/MCU/GPU pack maturity, benchmark campaigns, Foundry, Challenge, and Genesis Model. Exclusion means separate compatibility-scoped release work, not abandonment.
  4. Historical fan-out rule, superseded by section 3.25. This audit activated GenesisBench Preview/Trust, F2.a-q calibration, and GenesisChallenge immediately after exact reviewed R9.4.f. The current DAG instead runs finite F2.a-r first and fans Bench/Challenge out only after the independently promoted algorithm-library release. The retained rule is that every active lane has one coherent WIP slot, separate owners, files, compute, custody, and promotion authority; shared-interface changes still merge serially and cannot rewrite a frozen consumed profile.
  5. Cross-product edges remain task-specific, not lane-wide. F2.a-r uses frozen GenesisCode semantics and independent checkers, so it does not wait for GenesisBench. F2.y, which explicitly integrates GenesisBench and the later model ecosystem, still waits for a compatible R8.5.s Trust release. GenesisChallenge now waits for F2.r but not GenesisBench Trust; its benchmark-validity category stays ineligible until a compatible benchmark baseline and destination owner exist. Genesis Model remains final and requires GenesisBench Trust plus independently reproduced MF through R8.5.t.
  6. The DAG must prove every handoff rather than merely describe it. The GenesisCode release contract rejects unfinished Benchmark, Challenge, Foundry, Model, universal-pack, and pack-assurance ancestors. The current contracts additionally reject Bench/Model ancestry in F2.r, require F2.r in Bench and Challenge, reject F2.s-z expansion ancestry in those downstream release roots, require Benchmark Trust for F2.y, and reject Challenge ancestry from Model readiness. Mutation controls must fail if any forbidden edge or prerequisite bypass is introduced.
  7. The pre-Core selector is continuation-complete. Ordered anchors cover every unfinished task in the exact R9.4.f ancestor closure, including each Core member of non-sequential workstreams. Sequential-anchor advancement is filtered to that closure, so trailing pack-only assurance cannot leak onto the Core clock. A missing Core anchor, post-Core anchor, or selector candidate outside the closure fails the roadmap manifest instead of requiring a future manual queue rewrite or silently producing an empty focus.

3.25 Foundry-first algorithm-library correction: 2026-08-11

The earlier post-Core fan-out correctly removed Foundry from the GenesisCode release dependency graph, but it treated calibrated discovery as an optional peer of GenesisBench and provided no finite path from retained research results to a supplied GenesisCode algorithm library. That left two avoidable defects: Bench could construct tasks before the project had a canonical first-party algorithm package, and useful Foundry outputs had only a late generic promotion task after model/Bench integration. This correction replaces that ordering without moving research into production trust.

  1. Foundry Foundation is the first post-Core outcome. Exact reviewed R9.4.f activates F2.a-r. F2.a-q builds and independently reproduces the bounded search/verification calibration; F2.r establishes the destination-owned package boundary and releases the first genesis/algorithms collection. GenesisBench Preview, GenesisChallenge operation, and Genesis Model work remain non-start-ready until F2.r. Optional platform packs and unrelated F1/F3/F4 research may still proceed under disjoint authority because they neither define nor consume the benchmark/library contract.
  2. The finite gate is not “finish Foundry.” Foundry is an enduring research program, so requiring F2.a-z before Bench would replace one sequencing mistake with permanent blockage. F2.r is a bounded foundation release with exact artifacts, independent reproduction, destination review, compatibility, and rollback. F2.s-z expand engines, abstractions, laws, domains, integrations, and long-term memory after that foundation and may run beside Bench under isolated resources.
  3. Search and library authority are split. foundry/ owns quarantined grammars, candidates, search state, failures, counterexamples, and proposal receipts. genesis/algorithms is an ordinary signed .gc package collection owned by package maintainers. Promotion copies a reviewed semantic artifact across a versioned boundary and reruns independent conformance without the search engine; the library cannot import Foundry code or mutable state, and Foundry cannot approve its own package proposal.
  4. The library is useful even when discovery is not novel. Its first release includes independently verified reference implementations and exact-domain specialized variants from the calibration family, with stable APIs, deterministic fallbacks, complexity contracts, provenance/licenses, examples, and supported-profile tests. A result may be valuable because it is correct, portable, explainable, or efficient; historical novelty is neither required nor implied.
  5. No same-release circularity is permitted. A Foundry run may use algorithm-library release N only as a frozen baseline and may propose release N+1. Its task, grammar, verifier, and acceptance are precommitted against N; it cannot mutate the baseline, oracle, or destination during search. Package N+1 is rebuilt and verified from its promoted .gc source with foundry/ absent.
  6. Bench consumes only disclosed releases. GenesisBench binds exact compatible GenesisCode and genesis/algorithms identities. Public library APIs may be part of the provided environment, task surface, examples, or baselines, but mutable Foundry state, hidden candidates, and search-specific oracle material are excluded. Scorers verify task behavior independently and never infer correctness from library membership or a Foundry receipt.
  7. Challenge and Model remain downstream. GenesisChallenge begins after F2.r so real algorithm/library and Foundry improvements can be contested under destination-owned review; benchmark-validity categories still wait for a compatible Bench release. Genesis Model still waits for Bench Trust and R8.5.t, and its research-transfer authority consumes the independently reproduced Foundry Foundation plus explicit promoted/rejected transfer records rather than mutable search history.
  8. The machine DAG enforces the correction. F2.r becomes the Foundry release root; every unfinished GenesisBench Preview leaf and the Trust root inherit it; Challenge authority inherits it; model transfer/readiness inherits it. Mutation controls reject a Bench, Challenge, or Model path that bypasses F2.r, reject Foundry Foundation ancestry on Bench/model implementation beyond the explicitly allowed a-r set, and preserve the prohibition on Bench/Model ancestry entering F2.r.

3.26 Foundry-to-applications correction: 2026-08-11

Section 3.25 correctly put a finite Foundry Foundation and supplied algorithm release before Bench, but it still jumped directly from a library artifact to evaluation and contribution programs. That omitted the public proof that ordinary agents and users can compose the language and library into useful systems. This correction refines the existing R8.3 flagship allocation rather than adding a fifth product, semantic authority, or unbounded prerequisite.

  1. Foundry and Genesis Algorithms remain first. Exact reviewed R9.4.f still activates F2.a-r and nothing in R8.3 can enter implementation WIP before F2.r. Examples consume a signed destination-owned library release, never candidate state or a mutable Foundry workspace.
  2. Examples and apps have different jobs. GenesisExamples are compact executable learning/reference projects; GenesisApps are maintained production-style systems. Conformance, benchmark-reproduction, and host-workflow fixtures remain fixtures. A catalog manifest records the classification and rejects marketing relabeling.
  3. The downstream gate is finite. R8.3.a-b publishes a bounded seed containing an algorithm/CLI, MCP server, real web server, separate full-stack website/application, and integrated workspace. It proves meaningful algorithm use, authentic execution, one-language source closure, maintenance, and operation with foundry/ absent. It does not wait for all 18 archetypes or ten universal flagships.
  4. Universal breadth remains independent. R8.3.c-e retains the prior ten-flagship, real-target, maintenance, source-audit, and independent-operation burden and advances as platform packs mature. Open-ended app breadth cannot block Bench, Challenge, Foundry expansion, Core maintenance, or Model readiness.
  5. Bench learns from surfaces, not answers. The public catalog may provide disclosed context, practice, capability taxonomy, and task-family evidence. Held-out lineages require independent provenance and cannot copy public source, oracles, traces, seeded defects, or acceptance secrets. Catalog membership never substitutes for scoring.
  6. Challenge may expand the catalog without controlling it. After R8.3.b, GenesisChallenge may offer creation, promotion, target-expansion, maintenance, simplification, and optimization work against frozen baselines. Acceptance remains with the example/app destination owner and never follows automatically from a prize, score, or solver claim.
  7. The machine DAG enforces the order. R8.3.a inherits F2.r; R8.3.b inherits the exact product profiles used by the seed; unfinished Bench Preview leaves, Bench Trust, and Challenge inherit R8.3.b; broad R8.3.c-e remain forbidden from their release ancestry. Mutation controls reject bypass, reverse dependency, mutable Foundry input, and accidental all-flagship serialization.

4. Non-negotiable invariants

These invariants apply to every phase and every execution tier:

  1. Pure kernel: G-lambda evaluation has no filesystem, time, randomness, network, process, environment, UI, model, or hidden global-state access.
  2. Unforgeable seals: UNHANDLED, EFFECT, and ERROR are observable only through Prelude-created seal tokens. User terms cannot forge them.
  3. No panic on user input: parse, evaluation, effect, bridge, package, registry, compiler, replay, and deployment boundaries return explicit internal errors or sealed language errors. Production panic!, unreachable!, unchecked indexing, and abort paths require mechanically proven non-user reachability.
  4. Deny by default: every effect, host bridge, plugin, package hook, deployment action, model call, and inbound service operation requires explicit authority.
  5. Deterministic facts: canonical terms/hashes, value hashes, ordering, paths, numeric behavior, effect logs, replay decisions, scheduling, artifacts, and bootstrap outputs are specified and versioned.
  6. Strict replay: replay checks every serialized fact or the specification explicitly marks a field derived and recomputes it canonically. Unknown/missing fields fail closed.
  7. Bounded execution: step, allocation, heap, payload, response, effect, concurrency, process, wall-clock, and artifact budgets have defined enforcement and failure semantics. Cooperative timeout alone is insufficient for blocking host work.
  8. Semantic identity across tiers: treewalk, compiled AST, bytecode, WebAssembly, and any JIT produce identical observable values, hashes, effects, errors, and resource-accounting semantics, subject only to explicitly versioned performance counters.
  9. Optimizers remain outside trust: optimization, specialization, JIT, and AI-generated rewrites require differential or translation-validation evidence against authoritative semantics.
  10. Compatibility is explicit: canonical identity and serialized formats never drift silently. Every incompatible change has a profile/version bump, migration, dual-read window where safe, and deprecation record.
  11. Evidence cannot bless itself: the implementation under test cannot be the sole authority that validates its own release claim. Critical formats have small independent verifiers and negative-control corpora.
  12. Self-hosting reduces semantic trust: migrating code to .gc must remove or demote the prior authority. Duplicating semantics without retiring an authority does not advance closure.
  13. Generated closure is atomic: canonical-input drift, generated-output drift, and incomplete transitive regeneration fail focused checks and CI. Read-only checks never repair state, and publication cannot proceed from a partially refreshed authority graph.
  14. Source and generated state stay distinct: ignored/generated roots are excluded from source topology only by explicit classification, ownership, quotas, and cleanup policy. They cannot hide source, tests, private custody violations, or release evidence.
  15. Host boundaries are quantity- and scope-correct: resource limits, cancellation, pipe lifecycle, locking, and persistent allocation name and enforce the exact promised process/thread/session scope on every supported host. Platform approximations, polling intervals, and unsupported hard guarantees are explicit; supported-host CI and contention stress are mandatory evidence.
  16. Research depends inward, never production outward: Foundry may consume frozen GenesisCode crates, schemas, artifacts, and binaries; no production crate, default runtime path, release gate, or language semantic authority depends on Foundry implementation code or mutable research state.
  17. Discovery cannot promote itself: candidate generators, portfolio controllers, LLMs, task generators, benchmarkers, and research registries cannot certify correctness, erase failures, alter hidden evaluation, or approve their own laws/packages/rewrites. Promotion is destination-specific, independently verified, reversible, and preserves every rejected ancestor and counterexample.
  18. Proof scope is exact: a proof, exhaustive check, differential test, cost model, benchmark, or theorem is bound to the exact candidate, lowering, semantic profile, assumptions, target, and verifier identity it covers. No theorem about an abstraction is silently laundered into a claim about emitted machine code or a wider input domain.

5. Release ladder and critical path

Milestone Intended release User-visible result Hard gate
M0 Truthful Green Door v0.2.x Fresh clones build and checks report facts without mutating them R0
M1-A GenesisCode Agent Preview GenesisCode v0.3 Compact SDK, stable diagnostics, local warm/MCP interface, generated authoring skill, safe collaboration, and typed English-intent preview support reproducible agent work R1.1-R1.3 and R1.5-R1.7; R0.4.j
M2 Runtime Beta v0.4 Interactive, resource-bounded execution with validated bytecode and bounded long-lived processes R2-R3 core
M3 Self-Host Authority Beta v0.5 Frozen core semantics above stage0 are GenesisCode-authoritative and bootstrap to a cross-host fixpoint R5.1 and R5.3 core freeze, then R4
M4 Platform Pack Betas independently versioned packs Web/UI, service/data, desktop/mobile, game/media, Embedded Linux, MCU/boards, and GPU/XR packs advance independently with authentic targets and one-language source closure relevant R5-R6 workstream per pack
M5-C GenesisCode v1 Core Candidate GenesisCode v1.0-rc Formal/fuzz/security evidence, frozen core semantics, self-host authority, agent interfaces, native/WASI plus authentic CLI/service proof, and independent operation meet Core SLOs R7.1.a-e, R7.2, R7.3.a-e, and R7.4.a-c; R8.1, R8.2.a-b, and R8.4
M5-P Universal Platform Maturity independently versioned pack and GenesisApps catalog releases Eighteen authentic agent-authored archetypes and ten maintained GenesisApps, including real boards, prove the growing one-language portfolio without blocking Core, Bench, or Challenge R8.2 plus R8.3.c-e; may complete after M6 and after the finite R8.3.a-b seed
M6 GenesisCode Trust Release GenesisCode v1.0 Core Reproducible artifacts, signed evidence, compatibility guarantees, and operational support are published; compatible pack and benchmark versions are named independently, and a model may be absent R9 Core scope
F1 Frontier Governance Lab post-v1 Bounded self-improvement under preserved v1 compatibility and independent roles F1
MF Genesis Foundry Foundation independently versioned post-v1 Foundry snapshot plus Genesis Algorithms v0.1 Deterministically rediscover, verify, compare, explain, and replay bounded mechanisms, then independently promote useful .gc reference/specialized algorithms into a signed supplied library with no research-runtime dependency or novelty claim R9.4.f, then F2.a-F2.r
M1-B GenesisBench Preview first post-v1 GenesisBench release; stable historical ID Four isolated tracks, transport-neutral task/artifact contracts, model-portable adapters and failure attribution, real-model commissioning, a useful easy/medium/hard private epoch with precommitted fast forms, reproducible results, and public Preview governance against immutable v1 Core, Genesis Algorithms, and finite GenesisExamples/GenesisApps seed releases MF, then R8.3.a-b and remaining R1.4.o/q/r/p
M5-B GenesisBench Trust Release independently versioned GenesisBench release A >=90-lineage temporal, maintenance, reconstruction, calibrated-difficulty, and model-portable frontier is independently reproduced across standard/full forms and three interface stacks, governed, discriminative, and useful to language and library adoption MF, M1-B, then R8.5.a-e and R8.5.s
M5-M Genesis Model Preview independently versioned model package After v1 Core, GenesisBench Trust, MF, and R8.5.t, a firewall-audited, lineage-bound, profile-bound local specialist is reproducible and useful without becoming a language, library, or benchmark dependency R8.5.h,i,t,f,g,j-r with applicable four-cell or frozen-language two-cell controls; necessarily after M6, MF, and M5-B
F3/F4 Frontier Systems post-v1 Optional validated specialization/JIT, distributed/hardware execution, independent implementations, and evidence-weighted governance F3-F4

Milestone IDs are stable historical authorities, not sortable execution priorities. Sections 3.24-3.26 govern activation: M6 opens Foundry Foundation plus independent platform-pack and unrelated F1/F3/F4 lanes; MF then opens the finite GenesisExamples/GenesisApps seed and continued Foundry expansion; R8.3.b then opens GenesisBench and GenesisChallenge. M1-B precedes M5-B inside GenesisBench; M5-B and MF are prerequisites of the later Genesis Model readiness decision.

Critical path:

R0 truth/evidence
  -> R2.2 host/resource lifecycle closure
  -> R4.1 self-host truth boundary
  -> R5.1 + R5.3 core semantic/profile freeze
  -> R4.2.a-e non-codegen semantic authority
  -> R2.1/R2.3/R2.4 + R3 execution prerequisites
  -> R4.2.f-g + R4.3-R4.5 self-host authority and bootstrap fixpoint
  -> remaining R1 agent product surface
  -> R5.2 + R5.4-R5.5 core language/extension completion
  -> R6.1-R6.4 package/build/registry/deployment core
  -> R7.1.a-e + R7.2 + R7.3.a-e + R7.4.a-c Core assurance
  -> R8.1 + R8.2.a-b + R8.4 core product proof
  -> R9 v1 Core release
  -> F2.a-r Foundry Foundation + Genesis Algorithms v0.1
  -> R8.3.a-b GenesisExamples/GenesisApps seed
  -> GenesisBench Preview/Trust and GenesisChallenge
  -> R8.5.t Genesis Model readiness -> optional Genesis Model

Platform packs: relevant R5.6-R5.9 -> R6.5-R6.7 -> R8.2/R8.3 pack proof -> independent pack release

Concurrency contract:

  • Before R9.4.f: one repository-changing GenesisCode Core transaction owns the implementation lane. Work described as “parallel” elsewhere in this historical roadmap is scheduling-ready only and cannot consume repository WIP, local compute, shared locks, or active-agent attention. One separately provisioned, read-only assurance observation may coexist only when it directly satisfies the selected Core task and has no authority to mutate or close it.
  • At R9.4.f: publish the signed immutable Core identity and activate F2.a-r Foundry Foundation first. Optional platform packs and unrelated F1/F3/F4 research may use isolated lanes, but GenesisBench, GenesisChallenge, and Genesis Model remain frozen. Shared schema, compatibility, release, and generated-authority changes remain serial merge transactions even when lane-local implementation is concurrent.
  • Genesis Foundry Foundation: execute F2.a-q calibration against frozen Core semantics and independent checkers, then F2.r promotes Genesis Algorithms v0.1 through a separate package authority. Foundry runs against an immutable prior library baseline, never the release it is proposing, and production builds remain complete with foundry/ absent.
  • At F2.r: activate R8.3.a-b as the finite GenesisExamples/GenesisApps seed transaction and continued F2.s-z expansion as an isolated lane. Bench, Challenge, and Model remain frozen until the seed is released. This preserves Foundry-first ordering without making open-ended Foundry or universal app breadth a permanent downstream blocker.
  • GenesisExamples/GenesisApps seed: publish exact catalog identities plus the algorithm/CLI, MCP, web-server, website/application, and integrated-workspace proofs against immutable Core and Genesis Algorithms releases. Every item builds with foundry/ absent, uses released algorithms materially, and separates public learning/reference content from independent held-out benchmark lineages.
  • At R8.3.b: activate independent GenesisBench Preview/Trust and GenesisChallenge lanes. Continued R8.3.c-e universal showcase work proceeds independently and cannot block their release clocks.
  • GenesisBench: complete remaining Preview and Trust work against exact immutable Core, Genesis Algorithms, and public example/app catalog identities, with independent custody, compute, adapters, scoring, and publication authority. Bench may consume only disclosed releases, never mutable Foundry state, published solutions as held-out answers, or search-specific oracles.
  • GenesisChallenge: establish and operate its separate non-ranking ledger after R8.3.b. A challenge category is eligible only when its destination has a ready verified baseline and independent owner; example/app expansion is an explicit destination class, and benchmark-validity challenges additionally require a compatible Benchmark release.
  • Platform and frontier lanes: optional packs plus F1/F3/F4 may advance after Core under separate resources and compatibility contracts. They cannot retroactively broaden Core or borrow M6 evidence.
  • Genesis Model: remains outside the initial fan-out. Architecture, corpus construction, teacher selection, training, distillation, packaging, and integration stay at zero WIP until R8.5.t records GenesisBench Trust, reproduced MF, firewall, feasibility, and independent-review readiness.
  • JIT research can start after the bytecode semantics are stable, but cannot displace critical-path work without failing a published performance decision gate.
  • Additional capability families remain experimental until existing families meet the maturity and domain-proof rules. Universal-product work grows primarily through libraries, components, target adapters, and BSP packages; it does not expand the kernel or bypass profile negotiation.
  • Foundry architecture/schema ideas may be captured before v1 only as non-authoritative inputs that are read-only with respect to production and benchmark semantics, consume no release capacity, and produce no algorithm or optimization claim. Ratification, executable search, and research-registry operation begin only after compatible v1 Core semantics/evidence interfaces are frozen. F2.r is the finite prerequisite for the GenesisExamples/GenesisApps seed; R8.3.b is the finite prerequisite for first GenesisBench and GenesisChallenge work; later Foundry and universal-showcase expansion never block GenesisCode maintenance, platform packs, or later compatible Bench releases. Independently reproduced MF and GenesisBench Trust both remain prerequisites of the optional Genesis Model readiness decision.

6. Program-wide budgets

Budgets are acceptance contracts, not aspirations. R0 records the reference machines, raw samples, and confidence method. Budgets may be tightened freely. Loosening one requires a dated rationale, before/after evidence, and an explicit milestone decision.

6.1 Agent experience budgets

ID Budget Preview target v1 target
AB-1 Core language card <= 4,000 tokenizer-independent approximate tokens <= 3,000
AB-2 Task-specific context bundle <= 12,000 tokens including capability cards and examples <= 8,000
AB-3 Warm request latency, no execution p50 <= 5ms; p95 <= 20ms p50 <= 3ms; p95 <= 10ms
AB-4 Changed-file edit/check loop p95 <= 2s on a 100-module workspace p95 <= 1s
AB-5 Reference-agent parse success >= 95% first attempt on public test split >= 98%
AB-6 Diagnostic-guided repair >= 85% recovery within two repair turns >= 95%
AB-7 Held-out capability coverage >=8/9 core task classes have an independently verified baseline solve across the >=45-lineage Preview, with uncertainty over lineages GenesisBench Trust covers 9/9 classes across two independent model families on compatible frozen GenesisCode profiles; product-archetype coverage is reported separately under UP/M5-P and never blocks the benchmark
AB-8 Structured-output stability 100% protocol/schema conformance or explicit typed failure 100%
AB-9 Benchmark run reproducibility 100% of published runs validate and deterministically rescore from their closed bundle 100% plus independent replay by two operators
AB-10 Benchmark-to-language feedback Every repeated failure cluster has a typed owner and disposition >= 90% of eligible clusters produce a language/tooling fix or an explicit non-goal
AB-11 English-intent safety 100% of requests become a typed goal/acceptance/authority record or explicit abstention; no model response directly mutates accepted state 100% across two independent general-model families
AB-12 Trajectory completeness 100% of retained model calls, retrievals, tool calls, attempts, diagnostics, patches, decisions, and outcomes are content-addressed or explicitly redacted by policy 100% plus independent bundle replay

Agent benchmarks always publish model/runtime version, decoding settings, context, seeds where available, attempts, failures, and cost. A model result never substitutes for deterministic compiler/runtime conformance.

6.2 Benchmark validity budgets

These budgets govern benchmark claims separately from language correctness and runtime performance:

ID Contract Acceptance
BQ-1 Closed-run reproducibility Every ranked run binds repository, runtime, cards, prompts, model/runtime identity, attempts, candidates, effects, scores, and host facts; deterministic rescoring is byte-identical.
BQ-2 Oracle and held-out isolation Zero private prompt/input/oracle bytes in distributed packs, model contexts, logs, diagnostics, training corpora, or public run bundles before authorized disclosure.
BQ-3 Contamination truth Every result is labeled temporal-clean, declared-uncontaminated, declared-contaminated, or unknown; only tasks committed after a model’s immutable release may receive the temporal-clean label.
BQ-4 Frontier freshness At least 25% of scored challenge weight is newly precommitted per active epoch; rotate within 90 days, on material leakage, or when saturation triggers, whichever occurs first.
BQ-5 Saturation response If the top three eligible model families each exceed 90% quality for two consecutive epochs, advance task depth or retire ranking weight without rewriting historical results.
BQ-6 External reproducibility M1-B requires one project-controlled reproducible baseline; GenesisBench Trust requires two independent operators to replay and rescore at least three model families from closed bundles.
BQ-7 Training-corpus admission 100% of admitted artifacts have compatible license/consent, complete lineage, passing semantics/obligations/policy/resource checks, independent review, deduplication, and zero held-out overlap.
BQ-8 Evaluation firewall Training builders cannot read active held-out custody material; evaluators cannot silently write training stores; cross-role access and artifact movement are capability-gated and audited.
BQ-9 Statistical unit The independent task lineage, not its context/tool/scaffold condition, is the sampling unit. Reports cluster all repeated conditions by lineage, publish denominators and uncertainty, and never count a 3-condition ablation as 3 independent solves.
BQ-10 Track isolation cold-acquisition, open-agent, genesis-adapted, and embedded-local results have distinct eligibility and ranking cohorts. No headline rank, aggregate, or confidence interval crosses track, profile, epoch, scaffold, or incompatible hardware class.
BQ-11 Challenge scale and balance M1-B Preview has >=45 independently committed and private-payload-verified held-out lineages, >=5 per core task class, across declared difficulty bands and >=2 independently disclosed author/generator authorities; GenesisBench Trust has >=90, >=10 per class, and no author/generator family controls >25% of ranking weight. Commitment metadata without custody-verified payload/oracle bytes does not count.
BQ-12 Reference scaffold and adapters Cold Acquisition fixes and hashes system prompt, retrieval, cards, tools, transactions, repair policy, budgets, and adapter conformance. The model revision is the only meaningful variable; any scaffold change creates a new cohort/version.
BQ-13 Useful novelty and maintenance Post-release overlays introduce useful precommitted packages/capabilities rather than riddles, and >=25% of mature ranking weight exercises a follow-on requirement, migration, defect, policy tightening, or resource change against a previously accepted artifact.
BQ-14 Public governance and correction Challenge/result admission, custody, conflicts, embargo/disclosure, security reports, appeals, invalidation, supersession, and operator rotation follow a signed versioned policy. Corrections append status and rationale without deleting accepted history or silently recomputing old cohorts. A public benchmark/leaderboard requires at least one disclosed independent non-project reviewer/operator; project-controlled evidence remains labeled Preview.
BQ-15 Orthogonal stratification Every scored lineage has one immutable epoch-local easy, medium, or hard calibration plus separately keyed context, run profile, task class, capability/product surface, track, language profile, scaffold/adapter, attempt policy, and hardware cohort. Reports publish every required stratum and never infer difficulty from context size or merge incompatible cells.
BQ-16 Fast-profile integrity Every fast form is content-addressed, precommitted before target access, selected at lineage granularity by a declared stratified algorithm, and accompanied by its sampling frame, coverage omissions, uncertainty, and at least one alternate form. Fast results remain a separate directional cohort and cannot satisfy Trust-release, full-coverage, or strongest-model claims.
BQ-17 Reconstruction authenticity Every migration or product-reconstruction lineage binds license/provenance, destination profile maturity, behavioral and seeded-maintenance acceptance, one-language source closure, disposable generated artifacts, capability/resource/security bounds, and authentic target execution. Unsupported target support is excluded explicitly rather than scored as solver failure.
BQ-18 Model/interface portability The common artifact-text matrix requires only bounded text input/output and produces the same candidate workspace and score through completion, chat, local-runtime, HTTP, and JSON-stdio adapters. Native-tool and canonical JSON-text-tool encodings are metamorphically equivalent where supported; vendor-only features create separate interaction cohorts and never become eligibility requirements.
BQ-19 Failure attribution Every attempt has one independently rederived terminal owner class. A semantic model failure requires proven request delivery, closed response capture, conformant adapter/scaffold behavior, successful neutral extraction, intact workspace, and model-independent scoring; unsupported capabilities and adapter/provider/infrastructure defects remain non-scoring typed outcomes.
BQ-20 Tokenizer-neutral context Canonical content and budgets are bound in UTF-8 bytes plus structural counts, exact delivered bytes never vary by tokenizer, and provider tokens/tokenizer IDs remain reported resource facts. Per-model truncation, summarization, instruction tuning, or content reordering creates a different scaffold and cannot share a cohort.

6.3 Runtime budgets

ID Budget Current E0 observation Target
PB-1 Interpreter fib(25) strict workload signed baseline upper median-confidence bound about 211ms; independent later observation about 220ms <= 250ms, then ratchet
PB-2 Bytecode fib(25) not implemented <= 40ms
PB-3 Optional validated JIT fib(25) not implemented decision-gated; <= 8ms if built
PB-4 1M-element vector construction about 6ms on smaller audited workload; normalize corpus <= 150ms on normative workload
PB-5 100k-entry map construction signed baseline about 489ms; failing <= 300ms interpreter
PB-6 Cold CLI parse/check trivial module process startup about 285ms; deterministic snapshot not implemented <= 30ms with deterministic snapshot
PB-7 Self-host parse throughput signed lower bound about 42.8 KiB/s; failing >= 1 MiB/s, then ratchet
PB-8 Warm daemon leak not established <= 5% RSS growth after 100k bounded requests and quiescence
PB-9 Semantic tier parity broad partial gates exist 100% normative corpus, all values/effects/errors/hashes
PB-10 Bootstrap fixpoint not established cross-host byte-identical stage2/stage3 on all tier-1 hosts

6.4 Engineering and gate budgets

ID Budget Target
GB-1 Static policy/doc checks <= 15s warm, no compilation
GB-2 Default changed-file gate <= 2 minutes warm; no network; <= 1 GiB additional disk
GB-3 Standard local pre-commit profile <= 12 minutes warm; <= 3 GiB additional disk
GB-4 Full release profile <= 45 minutes on reference CI shard set; peak workspace artifacts <= 20 GiB
GB-5 Clean development footprint <= 8 GiB after a normal workspace build; genesis clean --cache-policy dev returns to <= 2 GiB generated state
GB-6 Evidence verification <= 5 minutes offline with no compiler invocation for an already-built release bundle
GB-7 Crate/module concentration no production Rust source file > 1,000 lines without exemption; no semantic crate > 20k lines by M3
GB-8 Fresh-clone prerequisites one checked manifest; no undeclared Python package/module assumptions

The 720,000ms GB-3 value is a provisional hard containment ceiling, not a statistically calibrated warm-performance claim. Local clean/fallback observations exposed that the prior 480,000ms bound was not valid for that different cache class, but they cannot establish a warm distribution or authorize a permanent SLO. R9.1.c owns release ratification of separate local warm, local clean/fallback, and hosted cold/shared-runner classes: retain five discarded warmups plus at least thirty conformant samples per claimed class; publish median, nearest-rank p95, MAD, and a distribution-free confidence interval; derive the hard ceiling from the declared host-class distribution and documented headroom; and add rolling trend alarms. Until then each overrun remains a typed hard failure under the existing containment ceiling. Failed samples remain preserved, cold evidence cannot be relabeled as warm, and an incomplete calibration cannot block pre-release implementation or authorize a performance claim.

6.5 Assurance budgets

  • Zero known P0/P1 trust-boundary defects at every milestone.
  • Zero user-reachable panic/abort paths in release profiles.
  • 100% negative-control rejection for seals, capabilities, replay tampering, package signatures, bootstrap artifacts, and evidence manifests.
  • At least 30 continuous days of sanitizer/fuzz/property soak with no unresolved high-severity finding before v1.
  • All tier-1 release artifacts reproduced by two independent builders; at least one uses an independently maintained verifier.

6.6 Genesis Model research-transfer budgets

These budgets govern what may be learned from Project Theseus, The ASI Stack, GenesisBench, and future research. They do not assert that any candidate architecture or training method is correct.

ID Contract Acceptance
MQ-1 Source and permission closure Every imported claim, report, artifact, dataset, trace, or design binds a clean source edition, path, digest, generator/command, environment, timestamp, owner, license/consent, and public/private class. Dirty worktrees, mutable latest aliases, dashboards, and uncaptured conversations are non-authoritative.
MQ-2 Evidence-state honesty Every candidate learning is classified as source-note-only, imported, artifact-missing, replay-ready, replay-failed, locally reproduced, stale, runtime-blocked, rejected, or archived lineage. Only locally reproduced evidence may directly support a Genesis Model design promotion; all other states remain visible.
MQ-3 Transfer replication Every promoted mechanism has a predeclared Genesis-specific workload, frozen baseline, simpler comparator, ablation, resource envelope, and deterministic acceptance where applicable. Stochastic claims report seeds, dispersion, uncertainty, failed runs, and hardware; one favorable run cannot promote a mechanism.
MQ-4 Firewall and privacy Zero active held-out bytes, private Theseus/ASI material without permission, secrets, evaluator instructions, or undeclared training rows cross into model research or training. BQ-7/BQ-8 and capability-separated storage remain mandatory.
MQ-5 Total-system efficiency Architecture, routing, memory, recurrence, test-time compute, distillation, and quantization changes are judged on verified task quality plus training/inference time, energy, memory, storage, latency, context, authority, complexity, inspectability, and rollback. A more complex method must beat the simplest adequate baseline on a predeclared Pareto criterion.
MQ-6 Claim-boundary preservation Theseus implementation receipts and ASI Stack synthesis may propose hypotheses but cannot establish Genesis Model capability, safety, self-improvement, or ASI claims. Every promoted claim names exact scope, assumptions, defeaters, non-claims, and independent verifier authority.
MQ-7 Negative-result retention Failed replay, null lift, regressions, stale evidence, missing artifacts, unsafe transfers, and no-promotion decisions remain content-addressed and searchable. Selection cannot silently discard them or train only on favorable descendants.
MQ-8 Currentness and supersession Refresh the transfer crosswalk before each model architecture freeze, training run, and release candidate. Changed upstream evidence creates a new candidate decision; it never rewrites the source edition, experiment, model checkpoint, or prior release history.

6.7 Universal product budgets

These budgets prevent broad platform names from substituting for usable products. Each target profile also publishes stricter hardware-specific limits.

ID Contract Acceptance
UP-1 One-language source closure Every canonical starter and flagship contains zero handwritten foreign application/build/deploy source. Generated foreign artifacts are inventory-bound outputs, reproducible from Genesis inputs, read-only to authors, and deletable without losing source.
UP-2 Authentic artifact execution 100% of claimed target artifacts install or start in the named real runtime and execute a nontrivial acceptance workload. Descriptors, empty exports, extension-only archives, host-side simulations, and scripts that merely hash bytes are ineligible.
UP-3 Product profile conformance Each product declares language, library, capability, target, SDK/BSP, artifact, lifecycle, and evidence profile identities. Unsupported combinations fail before build; no target silently falls back to a host interpreter or cloud service.
UP-4 Cross-target semantic identity Shared pure logic, schemas, state transitions, and replayable effects agree across every promoted target. Platform-specific behavior is isolated behind typed capabilities and explicit differential tests.
UP-5 Web quality Promoted web profiles pass real-engine navigation, responsive layout, keyboard/screen-reader semantics, security policy, offline/PWA where claimed, SSR/hydration equivalence, and declared startup/transfer budgets from Genesis source alone.
UP-6 Native/mobile quality Promoted desktop/mobile profiles pass install, launch, suspend/resume, termination, permissions, accessibility, input, storage, network loss, update/rollback, and crash-recovery tests on supported simulators plus named physical devices.
UP-7 Interactive quality Promoted game/graphics profiles meet declared frame-time, input-latency, audio-underrun, memory, asset-load, deterministic-simulation, and save/replay budgets on representative scenes; headless plan identity remains a prerequisite, not the product proof.
UP-8 Service quality Reference services pass typed protocol, authentication/authorization, persistence/migration, concurrency, backpressure, observability, secret, rolling-upgrade, rollback, fault-injection, and sustained-load SLOs in a self-hosted deployment.
UP-9 Embedded static closure Every MCU image has statically proven flash, RAM, stack, heap/arena, interrupt, task, peripheral, and energy/timing envelopes for its board profile. Unsupported dynamic language features fail at compile time with repair guidance.
UP-10 Hardware authenticity Embedded Linux and MCU claims include emulator/simulator evidence plus repeated hardware-in-the-loop boot, peripheral, reset, brownout/power-loss, watchdog, flash/debug, and OTA/rollback results on named boards.
UP-11 Product iteration After caches are warm, a representative one-file logic/UI/firmware edit reaches a runnable local target within the target’s published p95 budget; the build graph reports where time, disk, memory, and external SDK work were spent.
UP-12 Agent autonomy A published general model, using only the released Genesis SDK and declared domain packages, completes each promoted product archetype from English request through build, run/device test, policy minimization, package, deploy, and maintenance without human-written code or foreign-source edits.

6.8 Genesis Foundry research budgets

These budgets govern research validity; meeting them does not authorize a production change or establish historical novelty.

ID Contract Acceptance
FB-1 Reproducible search Same task, grammar, catalog, semantic profile, engine/version, seed, budget, and scheduler identity produce byte-identical candidate order, lineage, counterexamples, receipts, and terminal status; nondeterministic engines are isolated into explicitly statistical profiles.
FB-2 Authority isolation Zero production-to-Foundry code dependencies, zero ambient candidate capabilities, zero generator access to hidden verifier state, and zero self-promotion paths; mutation controls prove each boundary fails closed.
FB-3 Complete lineage 100% of generated, deduplicated, pruned, timed-out, rejected, superseded, and promoted candidates have content-addressed disposition records or a deterministic aggregate commitment that can prove membership without retaining unsafe payloads.
FB-4 Verifier false-accept defense Every task pack includes independently authored invalid candidates, forged receipts, wrong-profile proofs, hardcoded-oracle candidates, timeout/resource attacks, and law-poisoning mutations; accepted false positives are zero.
FB-5 Calibration before novelty The first snapshot rediscovers predeclared known optimal or near-optimal bounded structures and known lower bounds where available, with no novel label. Failure to recover the calibration set blocks broader search claims.
FB-6 Certified quotienting Every deduplication beyond canonical syntax cites a checked equivalence witness under an exact semantic profile. Heuristic similarity may prioritize search but never merges lineage, proofs, or identities.
FB-7 Search value Each engine and abstraction reports coverage, valid-candidate yield, counterexample yield, frontier quality, wall/CPU time, peak memory, disk, and review burden against exhaustive/random and no-abstraction baselines. Added complexity requires a predeclared material Pareto improvement.
FB-8 Cost honesty Abstract operation counts, measured host samples, target-specific models, uncertainty, warm/cold state, compiler/runtime identity, and applicability regimes remain separate. No scalar score fabricates a universal winner across incomparable profiles.
FB-9 Proof-to-artifact binding 100% of proof claims bind task, candidate graph, lowering, semantic profile, assumptions, theorem/checker identity, and emitted artifact where claimed; changed input invalidates the receipt rather than inheriting status.
FB-10 Counterexample permanence Minimized counterexamples and failing property classes enter an append-only curriculum before candidate repair; later candidates must defeat the entire compatible set, and supersession never erases the original failure.
FB-11 Research resource ceiling Every run has predeclared logical and host CPU, wall, memory, process, artifact, checkpoint, and energy-proxy limits; timeout/kill/reap is bounded and no descendant or lease survives terminalization.
FB-12 Independent reproduction FD4 requires a separately provisioned checker/operator to replay correctness and cost evidence from a closed bundle without the search engine, model, mutable research registry, or producing CLI.
FB-13 Library promotion firewall Every genesis/algorithms export has destination-owned API/semantic/complexity/applicability/provenance/license/compatibility review, an independently rerunnable test or proof bundle, rollback/deprecation policy, and a production build that succeeds with foundry/ absent. Search identity or FD status alone never admits an export.
FB-14 Library utility and portability Every public export is implemented in .gc, documented for humans and agents, covered by examples and malformed/boundary controls, and checked on every claimed Core profile. Specialized variants have deterministic applicability tests and a correct supported fallback; the release reports regressions against the prior version and simplest adequate baseline without requiring novelty.

7. Execution protocol

  1. Work the critical path in phase order unless a task explicitly names an earlier-safe parallel lane.
  2. Before implementation, add or identify the failing test, negative control, benchmark, or proof obligation.
  3. Prefer library behavior over CLI special cases and GenesisCode-authored semantics over new host semantics.
  4. Keep the pure reference evaluator readable and authoritative until a published self-host/TCB decision changes that authority.
  5. Run focused positive and negative checks first, then task-specific guards, then exactly one V2 integration route for the exact revision. Inspect bash scripts/test_changed_fast.sh --dry-run before launch; use its impacted route for ordinary source changes, or use one canonical generated-authority transaction when that transaction owns the source-to-derived closure and its declared read-only check set subsumes the changed-file route. Never run both for the same facts. Full workspace/release profiles run only where their manifest and validation stage require them.
  6. Never hand-edit generated evidence or generated matrices. Use one explicit canonical update transaction, inspect the complete diff and terminal check set, and rely on that transaction’s staged fixed-point/read-only verification. Do not rerun the updater merely to prove that the updater just reached its fixed point.
  7. Record wall time, peak RSS, disk delta, cache state, and network access for every governed gate profile when it runs; telemetry does not require an extra invocation.
  8. If a gate is broken, add it to upgrade_plan.md as an active P0/P1 defect with a roadmap back-reference. Strategic unfinished work stays here.
  9. If implementation reality contradicts this plan, fix the plan and evidence ledger in the same reviewed change. Do not preserve an inaccurate checkbox.
  10. Commit every coherent green checkpoint in task-scoped chunks and push it to a reviewed remote before starting the next high-risk tranche. Dirty-state evidence may be useful, but the local worktree may never be the only copy. Do not combine semantic format changes, migrations, and unrelated optimization.
  11. Default test profiles never invoke another governed aggregate pipeline and never absorb stress, soak, or performance loops. A test that fails under supported parallel load is a defect to reproduce and fix, not an ignore candidate.
  12. Until exact reviewed R9.4.f, 100% of repository-changing implementation WIP serves the GenesisCode Self-Hosted Core critical path: R1-A agent usability, R2-R3 runtime/tiering, R5.1/R5.3 contract freeze, R4 semantic authority/bootstrap closure, Core R5-R9 delivery, and executable product proof. New benchmark, Challenge, Foundry, model, governance, or optional platform-pack breadth cannot displace that path; new governance entrypoints require proof that an existing schema, gate, or negative-control matrix cannot be extended and normally replace or consolidate at least one equivalent mechanism.
  13. Enforce milestone scope neutrality and one high-risk WIP item globally before R9.4.f. After the signed Core handoff, permit one coherent item per isolated product/research lane, but serialize every shared-interface, compatibility, generated-authority, and release merge. A proposed task names the milestone and critical-path edge it serves, its falsifiable evidence, and the task or scope it refines, defers, consolidates, or retires; otherwise record it in section 9 without expanding the active queue. Changing milestone capacity or exit criteria requires an explicit dated decision, and product lanes never borrow completion claims from one another.
  14. The machine-selected execution frontier is stricter than the set of graph-ready tasks: one repository-changing GenesisCode production task is focused; queued priorities do not begin early; curated GenesisExamples/GenesisApps, GenesisBench campaign/publication, GenesisChallenge, Foundry implementation, optional-pack breadth, and Genesis Model implementation WIP are zero before R9.4.f. After R9.4.f, Foundry Foundation F2.a-r is the first downstream product path while packs and unrelated F1/F3/F4 may use isolated contracts. The R8.3.a-b seed waits for F2.r; GenesisBench and GenesisChallenge wait for R8.3.b; F2.y still waits for R8.5.s; Genesis Model implementation still waits for R8.5.t.
  15. No new feature family, target, schema, registry, gate, crate, or product archetype enters active WIP by default. It must close an observed acceptance gap for the current milestone and fit an existing authority; otherwise place it in the deferred frontier or a platform-pack backlog without changing the core release clock.
  16. Elapsed time, token use, commit count, repeated review, or an unavailable external service never activates a different product lane before R9.4.f. Continue the selected task’s local implementation or independently useful prerequisites; if its next edge is genuinely external-only, let the generated selector choose the next start-ready Core prerequisite rather than target-model access, benchmark custody/commissioning, campaign publication, Challenge, Foundry, pack breadth, or Genesis Model work. A concurrent read-only assurance observation must use separately provisioned resources and require no active-agent supervision.
  17. One successful invocation is the development limit for an identical command, revision, inputs, host class, cache state, and seed. One additional identical invocation is permitted only to reproduce or confirm a recorded flake or nondeterminism hypothesis; record that hypothesis before launch. A source or test change creates a new exact revision and may justify one new invocation.
  18. Repetition is evidence only when a declared independent variable changes or when a bounded autonomous harness is itself the subject. Microbenchmark samples, fuzz seeds, lifecycle fault cases, host/target matrix cells, and release soak intervals may repeat inside one governed harness; the surrounding whole-repository profile runs once after that harness passes.
  19. Before any campaign, record its decision, subject-readiness predicate, independent variable, observation-reuse rule, resource budget, stopping rule, and terminal artifact. Reuse compatible observations by exact identity. Statistical campaigns use a predeclared decision rule with anytime-valid uncertainty or an equally rigorous sequential method, a minimum identifiable sample, and a hard maximum budget; they stop as soon as the decision is resolved and report inconclusive rather than filling a quota or overstating an underpowered tail estimate. Do not start or continue a campaign whose result cannot alter an implementation, containment, promotion, or release decision.
  20. Long-running commands and campaigns run autonomously. Agents do not spend model turns polling unchanged output; they report launch, a material state transition, terminal result, or required intervention. Stop when the declared rule is met rather than filling an arbitrary sample quota after the decision is already resolved.
  21. Validation cannot outrun implementation maturity. Do not benchmark known-incorrect behavior, soak a stub, run cross-host qualification for an unsupported host, or exercise a release matrix before a release candidate exists. Preserve a truthful unsupported/not-ready result and advance the implementation that owns it.
  22. Expensive evidence never becomes a retroactive prerequisite for a completed task. If a stronger later assurance requirement is discovered, assign it to the earliest not-yet-completed milestone whose decision it can change; preserve the earlier evidence and reopen a completed task only for a reproduced regression in its original acceptance boundary.
  23. Machine-selected guards must test the selected task’s acceptance boundary. When one workstream contains materially different resource or evidence classes, bind the task to a narrower named execution profile rather than inheriting irrelevant benchmark, proof, or release commands. A checker must reject regression to the broader profile.
  24. A changed-file dry run is a launch decision, not ceremonial output. If its fallback is caused only by generated-authority outputs, dependency fanout already validated by the canonical transaction, or path-classification gaps, do not launch the aggregate. Reuse the exact transaction observations, repair the impact classification if necessary, and reserve fallback profiles for unresolved hand-owned semantic/configuration impact.

Validation ladder:

Stage Activation Required evidence Explicitly forbidden
V0 Contract Before implementation One falsifiable failure, negative control, or proof obligation Pretending an absent implementation passed
V1 Focused Smallest runnable slice exists Focused success, malformed/denied/failure path, deterministic rerun only when semantics require it Whole-suite loops and performance claims
V2 Integration Focused checks pass Task guards plus exactly one dry-run-selected impacted route: changed-file profile or subsuming canonical generated-authority transaction Running both routes for the same facts or repeating a green aggregate to accumulate confidence
V3 Assurance Relevant feature/workstream is functionally complete Declared property/fuzz/stress/cross-host matrix in one bounded autonomous harness Soaking known functional failures or unsupported products
V4 Release R9.1 release candidate exists Full host/reproducibility/security/performance matrix, class calibration, independent reproduction, and soak Using pre-candidate repetition to block implementation

R0. Truthful Green Door and evidence authority

Goal: make the repository, planning state, checks, evidence, versions, and resource costs tell the truth from a clean clone.

R0.1 Canonical planning and status

  • Evidence 2026-07-16: the canonical ledger now separates 29 foundation claims from 16 independently governed product/target claims spanning shared application/UI, static web, interactive/SSR/PWA web, self-hosted service/data, macOS/Linux/Windows desktop, iOS, Android, games/media, data/ML, Embedded Linux, closed-world MCU AOT, representative BSP/HAL families, device flash/debug/OTA, and physical hardware-in-the-loop. Every target is truthfully L0 and unsupported: current CoreForm packages, metadata-only mobile archives, empty Wasm exports, headless frame plans, host-operation routing, synthetic launchers, and simulator-only observations cannot raise target maturity. Each target records required and optional browser/host/runtime/simulator/device/board/lab scopes, mechanically coupled maturity and release state, an authentic execution predicate, a deletion-and-regeneration one-language predicate, owner, exact roadmap gaps, evidence IDs, limitations, and related foundations explicitly marked non-inheriting. docs/spec/PRODUCT_TARGET_MATRIX_v0.1.json and the product/target section of feature_matrix.md are generated alongside the five prior status views and are indexed in the documentation topology, Quarto site, source catalog, roadmap evidence bundle, and release notes without increasing the active documentation-leaf budget. The validator requires the broad aggregate boundary, at least one release-required scope, valid foundation links, one-to-one check evidence, no maturity/status escalation, immutable E3/E4 evidence for L5, and no stale gaps; 19 negative controls mutate these boundaries. Release notes independently reject product-target escalation, gap hygiene counts both foundation and product-target gaps, and no product or target is authorized for release.

R0.2 Evidence lifecycle

R0.3 Hermetic bootstrap and versions

R0.4 Gate architecture and disk lifecycle

  • Replace the contradictory v0.1 serial cold/warm pair with a v0.2 release-evidence DAG rather than increasing any limit. The closed DAG declares, for every required node, its governed gate identity, complete semantic and tool inputs, environment/host class, evidence class, cache-state sensitivity, repeat policy, isolation class, resource bounds, dependencies, and authorized consumers. Exact-revision invariant nodes execute once in the workflow and their authenticated same-run report may fan out to any number of manifest-bound consumers; this is evidence reuse, not result caching. No invariant node reruns merely because a second consumer, sample class, or aggregate needs the same fact. Cross-run .genesis/perf state, an unbound CI success, and a report from a different source, toolchain, environment, host class, invocation, or evidence class remain inadmissible.

    Cache-sensitive nodes use independent matched cohorts, not falsely paired samples from different hosted machines. Run at least three isolated cold workers and three isolated warm workers concurrently. Every cold worker starts from an empty, nonce-owned cache root. Every warm worker first performs a deterministic, bounded, untimed precondition into its own empty root, then proves the exact cache-key, toolchain, source, feature, artifact-inventory, ownership, and no-network state before starting the measured node set. Cold and warm cohorts execute the same closed cache-sensitive node set and match on source, workflow attempt, host profile, toolchain, workload manifest, limits, and GPU policy; they report distributions independently and never claim within-host deltas. Non-cache-sensitive stress or performance nodes execute only in their own contract-declared odd cohort of at least three exclusive workers. A node belongs to exactly one measurement class unless its normative question genuinely requires multiple classes, in which case the DAG records the reason and the aggregate rejects undeclared duplication. Optimizers, producers, pair/cohort workers, and prior reports cannot authorize or rewrite the aggregate.

    The read-only aggregate verifies complete DAG coverage, exact worker cardinality, disjoint owned roots, declared warm preconditions, no overlapping exclusive lanes, source/environment/target agreement, every artifact hash, cleanup recovery, bounded failure tails, and the absence of both missing and undeclared duplicate execution. It derives median, nearest-rank p95, dispersion, failure/abstention counts, peak RSS, generated-disk attribution, and critical-path wall time without treating correlated consumers or repeated reads as independent samples. Add adversarial controls for stale or cross-class report reuse, mismatched environment identity, forged warm state, nonempty cold state, hidden precondition time, duplicated invariant nodes, missing DAG leaves, overlapping exclusive workers, absent runtimes, impossible runner labels, cancelled-only history, missing schedules, wrong-head success, watchdog self-reporting, synthetic relabeling, and false success after an expected blocker.

    Keep node bounds, the per-worker GB-4 decision, job envelopes, the five-minute aggregate deadline, and the full-workflow watchdog distinct and ordered. Each worker must terminate and reap every child, retain complete logs plus a bounded portable failure tail, and leave diagnostic-publication time before external cancellation. Completion requires every required DAG node, not merely a prefix or one worker that reran the monolith, to pass on the exact reviewed main revision and named reference set; each measured worker is <=45 minutes and <=20 GiB, the successful DAG critical path plus aggregate is <60 minutes, and the disjoint watchdog recognizes a successful full run no older than forty-eight hours. The runner-absent path terminates within five minutes with exact setup, infrastructure, or unsupported-profile diagnostics. Append-only dispositions cover the 2026-07-18 through 2026-08-05 blackout and every subsequent failed full-profile design counterexample through the closing run. Capacity decision: this task owns P1.6, consumes the existing R0.4.f operational-hardening allocation, and must close before M1-A or M1-B publication; it defers no R6 implementation and does not make mobile, edge, service, GPU, or other platform-pack product claims early. Evidence 2026-08-05, round-eight disposition: the review correctly identified that hosted full had no successful authority and that the earlier monolithic run concealed its child failure, but its proposed single test_suite matrix and artifact diagnosis are now superseded by measured topology. Exact main revision b89fa08a947ff1edd4144553bc8694bf1af3f70b ran three disjoint governance/runtime/platform test lanes, four named target-readiness shards, two isolated release pair workers, and one read-only aggregate in Actions run 31023619554. All target shards and ordinary test lanes passed; the platform lane exposed a relocated-CARGO_TARGET_DIR Web binding bug, pair 2 exposed a 374159ms empty-target backend build against an incorrectly inherited 360000ms subordinate bound, and pair 1 reached the later stress sequence before the unchanged 2700000ms GB-4 supervisor terminated it. The corrective transaction resolves Web artifacts through the configured Cargo target, binds the release-only backend producer to its already-declared 600000ms gate envelope, reuses the exact release-full native gauntlet in native/WASI parity, and uses three stress repetitions in each independent pair rather than six repetitions in each worker. None of these corrections raises GB-4, changes a product claim, accepts stale evidence, or treats the review itself as completion. R0.4.j closes only after the exact promoted main revision satisfies the existing full-profile, pair, aggregate, target, watchdog, and append-only incident predicates above. Evidence 2026-08-05, exact-main serial-pair counterexample: PR #22 promoted the manifest-bound source-decomposition repair through all protected checks to exact main revision a4269136f17ca2e11c6b8f0bd296a2c3691d8138; its temporary branch was deleted locally and remotely. Full run 31063538202 then passed governance, runtime, platform, WebXR, deterministic GPU, all four named target-readiness shards, and both complete cold release-full executions. Pair 1 cold passed in 2568388ms; pair 2 cold passed in 2556616ms, each below the unchanged 2700000ms GB-4 ceiling with valid reports. Both workers nevertheless failed their warm child at the unchanged 3000000ms pair envelope: cold left only 430674ms and 442492ms, while warm logs independently reproduced the native gauntlet and roughly 313000ms backend matrix before reaching host-bridge stress. The source-decomposition gate no longer reran parity and was not the failure. This is a deterministic counterexample to the v0.1 requirement that two complete profiles execute serially inside fifty minutes: two individually legal 45-minute profiles cannot have a truthful 50-minute serial envelope, and the mandatory non-cache-sensitive work alone prevents a sub-eight-minute warm profile. Raising GB-4, the pair envelope, or the watchdog would conceal the contradiction. R0.4.j therefore requires the v0.2 DAG/cohort migration above, an append-only disposition for runs 31055018208 and 31063538202, permanent controls that reject the old impossible topology, and a successful exact-main full run plus independent watchdog recognition before closure. Evidence 2026-08-07, first exact-main v0.2 disposition: exact main revision 492aa3c5552b4584131ce82f394239cc5544a0d1 passed standard run 31169437484. Full run 31174700994 then passed governance, runtime, WebXR, deterministic GPU, all target-readiness shards, six of seven cache-sensitive workers, two of three stress workers, the invariant worker, and the GPU release gate. It exposed two independent orchestration defects rather than a language-semantic failure. First, the fan-out producer ran hosted image ubuntu24/20260720.247.2 with Node 22.23.1, while stress worker 3 ran rolling image ubuntu24/20260804.265.1 with Node 22.23.2; strict toolchain identity correctly rejected the consumer, but the diagnostic named no differing field. Second, the ordinary platform lane duplicated the DAG-owned native/WASI agent parity producer and failed only its tighter duplicate timing budget: both backends passed all 27 workflows and produced the same replay hash, while WASI took 2127ms against the duplicate 2080ms ceiling. The corrective transaction pins the hosted Node patch across every lane, rejects floating or inconsistent pins, preserves exact executable/version identity checks with field-level mismatch diagnostics and ten negative controls, and makes all release-evidence-owned direct updater steps standard-only so full CI executes each producer solely through its declared v0.2 DAG selector. No budget, evidence identity, parity requirement, or product claim is weakened. R0.4.j remains open pending a successful exact-main standard/full pair and independent watchdog recognition. Evidence 2026-08-08, historical closure: exact reviewed main revision da9359fd35221421563ea77b7dbe17e0a33d0e12 passed standard run 31249172621 in 4659s and release-full run 31249173667 in 2602s; both remained inside the independent 7200s standard-disposition and 3600s full-termination limits. The v0.2 aggregate 68fab0fa55bdf560a81f45b35168bb55123c9f1370b56c01c24ddb357fdd4345 derived its verdict read-only in 1176ms, authenticated all three cold, three prewarmed, one invariant, and three stress workers, rejected no declared condition, and reported status=pass, profileOperational=true, and productReleaseQualified=false. Cold p95 was 1293797ms, 6868234240 artifact bytes, and 3132882944 peak-RSS bytes; warm p95 was 823850ms, 6868246528 artifact bytes, and 3133816832 peak-RSS bytes. Every worker remained inside the unchanged 45-minute/20-GiB bounds with complete cleanup, and Android, edge, iOS, and service runtime each remained explicitly unsupported-product with releaseQualified=false. The append-only release-full chronology now retains all 20 scheduled or dispatched full runs through the closing success, including the explicit cancellation of redundant prior-revision schedule 31248899909; its canonical records digest is 9f38c2abba4f7ec8060cd5b7743ac8bd7a4718c383efdea81348d191b423c9d4 with 11 failures, 6 cancellations, and 3 successes. Independent watchdog run 31250800213 verified both retained ledgers, exact head and freshness, observed three canonical successful full runs, emitted zero violations, and preserved releaseQualified=false. This evidence remains valid for those runs but does not satisfy the later P1.6/P1.7 counterexamples; input: ci-release-full-closure-bundle-sha256:c75f96cacd842faa656845129de79ad49f41ec736fd4c5b2042d92a59c2a39d8. Reopened 2026-08-08 under P1.6: the closing standard run remains valid direct CI evidence, but the watchdog recognized exact-head push-fast and full histories rather than independently classifying that standard dispatch. Exact revision 6d22bc1e1eebee6ac76b8c20fb513f7d14b37d5b later demonstrated a failed standard run beside a successful full run and green watchdog. The correction required the standard classifier and negative controls, a passing exact-main standard/full pair, and independent watchdog recognition of both without rewriting the retained closure or failure chronology. Progress 2026-08-09 (local exclusivity correction; preserved R9.1.c input): the first local campaign proved that a point-in-time ps ... comm preflight could miss a shell-wrapped Quarto render and could not detect a competing build that began after admission, while still serializing competingLaneState=exclusive. Its six semantic observations remain preserved as superseded E0 diagnostics under ignored storage and cannot enter calibration. The corrected collector classifies complete process arguments for direct, shell, env, rustup, Deno, and Node wrappers; computes the measured workload’s complete descendant closure across nested process groups; polls once per second through workload termination; and kills, reaps, and retains any external match as a typed competing-lane hard failure with a positive maximum competitor count. Semantic passes and every other failure class require zero competitors. Twenty-nine collector controls, twenty-two independent calibration controls, closed schemas, static Python checks, and a bounded live macOS wrapper/ownership fixture pass locally. This collector work is preserved for R9.1.c; it is not an R0.4.j closure predicate and grants no SLO or release authority. Closure disposition 2026-08-10: P1.7 and P1.6 had already passed the operational acceptance boundary on exact reviewed main. Standard run 31284542152, release-full run 31284543591, aggregate 6c5c4e3da4ce4ee04fd744caaa7d088746aa814463061021d702a8a4244733b9, and watchdog run 31287266902 prove the bounded v0.2 DAG, cleanup, typed unsupported-target dispositions, and independent standard/full recognition while preserving productReleaseQualified=false. Section 3.23 assigns later multi-class SLO ratification to R9.1.c, closes this implementation task without weakening the provisional ceilings, and preserves every collected success and failure.

R0.5 Baselines and release hygiene

R0 exit criteria: a published clean checkout passes planning/topology/hygiene/complexity, root-lock, version, generated-artifact, warning-denied Clippy, no-user-panic static, and TCB static checks under declared tools on local and GitHub CI profiles; all checks are read-only; replay parity includes denied effects on native and WASI; the evidence verifier rejects every adversarial fixture; default tests neither nest aggregate pipelines nor contain stress/performance loops and remain stable under supported parallel load; default changed gates and generated-state caches meet GB budgets; generated status views exactly match the capability ledger.


R1. Agent Preview and preserved GenesisBench foundations: reproducible authoring and intent

Goal: make GenesisCode v0.3 a stable, self-contained target that the user’s AIs and other agents can learn, write, repair, test, and operate locally without hidden repository knowledge. Preserve completed GenesisBench protocol, fixture, scorer, and adapter foundations as non-expanding inputs; section 3.18 defers every unfinished campaign, custody, publication, governance, and Preview-release task until the GenesisCode v1 Self-Hosted Core release. Shared fixtures and schemas do not couple either product’s release authority.

R1.1 Versioned agent language profile

R1.2 Diagnostics as an agent API

R1.3 Warm daemon and generated MCP interface

R1.4 Training and evaluation corpus

R1.5 Authoring skill vNext

R1.6 Agent-safe collaboration

R1.7 Typed English intent and trajectory preview

R1-A GenesisCode Agent Preview exit criteria: GC-AGENT-v0.3 is frozen and versioned; compact cards meet AB-1/AB-2; diagnostics are structured across every user boundary; warm/MCP meets AB-3/AB-4 and survives cancellation, handle-cleanup, and leak stress; the generated authoring skill passes clean/offline installs and negotiates domain/target support without foreign-source escape; collaboration preserves transactional authority; typed English intent meets AB-11/AB-12 and emits closed product acceptance graphs without bypassing deterministic acceptance; pinned agent-product runs meet AB-5, AB-6, AB-8, AB-11, and AB-12 without private repository context. Shared public benchmark fixtures may validate this release, but private-epoch commissioning, a public leaderboard, or a Genesis Model package cannot block it.

R1-B GenesisBench Preview exit criteria, activated only after R8.3.b: GenesisBench binds exact immutable GenesisCode Core, Genesis Algorithms, and GenesisExamples/GenesisApps seed releases and publishes four isolated tracks, lineage-correct statistics, >=45 independently authored/governed and private-payload-verified Preview lineages, useful temporal overlays, a fixed reference scaffold, canonical CLI/adapters, self-hostable result registry, benchmark card/methods/first-results commissioning report, closed baseline bundles, and signed Preview-governance rules. The private epoch publishes immutable easy/medium/hard calibration, separately keyed context conditions, and at least two precommitted lineage-stratified fast forms whose results are directional and cohort-isolated. A common tokenizer-neutral artifact-text smoke/fast matrix runs through at least two independently implemented interface stacks and materially different model families; adapter conformance and terminal-owner derivation prove that provider, harness, extraction, unsupported-feature, and model failures remain distinct. The preserved historical Codex CLI/Luna xhigh reality gate remains labeled by its original identities; a new complete 27-condition follow-up and local-model cohorts are predeclared only after the seed and bind the immutable released language/library/catalog profile. Every expected cell is a replayable bundle or explicit missing record, mutable Foundry state and copied public solutions are absent, and the deployment invalid has a typed general language/card/tooling disposition with non-benchmark regression evidence. Published runs meet AB-7, AB-9, AB-10, BQ-1 through BQ-6, BQ-9 through BQ-16, and BQ-18 through BQ-20. BQ-7/BQ-8 remain hard blockers for training-corpus admission; BQ-17 activates only for reconstruction lineages against compatible L4 destination profiles.


R2. Runtime and resource foundation

Goal: finish the reference runtime, make memory/resource behavior part of semantics, and deliver a fast long-lived agent loop before adding aggressive execution tiers.

R2.1 Complete the compiled interpreter

R2.2 Heap and lifetime semantics

R2.3 Startup and incremental execution

R2.4 Runtime observability without semantic leakage

R2 exit criteria: PB-1, PB-4, PB-5, PB-6, PB-7, and PB-8 pass on normalized workloads; closure/tail/heap/handle stress tests pass; all exhaustion paths return sealed or explicit errors; incremental invalidation negative controls pass; treewalk and compiled modes retain PB-9 semantic identity.


R3. Validated execution tiers

Goal: gain predictable speed without enlarging the semantic TCB or fragmenting effects, errors, resource accounting, and canonical identity.

R3.1 Bytecode contract and verifier

R3.2 WebAssembly/stage2 closure

R3.3 Optimizer and tier controller

R3.4 Conditional JIT decision gate

R3 exit criteria: versioned bytecode and independent verification pass adversarial tests; PB-2 and PB-9 pass; stage2 coverage has no undeclared fallback; resource/capability behavior is identical across tiers. R3.4 may close with a documented “not required for v1” decision if product SLOs are met.


R4. Self-host semantic authority and bootstrap closure

Goal: make the toolchain above a minimal host genuinely GenesisCode-authored, independently verifiable, and reproducibly bootstrapped.

Demonstrable critical-path checkpoints: SH-A is the complete R4.1 truth/boundary package; SH-B is R4.2.a frontend/canonicalization at H2 on native and WASI with strict no fallback; SH-C is R4.2.b-e non-codegen semantic authority at H2 with legacy-route deletion controls; SH-D is R4.2.f-g plus R4.4.a-b and a hermetic one-reference-host stage2/stage3 rehearsal. Each checkpoint emits a runnable content-addressed artifact, ownership-ledger delta, differential and adversarial evidence, explicit remaining H-levels, and the next blocking edge. Only complete R4.4.c-e and R4.5 evidence can establish cross-host H3/H4 or close R4.

R4.1 Define self-host truth precisely

R4.2 Migrate semantic authorities

Each component follows the same sequence: normative corpus -> closed decision and call-site inventory -> GenesisCode implementation -> differential verifier -> production switch -> strict no-fallback profile -> Rust authority removal/demotion -> ownership-ledger update. A task cannot close while an applicable ledger row names only completed migration tasks but remains below its required level; every residual decision must name a still-open owner, an explicit stage0 disposition, or a separately justified non-applicability record.

R4.3 Reduce host complexity

R4.4 Bootstrap fixpoint and trusting-trust defense

R4.5 Replaceability and conformance

R4 exit criteria: every production semantic decision has an ownership-ledger row; no row below its required level points only to completed migration tasks; all required components reach H2; stage0 is contractually minimal; stage2/stage3 are byte-identical on tier-1 hosts; DDC and negative controls pass; legacy authorities are unreachable or removed; an independent verifier/conformance implementation validates release artifacts.


R5. Stable language and platform contract

Goal: freeze a coherent v1 language profile and complete the semantics needed by real agent-authored systems without uncontrolled surface growth.

R5.1 Audit and freeze existing semantics

R5.2 Deterministic structured concurrency

R5.3 Modules, contracts, and compatibility

R5.4 Standard library and capability surface

R5.5 Safe FFI and component extension

R5.6 Unified application, web, and UI model

R5.7 Batteries-included service, data, and AI platform

Service/data dependency boundary: R5.7 is a parallel capability family, not a scalar release gate. Secure Core service proof requires typed services, durable data, coordination, observability, and secure application foundations (R5.7.a-d,f). R5.7.e gates only data/numeric/ML profiles and products that consume it. A service, Core, benchmark, or research release cannot be delayed merely because an unrelated accelerator or ML backend is unfinished, and no data/ML pack may borrow the Core service proof.

R5.8 Games, simulation, graphics, and creative media

R5.9 Embedded language and hardware contract

R5 exit criteria: the v1 profile has no spec/implementation matrix gaps; effect rows and concurrency meet L3; numeric/text/path behavior is cross-host deterministic; stdlib APIs satisfy ownership/resource rules; compatibility and migrations work on the full corpus; extension negative controls pass; web/UI, service/data, game/media, and embedded profiles each have closed semantics, resource and capability models, generated agent cards, unsupported-feature diagnostics, and at least L2 reference implementations before target packaging may claim support.


R6. Packages, registry, builds, and deployment

Goal: let agents ship verifiable systems to the declared target matrix using infrastructure that can run entirely on the user’s own machines.

R6.1 Package and workspace completion

R6.2 Self-hosted registry trust

R6.3 Build target matrix

R6.4 Operations and lifecycle

Shared delivery authority boundary: R6.3-R6.4 complete the profile-parametric build, artifact, deploy, rollback, provenance, and incident authorities plus their declared native/WASI CLI and service Core cases. Their generic harnesses must reject unsupported or synthetic targets, but completing them does not qualify every listed platform. R6.5-R6.7 and the corresponding R8.2 proof own the authentic execution, lifecycle, generated-glue, and physical evidence for each independently released pack.

R6.5 Web and full-stack product delivery

R6.6 Desktop and mobile product delivery

R6.7 Embedded Linux, microcontroller, and board delivery

R6 exit criteria: deterministic offline resolution and package builds pass tier-1 hosts; a self-hosted registry supports publish/mirror/revoke/recover; target builds meet reproducibility, semantic-parity, authentic-execution, and one-language-source gates; native CLI/service, WASI/component, web static/interactive/SSR, OCI/system service, desktop, iOS, Android, Embedded Linux, and at least two materially different MCU families reach L4 or are explicitly excluded from the named release; deploy/rollback/secret/device negative controls pass; no descriptor, empty export, synthetic hash launcher, simulator-only result, or generated-glue edit is counted as product support.


R7. Assurance, formalization, and security

Goal: turn the trust story into independently checkable evidence focused on the smallest and highest-impact boundaries first.

R7.1 Layered test strategy

R7.2 Formal models and proofs

R7.3 Host and supply-chain hardening

R7.4 Reliability and scale

R7 exit criteria: all trust boundaries have threat models and mutation-tested negative controls; fuzz/sanitizer/property soak has 30 clean days with no unresolved P0/P1; required formal statements are checked and coverage limits are published; independent review findings are closed or explicitly release-blocking; supply-chain/recovery drills pass.

Core assurance boundary: M5-C/M6 require R7.1.a-e, R7.2, R7.3.a-e, and R7.4.a-c for the finite Core subjects. R7.1.f, R7.3.f, R7.4.d, and target-specific product/device evidence gate only the platform pack that claims those targets. A missing pack witness cannot block Core, GenesisBench, Foundry, or Genesis Model, and a Core result cannot satisfy a pack’s generator, supply-chain, physical, or lifecycle evidence.


R8. Product proof and adoption

Goal: prove GenesisCode is useful rather than merely feature-rich, especially for the user’s own AI systems.

R8.1 Reference agent integration

R8.2 Eighteen-archetype universal product gauntlet

R8.3 GenesisExamples and GenesisApps

This workstream refines the existing flagship-system allocation rather than creating a fifth product or a new semantic authority. GenesisExamples is the compact executable learning/reference catalog; GenesisApps is the maintained production-style showcase. R8.3.a-b form a finite post-Foundry seed checkpoint before GenesisBench or GenesisChallenge. R8.3.c-e continue toward universal platform maturity without blocking those downstream programs.

R8.4 Minimal human surface

R8.5 Benchmark-to-adoption and native-model flywheel

R8.5-A GenesisBench adoption lane

R8.5-B Genesis Model lane

R8.5-C GenesisChallenge contribution-contest lane

R8-C GenesisCode v1 Core proof lane: R8.1, R8.2.a-b, and R8.4 close when the relevant AB-5, AB-6, AB-8, AB-11, AB-12, and Core-applicable UP budgets meet v1 targets; independent agents complete the native/WASI CLI and authentic service loops without human code or foreign-source edits; the human quickstart/tooling meet their SLOs; and failures and unsupported domains are published. This lane alone may advance GenesisCode v1 Core to release candidacy.

R8-P Universal platform maturity lane: R8.3.a-b first publish the finite post-MF GenesisExamples/GenesisApps seed and unblock Bench/Challenge without claiming universal maturity. The remaining R8.2 work and R8.3.c-e close when all eighteen archetypes pass for their independently versioned packs; ten GenesisApps reach L4; seeded cross-target and physical-device maintenance succeeds; every one-language source audit passes; and independent operators reproduce each claim. This broader lane advances M5-P and individual pack releases, not the v1 Core, Bench, or Challenge release clock.

R8-B GenesisBench trust-release lane: After MF and M1-B, R8.5.a-e,s close when AB-7, AB-9, AB-10, BQ-1 through BQ-6, and BQ-9 through BQ-20 meet v1 targets on held-out, temporal, maintenance, authentic-target, reconstruction, stratified-difficulty, model-portable, and isolated-track tasks; exact compatible GenesisCode and Genesis Algorithms releases are bound; three model families, three independently implemented interface stacks, and two independent operators reproduce the common closed artifact-text standard runs; native/JSON-text tool encodings agree where supported; fast forms remain directional; benchmark-task/result governance and adoption evidence pass; and no language, library, Foundry-search, contest, provider, or model release evidence is borrowed. This lane alone may advance GenesisBench after the compatible language-and-library environment is frozen.

R8-M Genesis Model lane: R8.5.h,i,t,f,g,j-r close only after GenesisCode v1 Core, GenesisBench Trust, reproduced MF including the exact Genesis Algorithms release, BQ-7/BQ-8, and MQ-1 through MQ-8 pass; R8.5.t is the single readiness decision and every implementation task depends on it; every admitted artifact satisfies corpus policy; the Theseus/ASI/Foundry transfer manifest is independently reproduced and reviewed; the profile-bound model meets one declared Embedded Local class without cloud fallback; and applicable four-cell or frozen-language two-cell studies preserve cold general-model accessibility across every profile the package claims. This lane alone may advance the Genesis Model package. Portfolio completion may report all four outcomes plus Challenge operation, but no outcome borrows another lane’s evidence or waits for an optional release clock beyond its explicit prerequisites.

R8-X GenesisChallenge operating lane: R8.5.u-v may start only after R8.3.b, which itself follows F2.r, and close a contest season, not a product release. The lane has independent challenge admission, verification, credit, dispute, and destination-promotion authority; it never changes a GenesisBench score, a Foundry receipt, or algorithm-library/example/app acceptance and blocks no maintenance or later release. Benchmark-validity challenges remain ineligible until a compatible independently released benchmark baseline exists. Accepted contributions advance only the destination task or release whose ordinary evidence contract they independently satisfy.


R9. v1 Core Trust Release

Goal: freeze, reproduce, attest, publish, operate, and support a finite GenesisCode v1 Core without overstating any claim. Core includes the frozen language/runtime/package/evidence/agent/self-host authorities, native and WASI execution, and authentic CLI/WASI plus service product proof. Platform packs, GenesisBench, Foundry, and Genesis Model retain independent release clocks and identities; compatible versions may be named, absent, experimental, or unsupported without blocking Core.

R9.1 Freeze and release candidates

R9.2 Reproducible distribution

R9.3 Operations and governance

R9.4 Final acceptance

R9 exit criteria: every R9 task is complete; all M6 GenesisCode v1 Core claims have E4 attestations; independent reproduction, authentic native/WASI CLI and service execution, and verification pass; operational drills and all four public release-view workflows pass; GenesisCode is installable and usable offline without GenesisBench, Foundry, or a Genesis Model; public claims exactly match the Core and platform-pack ledgers; the one-language source audit passes for every Core proof artifact; and every linked pack, GenesisBench, scaffold, epoch, Foundry, or model identity is independently releasable and truthfully marked required, optional, compatible, incompatible, preview, or absent. Physical-device and wider platform witnesses are mandatory for the pack that claims them, not for Core.


F. Post-v1 frontier program

Post-v1 autonomous promotion is intentionally outside the v1 critical path. Sandboxed task generation, solver/verifier research, non-normative Foundry design notes, and R8.5.r controlled co-evolution experiments may run earlier, but they cannot complete frontier tasks, change compatibility, evaluator, corpus, model, trust, or release authority, or weaken maintenance of the released profile. Foundry ratification and execution begin from the preserved v1 semantic/evidence baseline, remain separately locked and removable, and earn expansion through calibration rather than scope optimism.

F1 Bounded self-improvement

F2 Genesis Foundry: proof-carrying algorithm discovery

Activation: F2.a-r Foundry Foundation work is non-start-ready until exact reviewed R9.4.f closes. It then consumes the frozen Core semantics and independent verifier interfaces without waiting for examples, apps, or GenesisBench. F2.r is the finite Foundry/Genesis Algorithms release gate required by the R8.3.a-b GenesisExamples/GenesisApps seed and Genesis Model transfer work; R8.3.b in turn gates GenesisBench Preview and GenesisChallenge. F2.s-x and F2.z may expand after F2.r under isolated resources; F2.y remains non-start-ready until a compatible GenesisBench Trust root R8.5.s also closes. Earlier material is design input only and creates no workspace, implementation WIP, research receipt, package release, example/app release, or claim.

Goal: establish a separately bounded computational-science system that composes typed mechanisms, searches only declared grammars, quotients only by certified relations, verifies candidates independently, compares correct candidates on honest Pareto frontiers, mines reusable abstractions, and proposes but never authorizes production changes. Convert independently accepted proposals into a supplied, stable, .gc-native Genesis Algorithms package through a separate destination authority. The initial objective is calibrated rediscovery, trustworthy negative evidence, and a useful algorithm-library foundation, not a claim that general algorithm invention or novelty detection is solved.

Boundary and authority foundation

Typed substrate and retained memory

Decisive bounded calibration

Calibration checkpoint: F2.a-F2.q are complete; FB-1 through FB-12 pass for the bounded calibration scope; the second operator reproduces the snapshot and every accepted/rejected receipt from declared inputs; all misses, counterexamples, resource costs, proof limits, and no-novelty/no-production claims remain visible; production builds pass with foundry/ absent. This checkpoint authorizes only destination review under F2.r; it does not authorize a package, model, optimizer, benchmark, or language change.

Promotion boundary and first supplied library

MF Foundry Foundation exit criteria: F2.a-F2.r are complete; FB-1 through FB-14 pass; the calibration snapshot is independently reproduced; every promoted and rejected package proposal is independently replayed; the public package is .gc-native, documented, versioned, installable, portable across its claimed Core profiles, and correct with deterministic fallbacks; no production, example/app, or Bench path depends on foundry/, mutable research state, or a search receipt as correctness authority. MF may feed R8.3.a-b and R8.5.i, but it does not authorize Bench, Challenge, a model, optimizer, language change, or F2.s-z result.

Expansion only after Foundry Foundation

F2 exit criteria: F2.a-F2.z are complete; MF Foundation was independently reproduced before expansion; FB-1 through FB-14 pass; every candidate and claim carries exact FD/proof status; Genesis Algorithms releases remain independently useful and production builds remain independent of foundry/; hostile candidates and false evidence fail closed; no model/search engine controls verification or promotion; broader packs and abstractions demonstrate measured value over simpler baselines; and any promoted artifact is destination-approved, reversible, compatibility-scoped, and independently verifiable from a closed bundle.

F3 Distributed and hardware-aware execution

F4 Ecosystem governance and replaceability


10. Sequencing and parallelization

GenesisCode:  R0 -> R2.2 -> R4.1 -> R5.1/5.3 -> R4.2.a-e -> R2/R3 prerequisites
              -> R4.2.f-g/R4.3-4.5 -> R1-A -> R5 Core -> R6.1-4 -> R7/R8 Core -> R9 v1 Core
                                                                                       |
                                     exact reviewed R9.4.f Core handoff               |
Foundry:                              F2.a-q calibration -> F2.r / MF + Algorithms v0.1
                                                               |
GenesisExamples/Apps:                  R8.3.a-b finite seed      |
                                                               |
               +-----------------------------------------------+-----------------------+
               |                                               |                       |
GenesisBench:  M1-B Preview -> M5-B                            |                       |
GenesisChallenge:                                              R8.5.u-v season         |
Examples/Apps expansion:                                       R8.3.c-e                |
Foundry expansion:                                             F2.s-z                  |
Packs/frontier:                                                                    pack releases + F1/F3/F4

Genesis Model: M5-B + MF + transfer/firewall/feasibility -> R8.5.t -> build/package
               | proposals return through deterministic GenesisCode acceptance

Practical order:

  1. Preserve the completed R0 publication/control-plane checkpoint. R0.4.k repaired transitive gate-input authority; P1.7 established bounded generated-publication ownership; P1.6 and R0.4.j closed on exact-main standard/full/watchdog evidence. Keep all incidents and timing observations, retain provisional hard ceilings, and defer class calibration to R9.1.c rather than rerunning green implementation gates.
  2. Close P1.5/R2.2.f through public cleanup-error propagation and one declared lifecycle/fault matrix per supported native host. Exercise success, malformed response, disconnect, cancellation, timeout, owner drop, restart, and repeated load in one bounded harness; do not rerun the surrounding workspace profile for every case.
  3. Complete R4.1’s stage0 contract, H0-H4 definitions, semantic-ownership ledger, truthful status, dependency firewall, and selected migration slice. This SH-A truth package must make every remaining host-owned semantic decision visible before migration work can claim progress.
  4. Freeze R5.1 and R5.3 against the correct current interpreter/resource semantics, then migrate R4.2.a-e by semantic risk and build leverage: frontend/canonicalization, type/effect checking, patch/refactor, obligation/policy, and package/registry/VCS. Each switch must make the GenesisCode implementation the reachable production authority and demote or remove the Rust semantic path; wrappers and parity-only demos remain H1 at most.
  5. Finish R2.1.h, R2.3, R2.4, and the R3.1-R3.3 execution prerequisites needed for self-hosted code generation. Performance tasks use one bounded normalized sampler rather than repeated whole profiles. Build bytecode only after semantic/resource contracts are fixed, and keep the optional JIT decision outside the self-host closure unless measured product evidence requires it.
  6. Close R4.2.f-g and R4.3-R4.5 end to end: compiler/optimizer/linker/build, CLI/agent orchestration, bridge generation, legacy removal, hermetic stage0->stage1->stage2->stage3 bootstrap, cross-host fixpoint, DDC, bootstrap witnesses, kernel/effect-host conformance, and an independent implementation or verifier. A .gc wrapper, matching local output, or self-issued receipt is not closure.
  7. Complete R1.3.f and R1.5-R1.7 against the self-hosted authority to release M1-A: compact negotiated authoring, collaboration, typed intent, cancellation, and warm/MCP behavior. Preserve prior GenesisBench bundles and invalids byte-for-byte, but do not run or expand a model campaign before v1.
  8. Complete the finite Core portions of R5.2/R5.4/R5.5 and R6.1-R6.4 against the self-hosted authority. Package resolution, registry operations, builds, deployment plans, documentation generation, and release assembly must operate self-hosted, offline where promised, with deterministic artifacts and no hidden Rust command semantics.
  9. Run applicable R7 assurance only after each subject reaches V3 readiness, then close R8.1, R8.2.a-b, and R8.4 Core product proof. Do not count labels, descriptor archives, empty exports, headless plans, or simulator-only demos as products. The native/WASI CLI and service workloads must be authentic, maintainable, independently operable, and authored under the one-language contract.
  10. Release M5-C and then R9/M6 only from independently reproduced E4 evidence. R9.1.c owns full timing calibration and release soak. The v1 Core release binds the exact stage0, self-host sources/artifacts, compatibility profile, package/build/docs/agent surfaces, supported hosts, support window, and rollback path. Unfinished platform packs, GenesisBench, Foundry, GenesisChallenge, and Genesis Model cannot block or lend evidence to this decision.
  11. At exact reviewed R9.4.f, publish the signed Core handoff and activate F2.a-r as the first post-Core product transaction. Allocate Foundry generator, independent verifier, package destination, and release roles separately; retain serial merge ownership for shared GenesisCode interfaces and generated authorities. Optional packs and unrelated F1/F3/F4 may use isolated resources, but Bench, Challenge, and Model remain frozen.
  12. Execute F2.a-q against frozen Core semantics, independently reproduce the calibration snapshot, then complete F2.r. Release genesis/algorithms v0.1 from destination-owned .gc sources, not the research workspace; bind exact contracts, fallbacks, profiles, provenance, proofs/tests, costs, compatibility, and rollback; prove clean build/install/use with foundry/ absent. Do not claim novelty or let the search engine verify or publish itself.
  13. At F2.r/MF, execute the finite R8.3.a-b GenesisExamples/GenesisApps seed before Bench or Challenge. Freeze the catalog/classification authority; publish the algorithm/CLI, bounded MCP server, production HTTP/WebSocket server, separate full-stack website/application, and integrated workspace; bind exact immutable Core and Genesis Algorithms releases; prove meaningful algorithm use, one-language source closure, authentic execution, clean build/deploy/replay/maintenance, and operation with foundry/ absent. Public examples are disclosed learning/reference inputs, never private benchmark answers.
  14. At R8.3.b, reissue the frontier for isolated GenesisBench, GenesisChallenge, broader R8.3.c-e application, and continued Foundry-expansion lanes. In GenesisBench, bind exact immutable Core, Genesis Algorithms, and public catalog releases and complete Preview R1.4.o/q/r/p, then Trust R8.5.a-e,s. Predeclare Qwen, local-model, Luna, and future-model campaigns; require private custody, difficulty calibration, fast forms, model-neutral adapters, failure attribution, independent operators, three model families and interface stacks, construct validity, adoption evidence, and reproducible standard/full forms. Never expose mutable Foundry state, copy public solutions into held-out payloads, rewrite historical attempts, or use scores to mutate the evaluated language/library/example profile.
  15. Concurrently after R8.3.b, establish R8.5.u and operate R8.5.v under the independent non-ranking GenesisChallenge ledger. Open only categories with ready verified destination baselines; include creation, target expansion, maintenance, simplification, and optimization of GenesisExamples/GenesisApps, while routing algorithm-library and Foundry improvements only through destination-owned acceptance. Defer benchmark-validity challenges until a compatible Benchmark release exists. The contest can stop without delaying any product or research lane.
  16. Continued Foundry F2.s-x and F2.z may proceed after F2.r under isolated resources and prior-release baselines; broader R8.3.c-e may proceed after the seed as relevant packs reach authentic support; F2.y waits for compatible GenesisBench Trust because it crosses the Bench/model integration boundary. Optional platform packs and F1/F3/F4 may also advance under disjoint resources and compatibility-scoped releases. Every later useful Foundry output crosses only through the F2.r/F2.z package or destination-promotion firewall.
  17. After both MF and GenesisBench Trust, freeze clean Theseus/ASI editions, bind accepted and rejected Foundry Foundation and algorithm-library transfer records plus the exact example/app catalog used for evaluation, and complete R8.5.i. Ratify R8.5.t only after the adversarial firewall, legal/privacy/compute/release feasibility, and independent review are also closed. No Genesis Model corpus, teacher, tokenizer, architecture, training, distillation, package, or integration implementation begins earlier.

Workstream dependency matrix:

Workstream Hard prerequisites May proceed in parallel Hard-blocks
R0.1 truth sources none R0.3 prerequisite discovery, R7 formal-model scoping every generated status/release claim
R0.2 evidence lifecycle R0.1 authority decisions R0.3-R0.5 M0, every L3-L5 promotion
R0.3 hermetic versions R0.1 source-of-truth rules R0.2, R0.4 reproducible checks, profiles, R1 card/version negotiation
R0.4 gate/resource architecture R0.1 ledger and R0.3 tool identities R0.2, R0.5 affordable continuous evidence and release execution
R0.5 normalized baselines R0.3 host/tool profiles and R0.4 telemetry schema late R0.2 verifier work every performance claim and JIT decision
R1.1 profile/cards R0 version/compatibility registry R1.2 diagnostic catalog R1.3 schema generation, R1.4 corpus, R1.5 skill
R1.2 diagnostics R1.1 profile IDs and R0 evidence schemas R1.1 cards R1.4 repair benchmark and M1-A/M1-B
R1.3 warm/MCP R1.1 schemas; R2.2 semantics required before final daemon sign-off R1.4-R1.6 prototypes agent product loop, AB-3/AB-4
R1.4 GenesisBench foundations and post-v1 Preview Completed foundations retain R1.1/R1.2 identities; every unfinished inference, custody, publication, portability, and governance task additionally requires F2.r and its exact Core/library identities No pre-v1 or pre-MF product lane; model-free correctness/security maintenance may accompany its owning GenesisCode gate Post-MF M1-B and R8.5 temporal benchmark/corpus candidates
R1.5 skill generated R1.1/R1.2 inputs, completed R1.4.n fixtures, and green R0.4.k authority R1.4.o/q/p and R1.6 M1-A agent readiness
R1.6 collaboration R0 evidence/identity; R1.3 transactional API for completion R1.4-R1.5 safe parallel-agent proof
R1.7 typed intent preview R1.3 transactional API and R1.6 authority separation R1.4-R1.5 AB-11/AB-12, responsible real-request trajectory capture, R8.5 specialist integration
R2.1 interpreter R0 normalized semantic/perf corpus R2.2 heap specification R2.3 snapshots, R3 execution tiers
R2.2 heap/resources R0 evidence and resource schemas R2.1 final R1.3 daemon sign-off, all R3 tiers
R2.3 startup/incremental R2.1 stable compiled representation and R2.2 accounting R2.4 AB-3/AB-4, R3 artifact caches
R3 bytecode/Wasm/tiering R2 semantic, resource, startup, and observability contracts R4.1 ledger, R5.1 semantic audit R4.2 compiler authority, M2, and the v1 JIT-or-no-JIT decision
R5.1 and R5.3 core freeze R0 compatibility registry; R2/R3 discrepancies resolved R4.1 ownership mapping every corresponding R4.2 authority switch
R4.1 ownership/TCB R0 truthful ledgers R2-R3 and R5 core audit all R4.2 migration acceptance
R4.2-R4.5 self-host closure relevant R5.1/R5.3 contract frozen; R3 verified artifact path remaining R5.2/R5.4/R5.5 implementation M3 and trusted self-hosted R6 orchestration
R5.2/R5.4/R5.5 completion R5 core profile plus R2/R3 resource semantics late R4 migration R6 compatibility and target claims
R5.6 web/UI and R5.7 service/data R5.1/R5.3 core, R5.2 concurrency, R5.4 stdlib, R5.5 extension boundary late R4 migration and R5.8/R5.9 R6.5 web/full-stack and shared product foundations
R5.8 game/media R5.1/R5.3 core, R2 resources, R3 tiers, R5.4/R5.6 R5.7 and R5.9 real interactive target/product claims
R5.9 embedded contract R5.1 numeric/text/profile semantics, R2 resources, R3 validated AOT inputs, R5.4/R5.5 R5.6-R5.8 R6.7 firmware and board claims
R6.1-R6.4 ecosystem/core delivery R5 compatibility/ABI freeze; self-host build authority for GA R7 continuous assurance target packaging, registry, and operations authority
R6.5 web/full-stack delivery R5.6/R5.7 and R6.1-R6.4 R6.6/R6.7 authentic web/full-stack flagships
R6.6 desktop/mobile delivery R5.6/R5.8 and R6.1-R6.4 R6.5/R6.7 authentic native/mobile flagships
R6.7 embedded/MCU delivery R5.9, R5.7 IoT services, R6.1/R6.3/R6.4 R6.5/R6.6 authentic Embedded Linux/board flagships
R7 assurance formal models and test scaffolding may start with their GenesisCode subject; release claims require frozen subjects the selected Core task only before R9.4.f; isolated lanes afterward every applicable M5-C/M5-B/M5-M gate and M6
R8-C GenesisCode proof M1-A interfaces; relevant R2-R7 features at least L3; authentic target and one-language source gates for claimed profiles no other repository-changing product lane before R9.4.f finite GenesisCode v1 Core acceptance through R8.1, R8.2.a-b, R8.4; separate M5-P maturity through the remaining R8.2-R8.3 tasks
R8.3 GenesisExamples/GenesisApps MF/F2.r for catalog start; R8.3.b additionally requires authentic Core, MCP, service, browser, and full-stack product profiles; R8.3.c-e require the corresponding R8.2/pack subjects continued Foundry expansion after F2.r; broader apps, packs, and F1/F3/F4 under isolated ownership finite public seed for Bench/Challenge, then M5-P universal showcase; never a semantic authority or Core prerequisite
R8-B GenesisBench trust release MF/F2.r; finite R8.3.a-b seed; M1-B; R1.6 authority; R7.1 assurance; R8.1 agent integration; exact compatible Core, Genesis Algorithms, and public catalog releases continued Foundry and R8.3.c-e expansion, Challenge, packs, and the independently gated model lane GenesisBench scope, independent reproduction, governance, and adoption; never blocks GenesisCode v1 or later Foundry/application maintenance
R8-M Genesis Model lane stable v1 Core; GenesisBench Trust; BQ-7/BQ-8; reproduced MF; MQ-1 through MQ-8; R8.5.i; independent R8.5.t decision post-v1 platform work and optional Foundry expansion independently reviewed research-transfer manifest, profile-bound model package, and Embedded Local evidence; never blocks GenesisCode, packs, or GenesisBench
R8-X GenesisChallenge contest R8.3.b plus R8.5.u policy/ledger authority; benchmark-validity category separately requires a compatible benchmark release post-seed product and research lanes under disjoint files/authorities one independently audited contest season and destination-owned accepted contributions, including example/app expansion; blocks no release
R9 GenesisCode Core release R0-R7 plus R8.1, R8.2.a-b, R8.4, and zero Core release blockers no other repository-changing product lane before R9.4.f; the handoff activates Foundry Foundation first GenesisCode v1.0 Core
F1 bounded improvement preserved v1 compatibility/governance baseline F2 specification and independent F3/F4 research no v1 milestone
F2 Genesis Foundry + Genesis Algorithms signed v1 Core semantic/evidence interfaces; F2.a-q calibration precedes F2.r independently promoted library release; F2.r precedes R8.3.a-b; F2.s-z are later expansion; F2.y additionally requires compatible GenesisBench Trust packs and independently sandboxed F1/F3/F4 during Foundation; examples/apps and later expansion after F2.r; Bench/Challenge after R8.3.b MF Foundation and Genesis Algorithms v0.1, then destination-specific proposals/releases; never a Core or production-runtime dependency; MF is an input to examples/apps and optional model readiness
F3 distributed/validated execution preserved v1 semantics plus relevant R3/R6/R7 profiles F1/F2/F4 under separate budgets optional post-v1 target/profile promotions
F4 ecosystem governance preserved v1 compatibility and independent verifier baseline all frontier lanes standards/replaceability decisions, never retroactive authority

11. Migration from the 2026-07-02 roadmap

No substantive objective from the prior R0-R11 plan is intentionally discarded. It is reordered and given stricter acceptance semantics:

Prior area New home Change
Prior R0 foundation repair R0 Reopened and expanded because root-lock hermeticity, evidence mutation, status accuracy, and gate cost remain unresolved.
Prior R1 interpreter overhaul R2 Retains completed inline-int/vector/parser/primitive work as baseline; focuses open work on closures, frames, maps, tails, heap, and startup.
Prior R2 bytecode/JIT R3 Bytecode remains required; JIT is conditional on measured need rather than assumed critical path.
Prior R3 warm/MCP/agent loop R1-R2 Moved before bytecode because it is the shortest path to useful AI adoption; warm exists and is hardened rather than rebuilt.
Prior R4 self-host closure R4 Strengthened with H-levels, semantic authority, DDC, independent conformance, and explicit stage0 truth.
Prior R5 AI completeness R1 and R8 Compact SDK/training readiness moves early; broad gauntlet and flagship proof remain late product gates.
Prior R6 language completeness R5 Existing effect rows/concurrency are audited/promoted instead of described as absent.
Prior R7 deployment R6 Adds offline/self-hosted infrastructure, profile compatibility, reproducibility, and operations.
Prior R8 ecosystem R6 and R9 Registry/package GA and release governance are separated.
Prior R9 verification R7 Adds mutation evidence, threat models, fault injection, DDC, and independent review.
Prior R10 human UX R8 Kept minimal and schema-generated, tied to real flagship workflows.
Prior R11 self-improvement F1-F4 Moved post-v1 so recursive automation cannot outrun evidence, governance, and release stability.

Completed work from the prior plan is preserved in git history and the audited baseline. It must be reclassified in the capability ledger rather than copied as unchecked green claims. If a prior task has durable evidence and satisfies the stronger definition of done, it may be marked at the corresponding new maturity level without reimplementation.


12. Milestone acceptance summary

Milestone Required phases Headline acceptance
M0 Truthful Green Door R0 Published clean clone and live remote CI; clean, hermetic, read-only checks; replay parity; load-stable default suite; warning-free build; bounded gate/cache disk and time; non-starving event/profile concurrency; bounded runner-unavailable disposition; independently watched standard/full evidence freshness; complete transitive generated/gate-input freshness; independent evidence verifier
M1-A GenesisCode Agent Preview R1.1-R1.3, R1.5-R1.7; R0.4.j Versioned compact SDK; structured diagnostics; bounded warm/MCP; generated skill; safe collaboration; typed intent; clean/offline agent use without benchmark/model release coupling
M1-B GenesisBench Preview F2.r/MF, then remaining R1.4.o/q/r/p First post-v1 benchmark release against immutable Self-Hosted Core plus Genesis Algorithms; four-track profile; transport-neutral artifact-text baseline; two independent interface/model-family portability proof; typed failure attribution; >=45 independently authored/governed and private-payload-verified held-out lineages; immutable easy/medium/hard calibration; precommitted lineage-stratified fast forms; fixed scaffold; CLI/adapters; registry; protected temporal overlays; complete newly predeclared commissioning cohorts; benchmark card, methods, first-results report, and Preview governance; no mutable Foundry state
M2 Runtime Beta R2-R3 core Resource-bounded runtime; deterministic snapshot/incremental loop; verified bytecode; semantic tier parity
M3 Self-Host Authority Beta R4 H2 toolchain authority; H3 cross-host fixpoint; DDC; independently verifiable stage0 contract
M4 Platform Pack Betas relevant R5-R6 workstream per pack Stable shared profile and independently versioned web/UI, service/data, desktop/mobile, game/media, Embedded Linux, MCU/board, and GPU/XR packs; each claimed target has authentic reproducible execution and one-language source closure
M5-C GenesisCode v1 Core Candidate R7.1.a-e, R7.2, R7.3.a-e, R7.4.a-c; R8.1, R8.2.a-b, R8.4 Formal/fuzz/security gates; frozen Core; authentic native/WASI CLI and service proof; one-language source closure; independent operation and maintenance; no R5.7.e, R7.1.f, R7.3.f, R8.2.c-r, or R8.3 dependency
M5-P Universal Platform Maturity R8.2-R8.3 18/18 authentic archetypes for claimed packs; ten maintained flagships including physical boards; independent operation and maintenance; may complete after M6
M5-B GenesisBench Trust Release MF, M1-B, then R8.5.a-e,s >=90-lineage independently reproduced temporal/maintenance/reconstruction benchmark over exact compatible Core/library releases; orthogonal difficulty/context/run profiles; common artifact-text standard matrix across three interface stacks/model families; native/JSON-text tool equivalence; smoke/fast/standard/full forms; signed public governance; discriminative validity; external benchmark-task/result operation and language/library adoption evidence
MF Genesis Foundry Foundation R9.4.f, then F2.a-r Independently reproduced bounded rediscovery with frozen Core semantics, external checkers, complete lineage/cost evidence, followed by destination-owned Genesis Algorithms v0.1 promotion, clean package rebuild/use with foundry/ absent, and no model, novelty, or production-authority claim
R8-X GenesisChallenge Season F2.r, then R8.5.u-v Independently governed non-ranking contest with ready language/library/Foundry destination baselines, signed attribution, adversarial controls, and destination-owned promotion; benchmark-validity category waits for a compatible Benchmark release
M5-M Genesis Model Preview post-MF/post-Trust R8 Genesis Model lane Stable Core, exact Genesis Algorithms, and GenesisBench Trust; reproduced MF; audited firewall; independent R8.5.t decision; admitted lineage-preserving corpus; clean pinned Theseus/ASI/Foundry transfer evidence; profile-bound Embedded Local model; applicable four-cell or frozen-language two-cell accessibility proof; rollback
M6 GenesisCode Trust Release R9 finite Core Reproduced signed GenesisCode v1 Core, E4 evidence, authentic CLI/WASI and service proof, offline verification/install without benchmark/research/model dependencies, compatibility/security operations, and independently scoped pack manifests

13. Risk register

Risk Why it matters Mitigation and trigger
Status theater Broad routing/tests can look complete while semantic ownership or cross-host proof is absent L0-L5 and H0-H4 ledgers; generated matrices; independent verifier
Local-only accumulation Large dirty worktrees are neither bisectable nor protected from accidental loss, and CI changes remain unexercised Task-scoped green commits, remote draft checkpoints, clean-clone reconstruction, required GitHub CI
Flaky assurance Load races and recursively nested suites turn trust checks into timing lotteries or encourage ignored failures Hermetic lane ownership, explicit fixture readiness, repeat-under-load controls, no retry-to-green acceptance
Validation theater Repeating an unchanged green aggregate or qualifying a stub consumes the implementation lane without changing a decision V0-V4 readiness ladder, one identical development success, autonomous decision-bearing campaigns, full calibration only at R9.1.c
CI evidence darkness Local/PR green can mask cancelled default-branch runs, stale full profiles, or jobs queued forever for absent self-hosted labels R0.4.j/P1.6 event-aware concurrency, hosted runner preflight, typed optional-profile nonclaims, disjoint watchdog, bounded standard/full freshness, append-only incident disposition
Generated-view partial refresh A canonical or transitive gate-input change can leave one protocol, manifest, card, schema catalog, index, site page, or model context silently stale while the graph itself reports fresh R0.4.k gate-input discovery, atomic transitive closure, one explicit updater, read-only freshness checks in local and remote gates, Rust/Python/helper/fixed-point dependency-completeness controls
Ignored output masks source drift Quarto renders, caches, build trees, simulator state, or benchmark output can pollute topology scans or hide untracked source/evidence behind broad ignore rules Explicit source/generated root classification, bounded generated-state inventory, cleanability proof, ignored-source canaries, and source-only hygiene scans derived from policy rather than ad hoc exclusions
Benchmark-shaped execution Workload recognizers can manufacture impressive numbers without improving the language runtime Remove production recognizers, anti-overfit source variants, property-based optimization preconditions, differential semantics
Roadmap scope overwhelms delivery A single maintainer can spend years on breadth before agents can use the language M1-A is the first language product and M1-B the independent benchmark wedge; critical paths block speculative breadth; each milestone is independently useful
Wrong post-Core serialization Starting Bench before a canonical supplied library and useful public applications loses Foundry and product-proof leverage, while requiring open-ended F2.a-z or all ten universal apps would block adoption indefinitely Exact R9.4.f -> finite F2.a-r Foundation/library gate -> finite R8.3.a-b example/app seed -> Bench/Challenge fan-out; continued F2.s-z and R8.3.c-e remain independent; mutation controls reject bypasses and accidental expansion ancestry
Optimizations change identity/effects Breaks the core trust model Reference semantics, differential corpus, translation validation, versioned formats, mutation tests
JIT dependency and security cost Can harm startup, platform reach, build size, and TCB without helping agent workloads R3.4 decision gate; bytecode/wasmi completeness; optional feature
Memory leaks or host OOM Warm agents and services become unsafe despite step/effect limits R2 heap semantics, logical metering, isolation, 100k-request stress, hard cleanup
Self-hosting duplicates rather than removes trust Two implementations increase attack and maintenance surface Ownership ledger; production authority switch; strict no-fallback; Rust demotion/removal
Self-host migration targets moving semantics Switching authority before a profile freeze causes duplicate churn and makes parity evidence ambiguous Freeze the relevant R5.1/R5.3 contract before each R4.2 production switch; version any later semantic change
Bootstrap trusting-trust attack A fixpoint alone can reproduce a malicious compiler DDC, independent builders/verifiers, signed source/evidence, conformance implementations
Agent benchmark contamination Inflates quality claims and trains to tests Temporal-clean labels only for post-release precommitments; split manifests, hidden oracles, commitments, rotation, canaries, contamination labels, model-agnostic scoring, immutable history
Statistical pseudoreplication Counting three context conditions of one task as three independent successes creates false precision and unstable model rankings Immutable lineage IDs, clustered/hierarchical analysis, BQ-9 denominators and intervals, predeclared statistics, no unqualified pairwise decimal ranks
Benchmark track conflation A raw model, heavily scaffolded agent, adapted specialist, and local appliance answer different questions; one rank would be uninterpretable Four closed eligibility tracks, scaffold/profile/hardware cohort keys, separate leaderboards and claims, BQ-10 negative controls
Small, saturated, or author-biased benchmark set Public-small near-ceiling results can erase frontier discrimination while a few private tasks, placeholder commitments, or one generator dominate rankings and teacher selection BQ-11 custody-verified 45/90 payload floors, per-class/difficulty balance, independent author/governor identities, family weight cap, medium/large priority, rolling epochs, stratified audits; one commissioning cohort never alone triggers BQ-5
Difficulty/context conflation and fast-set selection bias Calling long context hard or selecting a cheap subset after seeing model scores can manufacture gradients, hide capability holes, and produce unstable ranks BQ-15/BQ-16 orthogonal keys, non-target calibration, precommitted lineage-stratified fast forms, alternate forms, visible omissions/uncertainty, standard/full claim gates, immutable epoch labels
Provider/scaffold overfit masquerades as model quality A benchmark built around one API, chat template, tokenizer, tool-call encoding, reasoning control, or output convention can reward interface familiarity and reject otherwise capable models R1.4.r transport-neutral authority, common artifact-text profile, capability negotiation, native/JSON-text tool metamorphism, tokenizer-neutral budgets, BQ-18/BQ-20, differential item review
Harness failure is blamed on the model Request rejection, hidden prompt injection, adapter extraction defects, unsupported context, or broken workspaces can depress a model’s score without testing GenesisCode reasoning Closed terminal-owner taxonomy, model failure only after delivery/capture/extraction/workspace proof, independent rederivation, adversarial blame-laundering controls, BQ-19
Benchmark gaming distorts the language Task-specific APIs, runtime recognizers, or easier suites can improve rank while reducing general value Useful-work task criterion, anti-overfit variants, construct-validity audits, saturation ratchet, independent challenge authors, no score authority in production semantics
Contribution contest captures benchmark or release authority Open repository work has moving baselines and visible tests; mixing it with held-out evaluation enables self-scoring, Goodhart optimization, collusion, and prize-driven unsafe merges R8.5.u-v separate GenesisChallenge ledger and roles, no leaderboard bridge, hidden adversarial controls, destination-owned review/promotion, append-only failures/disputes, vector rather than scalar deltas
Native-model co-adaptation A specialist and language may improve only together while GenesisCode becomes harder for every general model and independent implementation Permanent Cold Acquisition track, applicable four-cell or frozen-language two-cell R8.5.r controls, total-agent-work accounting, compatibility/TCB costs, reject combined-only gains
Research-transfer cargo cult Appealing Theseus implementations or ASI synthesis can become architecture authority without surviving Genesis-specific replication, laundering claims and complexity into the model MQ-1 through MQ-8, R8.5.i clean editions and crosswalk, simplest-baseline/ablation gates, independent decisions, preserved null/negative/rejected outcomes
Premature specialist training Training before corpus diversity, lineage, consent, and held-out isolation wastes compute and irreversibly contaminates evaluation BQ-7/BQ-8 hard gate, tokenizer/base audit first, quarantine-only R1.7 trajectories, checkpoint manifests, no active-epoch selection
Teacher-output feedback collapse Training only on top-model answers can amplify common errors, erase diversity, and contaminate evaluation BQ-7/BQ-8 admission and firewall, lineage/time/model-family splits, deduplication, verified counterexamples, teacher diversity, independent held-out evaluation
Context bundle drift Agents learn APIs that no longer match runtime behavior Generated cards/symbol index, parse/run goldens, profile negotiation, stale-card rejection
Gate explosion and disk growth Slow checks discourage use and can consume tens of gigabytes Gate manifest, impact selection, shared caches, GB budgets, deterministic cleanup
Capability convenience erodes security Agents may solve errors by asking for broad authority Structured repair rules, policy diffs, minimization score, negative controls, independent review
Cross-platform nondeterminism Breaks hashes, replay, builds, and bootstrap Numeric/text/path specs, tier-1 host matrix, normalized archives, exact mismatch reports
Registry/signing compromise Invalidates ecosystem trust Offline roots, threshold/rotation/revocation, transparency, mirrors, rehearsed recovery
Formal work becomes decorative Proofs may cover a toy subset while claims imply full runtime Coverage ledger tied to executable forms; published assumptions/gaps; proofs hard-gate only their stated claims
Public claims outrun product proof Damages credibility and user trust E4-linked claims, explicit unsupported domains, flagship maintenance exercises
Artifact-suffix theater An .ipa, .aab, .wasm, app bundle, firmware image, or launcher can look complete while containing no executable product logic UP-2 authentic-runtime execution, R6.3.f real install/launch/effect workloads, target-native inspectors, independent builders, and explicit descriptor labels
Generated glue becomes a second source language Agents may be forced to patch HTML/CSS/JavaScript, Swift/Kotlin, C/linker scripts, shaders, shell, or YAML, defeating the one-language promise and reproducibility UP-1 inventories, read-only disposable generated roots, source maps, deterministic regeneration, R8.3.d deletion/rebuild audit, fail instead of manual repair
Universal scope collapses delivery Web, mobile, games, data, and boards can turn one language into many unfinished frameworks Shared app/data/capability/extension contracts, profile promotion by L-level, authentic flagship gates, package-first breadth, no kernel growth for domain convenience, explicit unsupported cells
Platform semantics leak into the kernel UI, cloud, device, or game convenience forms could bloat the TCB and make portability impossible Keep product APIs in libraries/effects/components, use typed target profiles, preserve pure shared logic, require compatibility and proof-cost review for syntax/semantic changes
Embedded profile silently changes GenesisCode Fixed-width arithmetic, static allocation, interrupts, and no-OS constraints can diverge from host semantics while still using the same name R5.9 explicit subset/profile, compile-time rejection, translation validation, target-specific numeric/fault rules, cross-profile corpus, never claim full-host equivalence
Simulator-only hardware confidence Mocks can pass while pin maps, timing, electrical behavior, resets, radio, flash, or OTA fail on physical boards R6.7 hardware-in-the-loop, named boards, repeated power/peripheral fault tests, measured envelopes, independent device witnesses, simulator evidence remains separate
Ecosystem escape-hatch erosion Missing packages may push agents toward arbitrary native code or broad process plugins R5.5 component-first adapters, signed high-risk native profiles, generated bindings, package provenance, explicit capability/resource isolation, application tree foreign-source audit
Single-project review authority One maintainer, key, host, or organization can accidentally validate its own benchmark, model, and release claims BQ-14 independent reviewer/operator, conflict records, role-separated custody, R9.2.c threshold attestations, and independently controlled mirrors
Foundry combinatorial/resource runaway Search can consume unbounded compute, disk, energy, or review effort while producing little knowledge Typed finite grammars, FB-7/FB-11 budgets, grammar-relative completeness, checkpoints, deterministic pruning evidence, per-pack admission, and automatic bounded terminalization without automatic promotion
Foundry verifier or proof laundering A candidate may pass a weak finite test or theorem about a simplified model and be advertised as generally correct emitted code Orthogonal FD/proof status, F2.j independent checker, FB-4/FB-9 mutations, exact artifact-to-theorem/lowering bindings, unknown status, and no promotion from benchmark evidence alone
Foundry law/catalog poisoning One false equality or over-broad mechanism can contaminate many later extractions and production proposals Quarantined speculative laws, stronger law obligations, versioned catalogs, dependency tracing, independent replay, rollback/canaries, append-only supersession, and zero direct mutation of gc_opt
Foundry novelty inflation and benchmark overfit New hashes, hidden hardcoding, or rediscovered literature can be mislabeled as invention Calibration-first MF, anti-hardcoding tasks, honest claim taxonomy, local/literature similarity limits, external review before novelty, complete negative-result publication, and regime-scoped claims
Foundry dependency/authority creep Convenient research code could become a hidden compiler/runtime dependency or allow a generator to approve itself One-way workspace architecture test, production build with foundry/ absent, separately locked releases, role separation, proposal-only bridges, and destination-owned semantic-patch promotion
Coupled product release clocks Waiting for model training can delay a ready language, while shared version numbers can make benchmark/model evidence appear authoritative for language semantics Independent signed release manifests, compatibility ranges, product-scoped blockers, optional model integration, and separate GenesisCode/GenesisBench/Genesis Model acceptance lanes

14. Maintenance rules

  • Keep Last audited current whenever priorities, facts, budgets, or release criteria change.
  • Never check a task without the full done annotation and durable evidence described in section 2.2.
  • Never delete completed work to make the queue look smaller. Move historical detail into a versioned evidence/changelog view when this file becomes unwieldy.
  • Generate status matrices from the capability/evidence ledger after R0; do not manually synchronize multiple truth sources.
  • Ratchet budgets downward. A loosening requires a dated decision record and evidence.
  • Promote repeated active failures to upgrade_plan.md as P0/P1; close them there only when their regression gate passes.
  • Keep governed check entrypoints one-in/one-out until M6 and the later M1-B activation. Prefer extending an existing authority and record consolidation savings; never add a gate merely to assert that too many gates exist.
  • Keep the R9.4.f GenesisCode Core closure free of unfinished curated Examples/Apps, Benchmark, Challenge, Foundry, Model, universal-pack, and pack-only assurance work. At the signed handoff, activate F2.a-r Foundry Foundation first; after F2.r, execute finite R8.3.a-b while continued Foundry, pack, and unrelated frontier lanes may remain isolated; after R8.3.b, activate Bench and Challenge without waiting for R8.3.c-e. Do not serialize unrelated lanes or parallelize shared-authority merges.
  • Keep program-level concepts frozen through R9.4.f as section 3.19 defines. A new product, program, benchmark track, research lane, milestone family, task family, or governed gate family requires a reproduced P0/P1 need and a one-for-one retired equivalent with explicit dependency and acceptance deltas; otherwise retain it only as a non-authoritative note.
  • Keep a published remote checkpoint for every coherent green tranche. An unpushed local worktree is not durable project state.
  • Regenerate declared views as one transitive closure. A canonical-input change may not be committed or published with a partial card, index, schema, site, context-bundle, or evidence refresh; the updater is explicit and all checks remain read-only.
  • Treat every governed gate’s recursively discovered source/helper/test identity as generated-authority input unless a reviewed fixed-point exclusion proves why it cannot be. A graph that reports fresh while any governed check fails only for stale identity is itself stale and blocks publication.
  • Keep source and generated state disjoint under one reviewed policy. Ignored or cleanable output roots may not contain authoritative source/evidence, and adding an exclusion requires a bounded-state owner plus a canary proving untracked source still fails hygiene checks.
  • Treat test-lane membership as governed data. Moving a test to stress, performance, platform, release, or ignored status requires a manifest diff, rationale, owner, trigger, expected frequency, and proof that required CI executes it; ignored tests never silently disappear from published counts or coverage claims.
  • Treat CI control-plane liveness as release evidence. PR supersession may cancel only the same PR’s obsolete run; pushes, schedules, and dispatches retain a terminal typed disposition, impossible self-hosted label sets are rejected by a hosted preflight before queueing, and a disjoint watchdog must detect stale or cancelled-only standard/full history without trusting the workflow it observes.
  • Never describe a task as outside a model’s training data merely because GenesisCode is new. Use the exact BQ-3 label supported by model-release, task-precommitment, custody, and disclosure evidence.
  • Benchmark task changes are append-only by epoch: preserve historical weights/results, precommit new private material before model access, disclose or rotate under policy, and never tune hidden tasks from a target model’s observed failures.
  • A benchmark condition is never promoted to an independent lineage because context, tools, scaffold, repair budget, prompt wording, or a semantics-preserving mutation changed. Preserve the parent lineage and use BQ-9 clustered analysis.
  • Keep benchmark axes orthogonal. Difficulty, context size, run profile, task class, capability/product surface, track, language profile, scaffold/adapter, attempt policy, and hardware cohort are independently keyed; changing one never silently changes or substitutes for another.
  • Keep the benchmark authority transport-neutral. Vendor SDK payloads, chat templates, special tokens, native tool encodings, reasoning controls, tokenizers, CLIs, and hosted services belong only in adapters or disclosed Open Agent scaffolds; adding a conforming model never requires changing task semantics or the scorer.
  • Preserve a minimum-common-denominator artifact-text condition for every eligible non-interactive lineage. Native tools, JSON mode, streaming, multimodality, and provider features are separate interaction conditions, never admission requirements or hidden advantages in the common cohort.
  • Attribute before scoring. No attempt is a semantic model failure until request delivery, response capture, adapter/scaffold conformance, neutral candidate extraction, workspace integrity, and model-independent scoring are independently rederived; ambiguous ownership remains invalid and non-claiming.
  • Bind canonical context in bytes and structure, not one tokenizer. Record provider tokens separately, never truncate or rewrite per model inside a cohort, and treat unsupported complete context as an explicit capability outcome with a separately predeclared smaller condition where useful.
  • Freeze every fast form before target-model access using a declared lineage-stratified algorithm. Publish its complete sampling frame, omissions, uncertainty, and alternate forms; fast evidence is directional and never closes standard/full, Trust-release, universal-coverage, or strongest-model claims.
  • Do not infer formal saturation from one model family, one mutable provider alias, one public commissioning epoch, or a cohort containing invalid cells. Such observations may reprioritize medium/large and private work, but only BQ-5’s predeclared eligible multi-family evidence changes saturation status.
  • Retire benchmark ranking weight prospectively. Preserve public/private anchors, historical difficulty labels, weights, attempts, and scores; equate new epochs through predeclared overlaps and uncertainty instead of rewriting the tide line.
  • Do not count a held-out commitment as a usable private lineage until custody can verify the payload, oracle, alternative-solution policy, provenance/consent, leakage status, and blind-execution contract without exposing it to the solver or training lane.
  • Never publish a cross-track rank. Cold Acquisition, Open Agent, Genesis-Adapted, and Embedded Local results require separate eligibility, cohort, analysis, and leaderboard surfaces; scaffold or hardware drift creates a new cohort.
  • Keep GenesisChallenge outside GenesisBench ranking and release authority. A contribution contest is never a fifth track, difficulty stratum, held-out epoch, benchmark adoption substitute, or corpus shortcut; accepted work enters only through destination-owned review, ordinary gates, and the R8.5 promotion/firewall rules.
  • Preserve GenesisChallenge attribution and failure history append-only. Solvers may earn signed credit for independently verified deltas, but no proposer, solver, model/provider, sponsor, contest operator, verifier/scorer author, benchmark custodian, or destination approver may admit, verify, reward, merge, or promote its own work unilaterally.
  • Never merge Foundry maturity with GenesisCode capability maturity, GenesisBench rank, or Genesis Model quality. FD and proof status are exact, orthogonal, artifact-bound, and can only advance through the named independent authority.
  • Foundry search history is append-only in meaning: pruning may compact payloads only through committed membership/disposition proofs; rejected candidates, failed laws, counterexamples, null ablations, and superseded promotion decisions remain discoverable and cannot be deleted to improve reported yield.
  • A Foundry pack, search engine, mechanism, abstraction, or law must beat its simplest adequate baseline under FB-7 before becoming a default. Added research breadth consumes an explicit resource/complexity budget and proposes at least one consolidation, retirement, or non-goal.
  • Production remains buildable, testable, releasable, and supportable with foundry/ absent. Any production dependency on Foundry code or mutable registry state is a P0 architecture violation; promoted artifacts must be copied through versioned destination-owned schemas and independently reverified.
  • Treat genesis/algorithms as an ordinary destination-owned package, never as a mounted Foundry registry. Every export binds exact semantics, applicability, complexity, provenance/license, supported profiles, independent evidence, fallback, compatibility, owner, and rollback; discovery, FD status, benchmark score, or challenge credit cannot admit it.
  • Prevent same-release circularity. Foundry may consume only a prior immutable Genesis Algorithms release as a baseline and may propose only a successor; tasks/checkers are precommitted before search, and the successor rebuilds and verifies with foundry/ absent before publication.
  • GenesisExamples/GenesisApps bind exact disclosed GenesisCode, target-profile, and Genesis Algorithms releases and exclude mutable Foundry state, hidden candidate history, and search-specific oracles. Every released item is independently runnable, classified, maintained, and evidence-bound; a dependency that does no workload-relevant work is rejected.
  • GenesisBench binds exact disclosed GenesisCode, Genesis Algorithms, and public example/app catalog releases and excludes mutable Foundry state, hidden candidate history, published solutions as held-out answers, and search-specific oracles. Later library/catalog releases create explicit compatibility cohorts rather than rewriting tasks or historical results.
  • Version GenesisCode, Foundry snapshots, Genesis Algorithms, GenesisExamples/GenesisApps catalogs, GenesisBench, reference scaffolds/adapters, task epochs, and Genesis Model packages independently. Every compatibility claim names all relevant identities; no release rewrites another outcome’s historical evidence.
  • Keep portfolio milestones observational and dependencies scoped. The machine policy must not make GenesisCode Core wait for Foundry/Algorithms, curated examples/apps, Bench, Challenge, or Model; must require finite F2.r before R8.3.a-b and finite R8.3.b before the first Bench/Challenge release; must not make later Bench wait for unrelated Foundry, universal-app, or model expansion; and must never let one outcome borrow another’s evidence.
  • Never use a non-sequential workstream name as an implicit scalar prerequisite. Fan in every required leaf task explicitly, keep parallel branches explicit, and require the roadmap-manifest self-test to reject terminal-task-only expansion.
  • Treat product compatibility as a typed relation, not shared branding: required, optional, compatible, incompatible, preview, and absent are explicit, and an optional Genesis Model can never become an undeclared GenesisCode build, verification, or operation dependency.
  • Admit model output to examples, packages, or training data only through R8.5 provenance, license, verification, review, deduplication, secret, and held-out-firewall rules.
  • Establish and adversarially verify BQ-7/BQ-8 capability, storage, operator, and audit separation before admitting the first training artifact or starting training; a later firewall audit cannot retroactively cleanse an exposed corpus or checkpoint.
  • Import Theseus or ASI Stack learning only through a clean signed content-addressed R8.5.i source edition and transfer manifest. Mutable latest state, dirty worktrees, prose summaries, dashboards, and uncaptured conversations may propose a candidate but cannot authorize model design, data, training, or release decisions.
  • Evaluate a coupled language/model pair change through R8.5.r’s four cells before promotion; evaluate a model-only change against the frozen language through its two non-degenerate cells. Preserve cold-model accessibility and reject benchmark-specific APIs, evaluator changes, or combined-only co-adaptation even when the candidate specialist improves.
  • Add a new capability family only with a spec owner, capability schema, resource model, threat model, host/platform scope, negative controls, generated bindings/cards, and at least one product workload.
  • A target name or file suffix is never capability evidence. Promotion requires install/start/flash and a nontrivial workload in the declared real runtime, plus resource/lifecycle evidence; descriptor, empty-export, synthetic launcher, and simulator-only results remain labeled below L4.
  • Preserve the one-language authoring contract. Generated foreign artifacts are disposable read-only outputs with source maps and identities; if a user or agent must edit one, record a product defect and remove that target from the qualified matrix until GenesisCode owns the missing abstraction.
  • Keep Embedded Linux, full-host, and constrained MCU profiles distinct. New boards arrive through BSP/HAL packages and conformance suites, never hidden compiler branches; every firmware claim names exact silicon, board, toolchain, memory map, peripherals, and physical evidence.
  • Change canonical hashes, logs, package/patch/evidence formats, bytecode, snapshots, or bootstrap envelopes only through a versioned compatibility proposal and migration corpus.
  • Review this roadmap at every milestone and at least monthly while active. Reordering must explain dependency or evidence changes, not preference alone.
  • The final test of the roadmap is not how advanced it sounds. It is whether an independent user and their agents can reproduce the claims, understand failures, remain inside explicit authority, and ship useful software without trusting hidden machinery.