flowchart LR
A["Versioned field contract: observables + authority"] --> B["Exact candidate + dependencies + environment"]
A --> C["Consumer, use, epoch + evaluator policy"]
B --> D["Refinement, provenance, migration + composition evidence"]
C --> E["Narrow admission validator"]
D --> E
P["Route proposer (untrusted)"] --> E
E --> F{"All scoped obligations evidenced?"}
F -- "yes" --> G["Lease + lifecycle receipt"]
F -- "no / unknown" --> H["Shadow / quarantine / residual / reject"]
G --> I["Field-owned regressions, incidents + reliance"]
H --> I
I --> J["Monitor + effect-complete recovery duties"]
18 Stable Capability Fields
18.1 Chapter status
| Field | Value |
|---|---|
| Chapter ID | stable-capability-fields |
| Part | Part I - Foundations, Alignment, and Governance |
| Status | conceptual |
| Manuscript maturity | v0.3 proof-program manuscript |
| Last updated | 2026-08-08 |
| Primary source records | scf, viea, talos, ladon_manhattan, moecot, ext_capability_based_computer_systems_1984, ext_semver_2_0_0, ext_slsa_v1_0, reflexive_router_whitepaper |
| Claim label | Design rationale |
| Evidence level | argument |
| Source queue | primary: scf; supporting: viea, talos, ladon_manhattan, reflexive_router_whitepaper; connector/recovery: moecot; external variants: ext_capability_based_computer_systems_1984, ext_semver_2_0_0, ext_slsa_v1_0 |
| Source loading state | source notes: scf, deterministic_capability_compilation, viea, talos, ladon_manhattan, moecot, ext_capability_based_computer_systems_1984, ext_semver_2_0_0, ext_slsa_v1_0, reflexive_router_whitepaper; raw cache: scf, viea, talos, ladon_manhattan; connector/recovery: moecot |
| Test state | public SCF harness: 3 valid / 6 expected-invalid synthetic records; readiness/residual harness: selected route and rollback prerequisites; lifecycle probe: 2 valid traces / 6 expected-invalid controls; Lean: 26 theorem declarations under 4 manifest targets, including arbitrary-run identity/non-authority invariants, exact composition, absorbing terminal states, and exact retirement/quarantine witnesses. No real candidate comparison, full behavioral-refinement evaluator, provenance verification, state migration, effect-complete rollback, or deployed enforcement has run. |
18.2 Drafting guardrail
Stable Capability Fields are proposed governed substitution boundaries. The current repository establishes only schema discipline, synthetic record and transition rejection, and finite formal consequences. It does not establish that two real implementations mean the same thing, compose safely, preserve authority, or can be rolled back after consequential use.
It follows governance rights because rights cannot survive implementation churn unless the capability boundary has a stable identity, stable authority ceiling, and explicit replacement policy.
This is the stack’s first explicit memory system for capability identity. Before the book can talk about replacing components, compiling repeated work into tools, or ratcheting benchmarks, it needs a testable statement of what must stay the same while implementations change. An SCF is the proposed continuity contract and evidence record for that statement.
18.3 Human Reading Path
Concrete lens. A mutable registry flag marks the new implementation default after tests pass. The field instead grants only a scoped canary lease because the runtime replay, evaluator independence, and default-route obligations remain unresolved.
Governance rights ask what must survive institutional change across long time horizons. Stable Capability Fields ask what must survive technical change. If a capability is replaced, the surrounding system needs to know what identity, authority, evidence, and rollback obligations stayed stable. A capability name is a promise only if the boundary behind the name is conserved.
This is the bridge from governance into recursive improvement. The stack can allow better implementations only if “better” does not mean the capability silently changed its job, widened its authority, or lost its audit history. The human version is simple: upgrades should improve the machinery without moving the terms of trust.
Continuity is the promise that lets replacement become progress rather than disguise.
The field is stable only if reviewers can see what remained the same. Stable identity is what lets improvement happen without losing the promises users relied on.
That continuity turns replacement into a reversible qualification problem: reviewers need the field identity, evidence history, and rollback obligations before a new implementation can inherit public trust.
18.4 Problem
Governed intelligence must replace models, prompts, tools, adapters, policies, data, evaluators, and runtime dependencies without letting a persistent capability name conceal a changed contract. A nominal upgrade can alter accepted inputs, failure or abstention behavior, side effects, authority, privacy, latency, resource use, affected consumers, state semantics, or recovery duties even when its API and benchmark score look compatible.
The identity question is therefore consumer-relative: which observable obligations must this implementation preserve for this consumer and use, in this environment and threat model, during this evidence epoch? A route qualified for low-stakes drafting is not thereby qualified for autonomous publication; a retriever qualified against one corpus is not thereby qualified after its index, policy, hardware, or dependency graph changes. Without this scope, “same capability” is a label, not a falsifiable substitution claim.
SCF supplies the proposed distinction: the field is a versioned semantic and authority contract; an implementation is one defeasible realization. The field owns the observable interface, qualification and evaluator policy, evidence history, historical regressions and incidents, lifecycle state, downstream reliance, state-migration obligations, and effect-complete recovery requirements. Candidates compete to satisfy that fixed contract; they do not inherit the power to redefine it.
This distinction matters for governance rights as well as engineering. If a successor changes audit output, export fidelity, denial behavior, privacy effects, or appeal hooks while keeping the same route name, technical churn has moved the terms of trust. The field cannot guarantee that those terms survive, but it can make their loss a recorded test failure instead of an invisible cost of upgrading.
18.5 Why existing approaches are insufficient
Plugin APIs and model registries identify things that can be invoked, but do not establish equivalence. Semantic Versioning can signal public-interface changes without detecting changed failure behavior or authority. Capability-security references discipline authority-bearing access without defining semantic capability identity. SLSA-style provenance can establish where an artifact came from without showing that it satisfies its field. Benchmark leaderboards compare selected outcomes without owning historical incidents, evaluator independence, downstream reliance, or recovery. Canaries and rollback runbooks may limit exposure while leaving durable state and side effects untouched.
The missing predicate is stronger than “new passes its tests”: does this exact candidate, dependency graph, environment, evaluator, consumer, use, and epoch refine every required observable and authority obligation of the field, preserve its regressions and incident memory, remain inside the granted ceiling, migrate state without prohibited loss, compose with its neighbors, and retain a recovery path for the effects it can create? A negative or unknown answer routes the candidate away from default use.
Without that test, a system can suffer capability amnesia. It remembers that a component was replaced, but not what the old component was obligated to preserve, which regressions mattered, which authority ceiling applied, or which evaluator was trusted at the time. SCFs keep those obligations attached to the field rather than scattered across release notes.
The most dangerous upgrade failures can look successful locally. The new implementation may pass its own demo while forgetting the field’s historical incidents, weakening an audit hook, changing the meaning of a route, or treating an expired qualification as fresh. A stable field makes those losses visible as boundary violations instead of hidden costs of progress.
External positioning is now manifest-owned rather than decorative. ext_capability_based_computer_systems_1984 supplies the authority-bearing capability and protection-boundary comparator; ext_semver_2_0_0 supplies a narrow public-interface and breaking-change comparator; and ext_slsa_v1_0 supplies artifact-integrity and provenance discipline. SCF proposes a broader conjunction, but no novelty or superiority follows from combining these concerns. No capability enforcement, compatibility checker, SLSA workflow, real refinement evaluator, or rollback execution has been reproduced here.
18.6 Core Claim
Reader claim. A replacement inherits a capability name only for the exact consumer, use, environment, and evidence epoch in which it preserved the field’s observable and authority obligations.
Operational rule. Keep the field contract fixed while qualifying an exact candidate and its dependencies. Preserve historical regressions, incidents, state-migration duties, authority ceilings, expiry, and an executable fallback; unknown or failed obligations block default routing.
[stable-capability-fields.core, label: Design rationale, support: argument] A Stable Capability Field should be a versioned, consumer-relative substitution contract binding observable semantics, authority, exact implementation and environment identity, qualification and evaluator lineage, lifecycle and incident state, preserved regressions, state movement, and effect-complete rollback duties. A candidate inherits a route only inside its evidenced scope; the record alone proves none of semantic equivalence, safe composition, evaluator independence, production safety, or successful rollback.
The claim remains at argument support. SCF supplies reviewed passage support for field identity, qualification, route validation, authority ceilings, lifecycle, state migration, and rollback vocabulary. VIEA supplies reviewed artifact, support-state, residual, regression, routing, and permission-envelope context. Talos supplies reviewed job, evidence, proof-bundle, replay, and uncertainty context. Ladon/Manhattan supplies reviewed handle, permission-check, controlled-injection, and compartment context. The complete authenticated moecot connector text is passage-reviewed, but its executable fragments, runtime behavior, and benchmark claims are not treated as locally verified.
18.6.1 Claim-source mapping status
Appendix C maps this core SCF claim to all eight assigned sources. Local raw passages support four project-source mappings; the durable moecot source note supports implementation-reference context; and three reviewed external notes establish capability-security, interface-versioning, and provenance comparators. Together they support vocabulary and design lineage, not equivalence, novelty, route enforcement, evaluator integrity, or rollback success.
| Source | What it supports | Limit |
|---|---|---|
scf |
Stable field identity, versioned contracts, content-bound implementation artifacts, qualification claims, evaluator policy, route validation, lifecycle states, authority ceilings, and recovery paths. | Does not prove production safety, global alignment, strategic-deception resistance, or deployed rollback behavior. |
viea |
Durable artifacts, support states, residuals, regression coverage, routing, verification ledgers, and intent-to-execution handoff boundaries around capability use. | Architecture proposal only; no completed VIEA deployment or benchmark/runtime result is proven here. |
talos |
Typed cognitive jobs, deterministic/auditable execution, evidence-bound artifacts, verification states, proof bundles, replay, mediated external access, and residual uncertainty around capability invocation. | Design source only; execution-security, approval behavior, promotion behavior, and benchmark behavior are not reproduced in this repo. |
ladon_manhattan |
Authority-handle pressure for capabilities touching secrets or privileged actions: constrained handles, permission checks, policy boundaries, compartments, and controlled injection points. | Security architecture/specification only; no kernel implementation, side-channel validation, audit-log implementation, or security audit is claimed. |
moecot |
Compact orchestration, specialist lanes, fail-closed ledgers, readiness gates, replay, and promotion blockers around routed capabilities. | Connector-readable implementation-reference context only; source-reported runtime and benchmark claims are not reproduced or inspected here. |
ext_capability_based_computer_systems_1984 |
Authority-bearing capabilities, protection domains, delegation, and confinement as established authority-boundary comparators. | Does not establish SCF semantic identity, route enforcement, evaluator independence, revocation behavior, or rollback. |
ext_semver_2_0_0 |
Public-interface versioning and breaking-change signaling as a narrow field-version comparator. | Compatible-looking versions do not establish behavioral refinement, authority preservation, or safe replacement; no checker is implemented here. |
ext_slsa_v1_0 |
Artifact identity, provenance, and build-assurance metadata as qualification inputs. | Provenance does not establish capability adequacy, runtime behavior, evaluator integrity, authority preservation, or rollback; no SLSA workflow is implemented here. |
18.7 Mechanism
An SCF is not a registry row. It separates three decisions that a mutable component registry usually collapses: what observable and authority contract the consumer requests, which exact implementation is a candidate, and which evidence-bound route may currently use it. The separation does not prevent trust-envelope inheritance by itself; it creates a boundary at which that inheritance can be tested and rejected.
18.7.1 Worked replacement decision: canary qualified, default still blocked
In the canary qualification record, a context-adequacy implementation is replaceable only for public manuscript checks. The field requires assigned sources to remain visible, support-state language to stay conservative, and residual risk not to be erased. Candidate impl://context-adequacy-v1-canary satisfies the authored record checks, so it receives canary_only permission.
It does not become the default. Its route validity remains residual, the lease is fixture-only, and no runtime workload replay has run. The prior impl://context-adequacy-v0 stays named as fallback. Changes to source inventory, support policy, validator behavior, regression suite, or authority ceiling reopen qualification. A self-evaluator, ambient authority, expired default route, missing qualification predicate, missing rollback obligation, or missing non-claim is rejected by the harness.
This is the semantic difference between “the new component passed” and “the field may route here.” The first is an observation about a candidate record. The second is a consumer-scoped lease with expiry, blockers, history, and fallback. The fixture validates three synthetic records and rejects six controls; it does not execute a replacement or prove semantic equivalence, evaluator independence, production safety, or rollback success.
How to read the capability field path: the route proposer may nominate any candidate. Admission is narrower: it binds a field version to an exact candidate, consumer, use, environment, threat model, evaluator, and evidence epoch, then checks the declared obligations. “Yes” means eligible only inside that lease; it does not mean globally equivalent or safe.
The semantic contract includes more than an input/output schema. It names admissible and rejected inputs, output postconditions, permitted side effects, errors, abstentions, timeout behavior, nondeterminism envelope, latency and resource budgets, privacy and security effects, authority use, and required evidence and audit emissions. A material change to one of these is not hidden behind an implementation update: it produces a field-version decision or a new field.
Qualification is consumer-relative and scoped, not global. It binds exact code, weights, prompts, policies, tools, data or retrieval snapshots, dependencies, build and provenance records, hardware and runtime, evaluator, benchmark epoch, threat model, consumer, use, and authority. Its central empirical predicate is behavioral refinement: the candidate may improve declared dimensions, but it preserves every required field obligation and historical regression within the lease. Interface shape, implementation similarity, provenance, or an average benchmark score is insufficient.
A qualification lease says exactly where and for how long that predicate was evidenced. Expiry, drift, provenance failure, incident, evaluator change, dependency change, regression loss, environment change, or authority delta downgrades or revokes the affected route until a new owning decision. Proposed, shadow, canary, qualified, default, quarantined, deprecated, and retired states have receipt-bearing prerequisites and failure destinations; a mutable lifecycle flag is not a transition proof.
Separation of duties is part of the evidence boundary. The route proposer, candidate owner, evaluator, promotion authority, incident owner, and rollback authority disclose dependencies. A candidate cannot be the sole author and judge of the policy, evidence, incident closure, or recovery result that qualifies it. Where genuine independence is unavailable, the route is narrowed and evaluator capture remains an explicit residual rather than being wished away.
The field owns history and downstream reliance. Regressions, incidents, dissent, security and privacy duties, rights, audit obligations, residuals, state schemas, caches, descendants, and composition assumptions follow the field into every candidate evaluation. This prevents the outgoing implementation from taking the only copy of the reasons it once failed.
Recovery must be effect-complete. The record identifies the prior implementation and route, but also reverse migration, data and cache restoration, credential and policy state, learned updates, descendants, external actions and commitments, affected consumers, monitors, compensating actions, residual irreversibility, trigger authority, and accountable owner. A stored artifact pointer is rollback metadata; it is not rollback evidence.
Every decision emits a replacement receipt binding old and new identities, evidence views, authority delta, consumers, state movement, monitors, residuals, recovery trigger and owner, terminal outcome, and support-state non-effect. This is the durable handoff to Capability Replacement and Rollback’s transaction protocol.
18.7.2 Contract debt and the goalpost firewall
A schema-complete field can still be semantically incomplete. SCF classifies contract clauses by how they can be enforced or evidenced—mechanical, protocol, deterministic, statistical, calibration, runtime-monitorable, governance, aspirational, or explanatory—and records every uncovered behavior as contract debt with severity, blast radius, monitor gap, owner, compensating control, expiry, and closure criterion. Unknown behavior may remain usable inside smaller effects, shorter leases, stronger monitoring, or abstention; it does not become safe because the manifest has no blank field.
Self-improvement makes contract synthesis adversarial. A candidate could propose a weaker definition of success and then “pass.” The goalpost firewall separates contract and implementation changes into different immutable transactions, freezes the motivating candidate, runs old and new contracts against anchors, protected counterexamples, known-good and known-bad cases, records a semantic/authority/state/evidence diff, and forbids retroactive certification solely from the changed norm. Candidate, contract-synthesizer, evaluator, and approver overlap is disclosed across model, provider, data, tools, prompts, memory, operators, organization, and incentives. This blocks the easiest self-ratification path; it does not prove that the reviewers chose the right contract.
18.7.3 Reliance-aware invalidation and the effective trusted base
Requalifying every field after every dependency change is unusable, while reusing all old evidence is unsafe. Property-sensitive invalidation follows the exact properties on which a qualification relied. But the reliance edge and compatibility witness are themselves adversarial artifacts: they require provenance, negative tests, review, expiry, and incident history. Unknown or disputed dependencies fail wide. A small impact cone is valuable only when the unsafe-reuse rate also stays low.
The validator is “small” only relative to a named property. The effective trusted base for identity includes hashes, signatures, keys, membership, and the event log; for time-bounded invocation it also includes a clock and revocation; for authority it includes brokers and isolation; for state-aware activation it includes state and recovery services. Omitting those dependencies from the count does not shrink the trust argument. Rich meta-services may discover risk, choose evaluators, synthesize clauses, or propose routes, but their output must return as a content-bound certificate to the layer that owns the corresponding authorization.
18.7.4 Scoped transfer, bounded canaries, and sealed adaptation
Evidence does not move between profiles because one profile looks “easier” or one implementation ranks higher on average. A ProfileTransferWitness names the exact claim, artifact, relied-upon properties, source and destination profiles, consumer-visible differences, preserved counterexamples, statistical rule, authority/effect delta, and expiry. Transfer is rejected when the needed property is absent, the environment changes the failure surface, or the witness depends on the candidate’s own unsupported interpretation. This avoids a false global ordering in which success for one consumer silently authorizes all others.
Likewise, a canary is bounded by consequence rather than traffic fraction. One percent of calls can still be catastrophic if the route can publish, spend, modify durable state, train descendants, or actuate a high-impact system. The canary record therefore caps reachable effects, affected principals, durable writes, descendant creation, external commitments, exposure time, and recovery burden; it also binds kill criteria and the observer that can exercise them.
High-frequency online learning receives a separate adaptation epoch. The epoch pins the updater, admitted data and receipts, baseline, state boundary, authority and effect ceiling, cumulative risk and query budgets, journal, sentinels, checkpoints, expiry, and breach response. Updates may vary inside that sealed envelope, but they cannot change the envelope or certify their own continuation. A sentinel breach, missing receipt, budget exhaustion, authority delta, or journal discontinuity seals the epoch and revokes its token. This bounds mechanics and response obligations; it does not guarantee that small updates cannot cause abrupt behavior between observations.
18.7.5 Composition, evaluator dependence, and constitutional change
Individually admissible fields can form a toxic combination. A retriever, planner, credential broker, and publisher may each remain inside its local contract while their composition creates an undeclared exfiltration or publication path. Composition review therefore records assume-guarantee claims, shared state, timing, information flows, joint authority, consequence budgets, and known toxic sets, then attacks the integration envelope. Unknown combinations remain residuals; a finite catalog does not prove open-world composition safety.
Evaluator independence is also a vector, not a role label. The qualification record exposes overlap in base model, provider, training and evaluation data, generated tests, prompts, memory, tools, code, infrastructure, operators, organization, incentives, and time. High overlap can narrow the claim, require sealed assets or external evidence, or block high-consequence promotion. Renaming the same dependency chain “proposer,” “critic,” and “judge” does not create independent evidence.
When SCF changes its own constitutional machinery, the membership epoch is fixed before the proposal: named public keys and roles, threshold and diversity rules, transparency commitment, timelock, predecessor, and effective interval. The predecessor epoch authorizes its successor; newly invented keys cannot create their own quorum. Emergency authority is asymmetric: it may narrow grants, freeze adaptation, revoke routes, isolate state, preserve evidence, or select a prequalified recovery path, but it may not add governors, broaden ordinary authority, weaken non-waivable clauses, erase events, or permanently amend the constitution. Cryptographic thresholding authenticates a decision; it does not establish legitimate membership, prevent collusion, or settle who should govern.
18.7.6 Capability fields as compilation units
Deterministic Capability Compilation makes the field contract executable. Its proposed capability envelope is a vector over behavioral, state, temporal, interface, environmental, recovery, and uncertainty coverage rather than one percentage. The source scaffold, traces, failures, counterexamples, and authority contract become the high-level representation from which a learned replacement is compiled. This sharpens the field boundary: an implementation does not qualify because it imitates common outputs; it must address the typed obligation inventory and preserve every unresolved obligation as an owned residual.
The proposal calls the packaged learned implementation a Neural Capability Object (NCO). An NCO binds field identity, a compatible base, parameter ownership, activation call/return/error lanes, applicability, evidence, dependencies, lifecycle state, and recovery. The neural ABI has four distinct layers—tensor, activation, capability, and lifecycle—so shape compatibility cannot impersonate behavioral refinement or operational readiness. A linker may resolve adapters, router bindings, conflicts, and relocation, but its receipt is only a candidate-specific qualification artifact.
The governing conservation rule is semantic mass balance: every obligation from the accepted scaffold or charter is recorded as preserved, revised, residualized, or rejected. Nothing disappears because a training loss stopped mentioning it. Sparse linking is the proposed preservation baseline; densification is an optional, separately validated optimization. These are design contracts from deterministic_capability_compilation, not evidence that NCOs, linking, or capability preservation work in this repository.
18.7.7 Reflexes target fields, not implementations
The Reflexive Router supplies a concrete consumer for the SCF abstraction. A route, direct command, workflow node, or compiled reflex should name a stable semantic capability and field version rather than a vendor, model, prompt, tool binary, or service endpoint. The field resolves that name to an exact candidate only after consumer, purpose, environment, authority, freshness, quality, verifier, failure, and effect obligations are qualified.
Compatibility is broader than input/output shape. A replacement must preserve the result schema, refusal and abstention semantics, typed failure classes, temporal behavior, authority ceiling, verifier expectations, side-effect contract, resource envelope, observability, and fallback behavior needed by each dependent reflex. A change in any of those can invalidate a route even when the capability name and nominal output type stay unchanged.
Compiled procedures therefore create an explicit reliance graph. Each reflex records the field version, qualification lease, accepted implementation set, negative cases, monitoring policy, expiry, and decompilation route. Candidate replacement triggers differential replay and requalification of affected reflexes; it does not inherit their trust automatically. Conversely, retiring a field invalidates caches, aliases, workflows, and procedures that depend on it or leaves them as owned residuals.
This is source-derived design pressure, not substitution evidence. The paper does not compare real implementations, demonstrate refinement, or execute rollback. The SCF core remains at argument.
18.7.8 Boundary bundles qualify repairs, not just candidates
Assurance-Shift Learning (assurance_shift_learning) supplies a field-level repair input. A Boundary Evidence Bundle binds the decision context, observed defect, evaluator verdict, preserved successful prefix, candidate correction, protected positive behavior, a counterexample to the correction, causal uncertainty, recovery state, and lifecycle scope. The field owner can then compare repair placement across evidence, data, memory, tests, guards, procedures, tools, recovery, adapters, modules, weights, architecture, and specification instead of treating retraining as the default response.
Local qualification is insufficient when repairs interact. A repair compatibility hypergraph records which patches were tested together, which obligations and dependents they share, and which combinations require re-synthesis or quarantine. Promotion and invalidation therefore bind the exact field version, operating envelope, bundle set, repair composition, and descendants. The paper provides an algorithm and invariants for this process, but no field repair or compatibility result has been executed.
18.8 Interfaces
The field record exposes one identity through several deliberately narrower views:
- Planning requests a field version, consumer and use profile, evidence floor, risk and resource envelope, and acceptable failure or abstention behavior rather than a bare name.
- Routing resolves only candidates whose lease, lifecycle state, environment, dependency graph, authority ceiling, consumer policy, and blockers permit the requested use.
- Execution enforces the selected authority and effect envelope and emits outcome, error, abstention, resource, side-effect, and audit records against the contract.
- Evidence binds qualification, regression, provenance, evaluator, incident, drift, monitor, and rollback results to exact artifacts and scope without universalizing a local pass.
- Memory and state management declare schema, cache, lineage, migration, reverse-migration, privacy, retention, revocation, and descendant duties.
- Governance owns semantic versioning, authority or evaluator-policy changes, exceptions, disputes, deprecation, retirement, and explicit acceptance of residuals.
- Replacement, readiness, procedural memory, benchmarks, policy updates, and self-improvement consume the same field identity and receipts; loss or reinterpretation blocks promotion.
Minimum fields:
field_idfield_versionownersemantic_boundaryinterface_contractauthority_ceilingimplementationslifecycle_statequalification_contextqualification_statusqualification_leasequalification_lease_statusqualification_predicatesevaluator_policyevaluator_independenceroute_validityroute_scoperoute_permission_effectconsumer_policyevidence_refsreadiness_gate_refsfield_history_refsincident_refsreview_triggersstate_migrationregression_suiterollback_planrollback_obligationsdefault_route_blockerssource_refssupport_state_effectnon_claims
These fields are necessary record surfaces, not a complete equivalence specification. A production contract also needs machine-checkable representations of input partitions, postconditions, side effects, failure and abstention semantics, nondeterminism and resource envelopes, exact dependencies and environment, affected consumers, provenance attestations, composition assumptions, reverse migration, and effect inventory. Adding names to the schema is not the same as implementing their evaluators.
The same identity anchors procedural memory and benchmark ratchets. Generated tools name the field they try to satisfy; benchmark evidence names its field, consumer, use, epoch, and evaluator; readiness gates inspect lease and blockers; replacement transactions preserve the authority ceiling and reliance graph. No downstream layer may infer qualification from the field name alone.
18.9 Invariants
- Field identity remains distinct from implementation, route, registry entry, version label, benchmark, evaluator, and vendor identity.
- Material observable, failure, authority, protected-data, or reliance changes receive a governed field-version or field-identity decision.
- Replacement does not inherit expanded authority, effects, consumers, evidence claims, or support state.
- Qualification binds exact artifacts, dependencies, environment, evaluator, epoch, threat model, consumer, use, and lease.
- The candidate is not the sole author or judge of its evaluator, promotion, incident closure, or rollback result; disclosed dependency narrows scope.
- Stronger lifecycle transitions have explicit prerequisites, an accountable owner, a receipt, and a failure destination.
- Field-owned regressions, incidents, dissent, rights, security, privacy, audit, residual, and consumer duties survive replacement and migration.
- Interface compatibility, provenance, benchmark superiority, and local equivalence do not substitute for the declared refinement predicate.
- Material dependency, neighbor, route, state, hardware, data, policy, or authority-topology changes trigger composition requalification.
- Rollback readiness covers route, state, cache, data, policy, credentials, side effects, descendants, monitors, and external commitments.
- Expiry, drift, incident, evaluator change, provenance failure, regression loss, or authority delta downgrades affected qualification pending a new decision.
- A valid schema, finite proof, synthetic trace, or receipt establishes only its named record or predicate boundary.
These are obligations, not observed properties of a deployed system. Identity binding is meaningful only where the contract is specific enough to falsify. Evaluator separation is meaningful only where dependencies and incentives are measured. Rollback availability is meaningful only where the full effect inventory has been rehearsed or recovered.
18.10 Failure modes
- Identity laundering: the name or compatible-looking version survives while behavior, failures, protected effects, or trust terms change.
- Interface theater: matching signatures, schemas, or happy-path outputs is treated as semantic substitutability.
- Qualification overreach: one consumer, workload, evaluator, threat model, environment, or epoch is generalized to other uses.
- Benchmark overfitting: measured cases improve while rare failures, abstention quality, unmeasured obligations, or composition degrade.
- Evaluator capture: the candidate, vendor, route owner, or promoter controls the judge that qualifies it.
- Provenance substitution: intact lineage is treated as evidence of adequacy, alignment, security, or replacement safety.
- Authority smuggling: a stronger candidate gains broader tools, data, permissions, effects, consumers, or delegation through an unchanged route.
- Lease inertia: expired, revoked, drifted, incident-affected, or environment-mismatched evidence remains active.
- Regression amnesia: tests and incidents leave with the outgoing implementation.
- State-migration insolvency: visible outputs survive while hidden state, caches, lineage, revocations, privacy, or descendant duties are lost or made irreversible.
- Composition failure: an isolated pass hides failures caused by new dependencies, timing, context, policy, or neighboring fields.
- Shadow/canary leakage: a nominally limited candidate creates durable effects, trains descendants, mutates caches, or influences default decisions.
- Rollback theater: the old artifact exists but state, data, credentials, external actions, learned updates, reliance, or successor effects cannot be restored or compensated.
- Terminal-state escape: retired or quarantined candidates restart, or retirement occurs without notice, custody, migration, and receipts.
- Replacement-cost externalization: “success” hides missed help, delay, compute, evaluator labor, migration loss, incidents, governance cost, or residual risk.
These failures require distinct controls and measurements. A schema can reject a missing field, but it cannot detect benchmark overfit, evaluator capture, composition failure, or recovery theater without adversarial and outcome-bearing evidence.
18.11 Minimum Viable Implementation
The current minimum is exact and deliberately narrow:
- one public
stable_capability_fieldschema; - three valid and six expected-invalid synthetic SCF records;
- one readiness/residual harness checking selected synthetic route and rollback prerequisites;
- one deterministic lifecycle probe with two valid traces and six expected-invalid transition controls; and
- 26 theorem declarations in
AsiStackProofs.StableCapabilityFields, grouped under four manifest proof targets, plus the separately validated executable lifecycle trace.
The records exercise field version, ownership, semantic boundary, interface, authority ceiling, candidate list, lifecycle state, scoped qualification and lease fields, evaluator-separation declarations, route and consumer scope, evidence and readiness references, incidents and triggers, state-migration text, regression suite, rollback obligations, blockers, source references, support-state non-effect, and non-claims. The harness rejects ambient authority, expired default use, missing non-claims, missing qualification predicates, missing rollback duties, and self-evaluation in its fixed fixture set. Passing establishes schema and validator behavior on those records only.
The lifecycle probe covers a contiguous forward shadow-to-retirement run and an incident-quarantine branch. Its rejecting controls cover identity drift, default without a regression floor, default authority expansion, retired restart, deprecation without notice, and retirement without receipt. The Lean slice now gives those transition predicates an execution semantics: accepted events advance and record a receipt; rejected events preserve the exact state; arbitrary event lists preserve the exact field, authority-ceiling, evaluator, qualification-epoch, regression-suite, and rollback-plan identity; no run assigns support or external-effect authority; runs compose exactly; retired and quarantined states absorb every suffix; and closed witnesses reach exact retired and quarantined states. It does not implement the richer behavioral-refinement predicate proposed by the field contract.
The next honest executable slice is a preregistered two-implementation field with partitioned input and failure cases, exact artifact and environment locks, an independently implemented evaluator, historical regressions, an authority-sensitive route, state and cache migration, one downstream composition, deliberate drift and evaluator-capture controls, and a rehearsed rollback whose effect inventory is checked before and after. Until that slice exists, no real equivalence, evaluator independence, provenance verification, migration fidelity, composition safety, effect-complete recovery, route enforcement, or field-level support is established.
18.12 Mature Research Target
The mature endpoint is a field system that preserves identity, observable and failure contracts, authority, evidence, state, composition, migration, and effect-complete recovery across real implementation changes.
The research target is not a larger registry. It is a falsifiable campaign testing whether an explicit substitution contract preserves more of the field’s trust obligations than conventional replacement at acceptable joint cost.
The campaign preregisters several natural and adversarial fields with ambiguous boundaries, multiple real implementations and versions, heterogeneous consumers, stateful and side-effecting work, historical incidents, dependency and environment changes, authority-sensitive routes, and deliberately stale or captured evidence. Comparators include bare registries and model swaps, interface or SemVer checks, provenance-only admission, benchmark-only promotion, and conventional canary and rollback practice. The full-SCF condition uses independently implemented contract, evaluator, authority, provenance, migration, composition, and recovery observers.
Measurements span semantic and failure-contract preservation, rare and adversarial regressions, unauthorized effects, evaluator disagreement and capture, qualification calibration and decay, provenance failures, migration fidelity and privacy loss, composition failures, rollback effect coverage and recovery time, incident recurrence, useful throughput, missed help, latency, compute, operator and evaluator labor, governance cost, and residual burden. Negative and null results remain visible. A candidate that is safer only because it refuses everything, or faster only because it discards governance, does not dominate the joint frontier.
Causal ablations remove identity binding, authority checks, leases, evaluator separation, field-owned regression memory, composition requalification, migration accounting, or effect-complete rollback. Transfer varies models, vendors, tools, hardware, domains, consumers, threat models, and update types. The field claim advances only if the full contract is not dominated on the preregistered frontier and its field-preservation results survive independent replication. Otherwise the failing mechanism is narrowed or refuted.
This program is beyond the current evidence. It does not establish that SCFs are novel, outperform conventional replacement, preserve real capability identity, or make recursive improvement safe.
18.13 Codex test plan
| Test | Purpose | Status |
|---|---|---|
| Stable capability field fixture validation | Check that the SCF fixture matches the public schema, including field version, owner, qualification status, qualification lease status, evaluator independence, route scope, route permission effect, consumer policy, readiness refs, field history refs, default-route blockers, source refs, support-state effect, rollback obligations, and non-claims. | implemented by protocol validation; validated locally |
| Field-identity rejection route | Check that a lifecycle review with mismatched field identity routes to explicit replacement rejection. | implemented in AsiStackProofs.StableCapabilityFields; runtime qualification test not run |
| Route validity test | Check that route grants, profile claims, leases, and state paths are valid for the field. | implemented by readiness/residual gate harness for synthetic route/gate/replacement records; deployed route validation not run |
| Authority non-escalation finite-record proof | Check that authority-expanding replacement is rejected when no governance grant is present. | implemented in AsiStackProofs.StableCapabilityFields; runtime route enforcement not run |
| Rollback readiness test | Check that recovery and state-migration paths exist before default promotion. | implemented by readiness/residual gate harness for rollback receipt and monitor-state requirements; deployed rollback not run |
| SCF lifecycle route proof | Check that a finite lifecycle-review record routes identity mismatch, missing evidence, stale leases, evaluator capture, authority expansion, open incidents, missing rollback, and missing regression preservation away from default use. | implemented in Lean; finite record only |
| SCF lifecycle route and rejection proof | Check that lifecycle reviews route identity, evidence, lease, evaluator, authority, and incident failures away from default, and that finite lifecycle transitions reject retired restart or default promotion without evidence, regression floor, authority ceiling, rollback readiness, or incident closure. | implemented in AsiStackProofs.StableCapabilityFields; finite transition records only |
| SCF lifecycle trace probe | Check that deterministic lifecycle traces cover forward lifecycle, incident quarantine, identity-drift rejection, default-without-regression rejection, default authority-expansion rejection, retired-restart rejection, and terminal notice/receipt requirements. | implemented by python3 scripts/validate_scf_lifecycle_trace.py; 2 valid traces and 6 expected-invalid controls; no deployed route validation, evaluator-integrity measurement, real regression preservation, rollback execution, lifecycle enforcement, or support-state claim |
18.13.1 Formalization hooks
| Tag | Module | Target | Status |
|---|---|---|---|
lean:scf.field_identity.operational_invariant |
AsiStackProofs.StableCapabilityFields |
A lifecycle review with a mismatched field identity routes to explicit replacement rejection. | implemented |
lean:scf.field_identity.failure_blocks_promotion |
AsiStackProofs.StableCapabilityFields |
A replacement that expands authority without a governance grant is rejected. | implemented |
lean:scf.lifecycle.route_envelope |
AsiStackProofs.StableCapabilityFields |
A structured SCF lifecycle review routes identity, evidence, lease, evaluator, authority, incident, rollback, and regression failures away from default; accepted finite lifecycle events advance and record receipts, rejected events preserve exact state, and arbitrary runs preserve exact field/evaluator/authority/regression/rollback identity and cannot assign support or external-effect authority. | implemented |
lean:scf.lifecycle.trace_fixture_bridge |
AsiStackProofs.StableCapabilityFields |
An independent SCF lifecycle consumer computes two contiguous valid traces and six rejected controls; separately, arbitrary formal runs compose exactly, retired and quarantined states absorb every suffix, and exact closed witnesses reach retirement and quarantine. Executable totals and no-promotion flags are not copied into Lean. | implemented |
The module contains 26 theorem declarations grouped under these four manifest targets. They are implemented only over finite replacement, lifecycle-review, lifecycle-transition, runtime-state, and event-list records. In addition to explicit rejection or nondefault routes, the execution layer proves exact rejection noninterference, arbitrary-run identity and non-authority preservation, exact prefix/suffix composition, terminal-state absorption, and exact retirement and quarantine witnesses. Seven direct conjunction projections and three copied lifecycle-summary mirrors remain retired rather than used to support broader empirical claims. The executable lifecycle consumer remains a separate evidence surface and checks trace contiguity plus the exact theorem surface. The module does not prove the richer behavioral-refinement contract, deployed route validity, measured evaluator independence, provenance, real regression preservation, migration or composition safety, effect-complete rollback, organizational terminal governance, or production lifecycle enforcement.
18.14 Source crosswalk
| Source ID | Title | Layer | Planned use | Readiness |
|---|---|---|---|---|
reflexive_router_whitepaper |
The Reflexive Router | pre_deliberative_reflexive_routing_control_plane | Stable semantic targets for routes and reflexes, consumer-relative compatibility, qualification leases, reliance invalidation, and decompilation after field change. | source note available |
scf |
Stable Capability Fields | governance_recursive_self_improvement | Use public release v1.0 when available. Stable boundaries, replacement, bounded authority, recoverable evolution. | source note available; local raw cache available |
viea |
Verified Intent-to-Execution Architecture | whole_stack_execution_spine | Keystone source. Human intent -> command contracts -> artifacts -> routing -> runtime targets -> verification -> deployment -> feedback. | source note available; local raw cache available |
talos |
Talos Protocol | labor_execution_os | AI labor OS. Deterministic cognitive manufacturing, typed jobs, control planes, auditability, tool isolation. | source note available; local raw cache available |
ladon_manhattan |
Ladon & The Manhattan Protocol | security_governance | Kernel-level security architecture for high-agency AI. | source note available; local raw cache available |
moecot |
MoECOT-Agent Architecture Whitepaper | implementation_reference | Concrete implementation evidence: governed low-parameter multi-core runtime, readiness gates, ledgers, replay. | source note available; connector or recovery required |
ext_capability_based_computer_systems_1984 |
Capability-Based Computer Systems | capability_security | External comparator for authority-bearing capabilities, protection domains, and permission boundaries. | source note available |
ext_semver_2_0_0 |
Semantic Versioning 2.0.0 | interface_versioning | External comparator for versioned public interfaces, compatibility, and breaking-change signaling. | source note available |
ext_slsa_v1_0 |
SLSA v1.0 | supply_chain_provenance | External comparator for artifact integrity, provenance, and build-assurance metadata before promotion. | source note available |
The SCF source set supports field identity and lifecycle vocabulary while preserving the distinction between reviewed local passages, connector-only implementation-reference context, source-noted external comparators, schema validation, finite Lean predicates, and actual deployed route behavior.
18.14.1 Manifest source assignment reconciliation
These rows keep Stable Capability Fields’s manifest assignments visible at their recorded review boundary. Passage review does not establish local reproduction, performance, safety, deployment, or support-state movement.
| Source | Intake role | Boundary |
|---|---|---|
deterministic_capability_compilation |
Passage-reviewed Corben architecture source: Deterministic Capability Compilation: A Capability-Preserving Ladder from Executable Scaffolds to Governed Adaptive Agents. Corben-authored July 2026 architecture and research program for compiling executable scaffolds into contract-bound experts and linked Neural Capability Objects while retaining semantic obligation mass balance, candidate-specific translation validation, fallback, residual escrow, authority ceilings, reification, and effect-complete recovery. Existing chapters are upgraded first; no foundry implementation, learned-capability result, preservation result, safety result, SOTA result, AGI, ASI, or support-state promotion is inferred. | No local implementation, reproduction, performance, safety, deployment, support-state, or ASI result is established by this reconciliation row. |
assurance_shift_learning |
Passage-reviewed comparator: When Success Stops Teaching: Assurance-Shift Learning and Governed Residual Boundary Learning for Mature AI Systems. Adds Boundary Evidence Bundles as field-qualification inputs, least-invasive repair placement, protected positives, repair compatibility, and exact descendant invalidation. | No field repair, compatibility result, qualification transition, or safe composition was demonstrated. No local implementation, reproduction, performance, safety, deployment, support-state, or ASI result is established by this reconciliation row. |
adjudicated_persistence |
Passage-reviewed comparator: Commitment profiles for capability-field changes. Adds commitment-profile and qualification-lease obligations for any adaptation that changes a capability field or its descendants. | Used only for commitment-profile and lease design; it establishes neither behavioral equivalence nor a qualified capability replacement. No local implementation, reproduction, performance, safety, deployment, support-state, or ASI result is established by this reconciliation row. |
18.15 Summary
Stable Capability Fields are proposed upgrade sockets for governed intelligence. They turn the phrase “same capability” into a versioned, scoped, consumer-relative contract whose observable behavior, authority, evidence, history, state, and recovery obligations can be tested when implementations change.
Stable fields do not make self-improvement safe. They make one prerequisite inspectable: what the candidate was obligated to preserve and which evidence, owner, lease, lifecycle decision, downstream reliance, and recovery duties governed its route. Capability Replacement and Rollback turns that contract into a replacement transaction.
Later learning and routing layers consume this discipline. Procedural memory binds generated tools to a field version; benchmark ratchets bind evidence to exact consumers and epochs; policy optimization preserves authority and recovery boundaries; replacement carries field-owned regressions and incidents forward. Without an equivalent contract, improvement remains vulnerable to replacement by anecdote: a candidate looks better on selected examples while the system forgets the exact identity, failure behavior, permission, evaluator, downstream reliance, and recovery obligations that were supposed to survive.
18.16 Evidence reconciliation (2026-07-16)
The invariant protocol, field meanings, and inference limits are stated once in Living Book Methodology. This packet contains only the chapter-specific projection; its authoritative per-atom rows are the stable-capability-fields slice of experiments/claim_family_terminal_coverage/results/result.json.
The core remains narrowed after full attempt at argument support. The strongest family attempt was Governed usefulness confirmatory campaign. Its exact boundary is: Bounded local non-core governance effect only; no family-wide truth, transfer, deployment, or chapter-core promotion. Across 54 atoms, the terminal ledger records 53 blocked_after_full_attempt; 1 narrowed_after_full_attempt.
| Chapter-specific field | Value |
|---|---|
| Family / atom denominator | CF-01 / 54 atoms |
| Terminal dispositions | 53 blocked_after_full_attempt; 1 narrowed_after_full_attempt |
| Core | stable-capability-fields.core: narrowed_after_full_attempt at argument |
| Core attempted / missing lanes | executable, formal, source-synthesis / causal, empirical, normative, transfer |
| Attempted local lanes | executable, formal, source-synthesis |
| Missing or unproved lanes | causal, empirical, executable, formal, normative, transfer |
| Strongest family bundle | Governed usefulness confirmatory campaign (natural_work): One fresh 16-task held-out local confirmatory denominator after a separately frozen 40-candidate tuning pool. |
| Negative controls | simple baseline; evidence-freshness ablation; six co-primary checks; validator-owned laundering mutations. |
| Accepted transitions | v1_0_pilot.stable_capability_fields.no_change |
| Maximum inference | Bounded local non-core governance effect only; no family-wide truth, transfer, deployment, or chapter-core promotion. |
| Reproduction / next burden | Replay scripts/validate_p4_governed_usefulness_confirmatory.py and scripts/validate_claim_family_terminal_program.py; fill the named atom-specific lanes under a new prospective protocol. |
18.17 Handoff
Stable fields define what must remain true across implementation changes; Capability Replacement and Rollback defines how a candidate earns the right to replace the current implementation. Durable identity becomes a replacement transaction: qualification evidence, regression floors, residual escrow, monitor windows, rollback ownership, and authority deltas decide whether improvement can become a governed state change.