Skip to main content

72  White-Box Evidence, Interpretability, and Activation Governance

72.1 Chapter status

Field Value
Chapter ID white-box-evidence-interpretability-and-activation-governance
Part Part IV — Evidence, Implementation, and the Living Book
Status conceptual
Manuscript maturity v0.5 source-role and method-comparison integrated reader chapter
Last updated 2026-08-02
Primary source records Four Corben architecture sources plus ten external method, challenge, and measurement comparators; see the source crosswalk below.
Claim label Design rationale
Evidence level argument
Source loading state source notes: deterministic_capability_compilation, kernel_english_residual_compiler, qcsa_whitepaper, platonic_world_model, ext_transformer_circuits_2021, ext_monosemanticity_2023, ext_scaling_sparse_autoencoders_2024, ext_circuit_tracing_2025, ext_probe_control_tasks_2019, ext_interpretability_illusion_bert_2021, ext_saebench_2025, ext_sae_benchmark_reliability_2026, ext_elk_report_2021, ext_influence_functions_2017
Test state The packet schema, record-shape fixture, independent semantic validator, twelve packet/protocol mutations, a 36-theorem Lean surface, and an independently simulated six-event lifecycle with 51/51 rejected mutations are implemented under two public targets. The claim-bearing campaign is protocol-ready and resource-isolated, not executed; no model-internal outcome or support transition exists.

72.2 Drafting guardrail

This chapter proposes how model-internal observations may enter the ASI Stack’s evidence and governance machinery. It does not claim that a discovered feature is a concept, that a circuit is a complete explanation, that an intervention is safe, or that mechanistic interpretability currently makes advanced models transparent. Source-reported methods and results remain external comparators until the exact artifacts are reproduced or independently challenged.

72.3 Human Reading Path

Concrete lens. The probe baseline names the correlated activation and treats ablation as causal confirmation. The packet adds matched controls, side effects, stability, and unexplained residuals.

Behavioral evaluation asks what a system did and whether that behavior survives challenge. White-box work asks what internal state helped produce it. Activation access does not make an interpretation true. A feature can split, merge, drift, or receive a convenient label that fits only selected examples.

Internal evidence should increase scrutiny before trust. A stable causal result may support a narrow decision. A visualization, probe, or story cannot widen authority, certify safety, or replace observed behavior.

Finding a pattern, naming it, showing that it predicts an output, changing it, and releasing the changed model are separate steps. Each needs its own record, controls, uncertainty, side-effect checks, and decision owner. Skipping those separations lets an attractive story carry more authority than it earned.

Value is measured against a strong behavioral baseline. Internal evidence may localize a regression earlier, reveal a dormant mechanism, or permit a more selective intervention. If it adds no decision value after compute, maintenance, privacy, and analyst costs, the simpler route should win. The goal is disciplined uncertainty about model internals, not a ceremonial claim of transparency.

72.4 Problem

Black-box tests are necessary but incomplete. They can miss dormant mechanisms, rare trigger paths, distributed representations, or internal changes that have not yet altered the sampled behavior. They also say little about why a model changed after training, fine-tuning, editing, routing, compression, or substrate replacement. For recursive improvement, checkpoint qualification, and activation steering, that blindness can matter.

White-box methods create a different danger. Once a representation is visible, humans are tempted to name it, draw a graph around it, and treat the resulting story as an explanation. Predictive probes can exploit information without identifying the computation that uses it. Sparse features can reconstruct some activations without recovering a privileged semantic basis. Attribution graphs can depend on replacement models and thresholds. Ablations can damage many unrelated computations. Steering can suppress a measured behavior while moving the underlying risk somewhere the evaluator does not look.

The stack therefore needs a custody and inference layer between internal access and governance. Its job is not to make every model legible. Its job is to keep the exact observation, interpretation, intervention, residual, and authority boundary visible long enough for other evidence owners to challenge it.

72.5 Why existing approaches are insufficient

Interpretability work is often discussed as though it had one output called an “explanation.” In practice it spans materially different claims:

  1. an internal tensor or event was captured;
  2. a method extracted a feature, component, pathway, or graph;
  3. a human or automated process assigned it a semantic label;
  4. the representation predicts a behavior or state;
  5. changing it alters an outcome;
  6. it is necessary or sufficient under a stated intervention;
  7. the recovered structure covers a meaningful share of the relevant causal process; and
  8. an activation policy built from it remains useful and safe in deployment.

Each step can fail while the earlier step remains true. Collapsing the ladder creates white-box laundering: an activation capture becomes a feature; a feature becomes a concept; a concept becomes a cause; a cause becomes a complete mechanism; and a mechanism becomes authority to steer or release.

Behavioral evaluation, formal specifications, security controls, and runtime monitoring do not solve that problem individually. Behavioral tests may not localize a mechanism. A formal proof can certify record routing without proving the interpretation faithful. Security can authorize activation access without making the analysis correct. Runtime monitors can observe an intervention without proving that its side effects are bounded. The missing object is a typed white-box evidence packet with an explicit maximum inference.

The external comparison set is deliberately plural. ext_transformer_circuits_2021 supplies a circuit-analysis vocabulary; ext_monosemanticity_2023 and ext_scaling_sparse_autoencoders_2024 supply dictionary-learning and scaling comparators; and ext_circuit_tracing_2025 supplies replacement-model attribution graphs and perturbation checks. Their source notes preserve approximation, reconstruction, feature-quality, and faithfulness limits. None is represented as local reproduction, a complete account of model computation, or proof that an activation intervention is safe.

72.6 Core Claim

[white-box-evidence-interpretability-and-activation-governance.core, label: Design rationale, support: argument] White-Box Evidence, Interpretability, and Activation Governance owns a model-, checkpoint-, substrate-, tokenizer-, population-, capture-, method-, feature-or-circuit-, label-, intervention-, comparator-, evaluator-, authority-, and time-specific Internal Evidence Packet: internal observations may influence scrutiny or activation policy only when their lineage, assumptions, reconstruction error, stability, causal status, behavioral cross-check, side effects, unexplained residual, uncertainty, expiry, and maximum inference survive review; no activation, probe, feature, label, attribution graph, steering result, source-reported mechanism, or formal record check alone establishes semantic truth, causal completeness, model safety, release readiness, support, transfer, AGI, or ASI.

Reader claim. An activation correlated with a behavior is a lead for scrutiny, not proof that the feature means what its label says or causes the behavior.

Operational rule. Bind each internal observation to exact model, checkpoint, tokenizer, population, capture, method, label source, reconstruction error, stability, intervention, comparator, behavioral cross-check, side effects, uncertainty, and expiry. Use it to route scrutiny unless causal evidence and residual analysis justify more.

72.6.1 Worked activation packet: a “deception feature” that tracks topic

An interpretability method finds one activation that is higher on a deception benchmark and labels it “deception.” Ablating the feature reduces deceptive answers, but it also suppresses discussion of hypothetical agents across both deceptive and honest examples. A matched topic-control set reveals that the activation tracks the subject matter more broadly than the behavioral label. The packet therefore routes those cases to extra review but blocks the claim that a deception mechanism was isolated.

The result keeps the capture population, feature identity, label provenance, reconstruction error, intervention, comparator, output changes, side effects, and unexplained residuals together. A later circuit-level intervention may strengthen or refute the hypothesis, but it receives a new packet. The record discipline does not establish semantic truth, causal completeness, whole-model safety, or readiness from an activation, attribution graph, or steering result.

The claim remains at argument. It states the governance burden for white-box evidence; it does not assert that present methods can meet that burden broadly.

72.7 The internal-evidence packet

The packet preserves six objects that must never be collapsed:

Object What it records What it does not establish
Capture record Exact model, checkpoint, site, tensor/event, input population, software, precision, and access policy That the captured state is semantically meaningful
Extraction record Method, hyperparameters, basis, replacement model, reconstruction, sparsity, thresholds, and seeds That the extracted unit is natural, unique, or causal
Interpretation record Label author, examples, counterexamples, alternatives, adjudication, uncertainty, and expiry That the label is the model’s own concept
Causal record Intervention target, control, necessity/sufficiency question, behavioral effect, dose, and uncertainty Completeness, safety, or transfer outside the intervention envelope
Coverage and residual record Reconstructed variance, unexplained behavior, missing paths, feature splitting/merging, and method disagreement That a scalar coverage estimate captures all relevant computation
Policy and release record Proposed monitor/steer rule, scope, authority, side effects, fallback, rollback, and qualification decision Authority to deploy merely because an intervention worked in a study

Every packet ends with one of a small number of permitted dispositions: observed, predictive_only, causal_bounded, contradicted, unstable, incomplete, expired, policy_candidate, or rejected. None is a synonym for “the model is understood.”

flowchart LR
  A["Exact model and capture lease"] --> B["Feature / circuit extraction"]
  B --> C["Labels, alternatives, counterexamples"]
  C --> D["Held-out stability and behavioral prediction"]
  D --> E["Ablation, patching, replacement, steering"]
  E --> F["Side effects, coverage, unexplained residual"]
  F --> G{"Evidence disposition"}
  G -->|"weak, unstable, contradicted"| H["Archive / raise scrutiny / residual"]
  G -->|"bounded causal evidence"| I["Activation-policy candidate"]
  I --> J["Independent qualification and release authority"]
  J --> K["Monitor, expire, rollback or retire"]
  K --> L["Re-evaluate after material model change"]

What the flow shows: white-box access begins an evidence process; it does not end one. Even a bounded causal result must pass separate policy, qualification, release, monitoring, and rollback decisions.

72.8 Ownership boundaries

Adjacent owner It keeps This chapter exclusively owns
Evidence States and Claim Discipline Support states, transitions, review, and claim identity Admissibility and maximum inference of model-internal evidence
Benchmark Ratchets Behavioral instrument validity, baselines, contamination, and score lineage Feature/circuit extraction and causal-internal validation
Adversarial Evaluation Elicitation, sandbagging, monitor sabotage, and behavioral deception tests Internal mechanism hypotheses and activation interventions
Replaceable Cognitive Substrates Architecture-family capability cards and adoption/rollback Cross-checkpoint and cross-substrate interpretation identity
Model Weight Custody Weight access, storage, transfer, and protected artifact release Scoped access to internal states and derived interpretability artifacts
Security Kernel Information-flow policy, isolation, secrets, and hostile access Whether an authorized internal observation is epistemically admissible
Runtime Adapters Permission and receipts for tool or external effects Permission and receipts for activation observation or intervention
Readiness and Safety Cases Deployment decision and structured assurance White-box evidence packet supplied to those decisions

72.9 Why this boundary earns a chapter

Interpretability is not merely another evaluation technique. It introduces a new protected object—the activation or internal-state derivative—and a new inference ladder from capture through extraction, labeling, prediction, intervention, coverage, and policy. Evidence States owns whether a claim moves; Benchmark Ratchets owns whether a behavioral instrument is valid; Security and Weight Custody own access. None of them owns the question this chapter must answer: what is the maximum decision-relevant inference that may survive from one exact internal observation or intervention?

That ownership remains useful even when white-box methods lose to behavioral baselines. The packet can end unstable, incomplete, contradicted, or rejected, leaving a reusable explanation of why the attractive internal story did not earn governance weight. Removing the chapter would scatter those obligations among evidence, security, benchmarks, and release policy, making it easier for a feature label or steering result to cross boundaries unnoticed. The terminal reader decision is therefore integrate at argument support, with empirical promotion kept in a separate, prospectively frozen lane.

72.9.1 Publication placement and preserved technical ownership

In the consolidated architecture reference, this chapter is the technical detail route beneath Adversarial Evaluation, Sandbagging, and Training-Time Deception. The parent owns the integrity of behavioral and training-time observations, including elicitation, selection context, monitor and reward provenance, cross-context discrepancy, alternative hypotheses, outcome audit, and re-evaluation. This route continues to own the distinct white-box object: the maximum inference permitted from exact internal-state capture, method-relative extraction, semantic labeling, predictive challenge, causal intervention, coverage residuals, and activation-policy qualification.

That placement does not merge the claims. An activation, probe, sparse feature, circuit, attribution graph, or intervention does not inherit a deception, sandbagging, intent, monitor-integrity, or safety conclusion from the parent. The parent does not inherit this chapter’s construct validity, method assumptions, causal status, residual coverage, source mappings, proof targets, tests, or support. This stable route retains its local owner identity, legacy URL, evidence exit, failure modes, and non-claims. The composition is editorial only and creates no empirical result, support transition, authority, readiness, release, AGI, or ASI inference.

72.10 Mechanism

72.10.1 1. Register the model and observation boundary

The packet binds the exact weights, checkpoint lineage, tokenizer, architecture, precision, adapters, router state, runtime code, and data population. A model edit, merge, quantization change, tokenizer change, substrate swap, or material distribution shift expires affected interpretations by default. “Same model family” is not sufficient identity.

Activation access receives its own purpose, audience, retention, disclosure, and derivative-use lease. Internal states can leak training data, user data, security-sensitive triggers, or proprietary weights; interpretability is not a privacy or custody exception.

72.10.2 2. Preserve extraction assumptions

An extracted object records the method that made it visible. For sparse autoencoders, that includes dictionary size, sparsity and reconstruction trade-offs, dead latents, optimization state, seeds, and feature-selection rules. For circuit or attribution work, it includes the computational graph, replacement-model assumptions, edge thresholds, pruning, and reconstruction error. For probes, it includes target construction, regularization, leakage controls, and whether the probed information is actually used by the model.

Method outputs remain method-relative. Stable identifiers may track a feature through one experiment, but they do not imply that nature supplied a canonical feature basis.

72.10.3 3. Separate labels from mechanisms

A semantic label is an adjudicated hypothesis. It carries positive examples, hard counterexamples, nearest alternative labels, annotator or model identity, agreement, selection process, uncertainty, and expiry. Automated explanation models cannot certify their own labels. A useful label may still be incomplete, anthropomorphic, or downstream of the computation that matters.

72.10.4 4. Demand predictive and causal challenge

The study first asks whether the internal hypothesis predicts held-out behavior better than simple behavioral signals. It then asks distinct causal questions:

  • Necessity: does removing or replacing the component reduce the behavior?
  • Sufficiency: does adding or activating it induce the behavior in an otherwise appropriate context?
  • Specificity: are unrelated capabilities and neighboring behaviors spared?
  • Dose response: does the outcome vary coherently with intervention strength?
  • Mediation: does the proposed path explain the effect better than plausible alternative paths?

Ablation, patching, counterfactual replacement, and steering have different intervention semantics. Agreement across them is stronger than any one method, but still does not establish whole-model causal completeness.

72.10.5 5. Keep coverage and residuals visible

Reconstruction loss is not semantic coverage. Variance explained is not causal coverage. A graph that reproduces one answer is not a complete algorithm. The packet therefore keeps reconstruction, behavioral prediction, causal effect, feature stability, unexplained residual, and method disagreement as separate quantities. Missing or contradictory paths remain first-class residuals.

72.10.6 6. Govern activation policies as model changes

An activation monitor or steering rule is a new component, not a free interpretation by-product. It receives a version, authority ceiling, trigger policy, expected benefit, false-positive/false-negative envelope, latency and resource cost, protected-capability set, side-effect tests, fallback, compensation or rollback route, monitoring period, and retirement rule.

White-box evidence may block, narrow, quarantine, or demand more review before it may widen authority. A release decision remains with readiness and assurance owners after independent behavioral qualification.

72.10.7 Probes, sparse dictionaries, circuits, and causal challenge

White-box methods answer different questions and should not be blended into one “interpretability score.”

Diagnostic probes ask whether information usable by a chosen probe family is present in an activation. Accuracy depends on layer, token selection, regularization, data, negative controls, and probe capacity. A powerful probe may learn a task from weak traces; a linear probe may miss information available nonlinearly. Selectivity controls, label permutations, matched random features, held-out concepts, and interventions are needed before performance becomes mechanism evidence. Hewitt and Liang’s control-task method makes the core confound explicit: a probe that predicts the target and a randomized word-type mapping may be demonstrating its own memorization capacity rather than a simple property of the representation. Report target accuracy and control accuracy together; do not let a slightly higher target score hide much lower selectivity.

Sparse autoencoders and dictionary learning approximate activations with a learned feature dictionary and sparse codes. The record includes training distribution, hook point, normalization, dictionary width, sparsity penalty, reconstruction error, dead and dense feature rates, seed and checkpoint stability, feature splitting and merging, and coverage of the target behavior. A named or apparently monosemantic feature is an interpretation hypothesis. Readable exemplars do not prove that it causes the behavior, exhausts the concept, or survives model change.

The same extractor can support a convincing but distribution-bound story. Bolukbasi et al. held a BERT direction fixed and found that its apparently coherent top-activating examples changed across corpora. The governance lesson is not that every direction is meaningless. It is that global, dataset-level, and local interpretations are competing constructs until materially different corpora and deployment-relevant shifts distinguish them. A random same-corpus split can preserve the very geometry that produced the illusion.

Circuit and attribution methods propose a computational subgraph, path, or set of components relevant to an output. Local linearity, patching distribution, baseline, decomposition, attention interpretation, and graph extraction are assumptions in the claim. Sufficiency and necessity are separate: preserving a circuit may recreate behavior without showing it is the only route, while ablation can break behavior through broad distributional damage rather than removal of the proposed mechanism.

Construct validity joins these methods to the deployment question. “Deception,” “uncertainty,” “planning,” “harmful intent,” or “memorization” is not validated by a convenient label set alone. The case supplies competing constructs, convergent and discriminant tests, relevant opportunities, known-positive and known-negative controls, and label provenance. Failure to detect a construct with an insensitive instrument cannot become evidence of absence.

Causal challenge changes the proposed mechanism while preserving as much else as possible. Activation patching, ablation, steering, feature clamping, counterfactual input edits, path intervention, targeted fine-tuning, and mechanistically chosen adversarial examples have different semantics. The record asks whether predicted behavior changed, unrelated capabilities moved, redundant paths compensated, the intervention was on-distribution, and an independently implemented evaluator agreed.

72.10.8 Comparative method matrix

No method occupies every rung of the evidence ladder. The selection rule begins with the question and admits the cheapest competent instrument; it does not begin with a favored visualization.

Method family Directly observes or estimates Strong comparator or control Characteristic residual Maximum pre-causal inference
Diagnostic probe Information decodable by a specified probe family from a specified activation Control task, label permutation, simple and regularized probes, matched random features Probe learning, leakage, layer/token selection, nonlinear miss The named information is decodable by this instrument in this distribution
Exemplar or automated label A compact hypothesis about inputs associated with a unit or feature Blinded counterexamples, alternate labels, multiple corpora, human and independent-model review Selection bias, dataset geometry, anthropomorphic or overly broad labels The label predicts selected held-out activations within its tested envelope
Sparse dictionary / SAE A reconstruction of activations through learned sparse coordinates Raw-neuron, PCA/ICA, supervised dictionary, architecture, width, sparsity, seed, and checkpoint sweeps Dead latents, absorption, splitting, merging, missed nonlinear structure The extractor supplies a useful coordinate system under stated reconstruction and sparsity conditions
Circuit / attribution graph A method-relative subgraph or path relevant to an output Behavioral baseline, random or frequency-matched graph, alternate decomposition, reconstruction check Omitted paths, replacement-model error, off-distribution patching The proposed components predict or reconstruct the scoped behavior under the extraction assumptions
Causal intervention Change in behavior after ablation, patching, replacement, or steering Sham and off-target interventions, dose response, mediation alternatives, protected-capability tests Broad damage, redundant paths, unnatural states, evaluator blindness A scoped necessity, sufficiency, specificity, or mediation result for the exact intervention
Activation policy A monitored or modified model component used in operation Behavior-only control, abstention, independent safety path, rollback and retirement drill Gaming, false alarms, latency, privacy exposure, drift, collateral capability change A bounded operational decision after independent qualification—not general understanding

Sparse-autoencoder comparisons need an additional measurement discipline. SAEBench demonstrates why reconstruction at a chosen sparsity cannot stand in for concept detection, interpretability, disentanglement, or downstream intervention behavior, and why those dimensions should not be averaged into one global score. Its practitioner guidance motivates width and sparsity sweeps with directly comparable baselines. Chanin’s later reliability audit then challenges constituent metrics through independent reseeds, training-trajectory discrimination, synthetic ground truth, degraded controls, and an oracle. These results are complementary: use a multi-dimensional benchmark, and require each dimension to prove that it can distinguish the cases it is asked to rank.

72.10.9 Evidence ladder and stop rules

The white-box ladder is deliberately noninheritant. Passing a lower rung does not imply the next one.

  1. Decodability: a specified instrument extracts a signal beyond matched controls.
  2. Stability: the signal survives seeds, batches, checkpoints, annotators, and appropriate method variation.
  3. Reconstruction: the extracted representation preserves the relevant activation or behavior with its residual visible.
  4. Semantic construct validity: positive, negative, convergent, discriminant, and cross-distribution tests support the named concept over plausible alternatives.
  5. Causal necessity or sufficiency: a controlled intervention moves the target outcome in the predicted direction.
  6. Mediation and specificity: the proposed path explains the effect better than alternatives while sparing off-target and protected capabilities.
  7. Cross-distribution transfer: the result survives the material environments, populations, model changes, and opportunities named in the claim.
  8. Intervention utility: the method improves a bounded diagnosis or action against a strong behavior-only baseline after collateral and total cost.
  9. Policy authority: an independent owner admits a versioned monitor or intervention with expiry, fallback, rollback, and a narrow authority ceiling.

A failed probe, SAE metric, circuit extractor, or intervention stops at the instrument, implementation, construct, or exact-mechanism level established by its controls. It cannot become a broad negative result about information, features, interpretability, or model safety. Conversely, a clean lower-rung positive remains diagnostic until the causal, transfer, utility, and authority rungs are separately earned.

The strongest result is a scoped chain: seeded mechanisms are recovered; plausible alternatives are distinguished; held-out behavior or failure is predicted; findings survive method and seed variation; causal interventions move behavior as predicted without disqualifying collateral damage; and the evidence improves a governed decision while preserving uncovered cases. Even that chain is evidence about one model, method, behavior, and domain—not transparent access to the model’s “true thoughts.”

flowchart TD
    M["Registered model and behavior"] --> P["Probe evidence"]
    M --> D["Sparse dictionary evidence"]
    M --> C["Circuit / attribution hypothesis"]
    P --> V["Construct-validity challenge"]
    D --> V
    C --> V
    V --> I["Causal interventions"]
    I --> O["Predicted and collateral outcomes"]
    O --> G{"Decision-relevant and stable?"}
    G -->|yes, scoped| A["Governed activation policy"]
    G -->|no or unknown| R["Residual / abstain / new test"]

72.11 Interfaces

No single interface is allowed to collapse internal observation into a release decision. Each handoff preserves the exact producer, consumer, artifact, authority, freshness, failure, and residual state. Missing identity or an unresolved material change closes the packet rather than inviting a best-effort interpretation from stale data.

The interfaces form an evidence chain rather than a pipeline of automatic promotion. The model registry tells the capture service which artifact was observed; custody and privacy constrain why the observation exists and who may see its derivatives; the feature registry retains method-relative identities; and the intervention runner reports actual behavioral and collateral effects. The evaluator then tests the result without inheriting the extraction method’s labels as ground truth. Only the evidence ledger may expose a bounded packet to readiness and release owners, and it exposes residuals and expiry alongside any favorable finding. A missing link therefore produces an incomplete or rejected packet, not an invitation for a downstream consumer to reconstruct the missing assumption from context.

Interface Required handoff
Model and checkpoint registry Exact artifact identity, lineage, architecture, tokenizer, adapters, and material-change events
Custody and privacy Activation-access purpose, audience, retention, disclosure, subject/data obligations, and revocation
Activation capture Sites, tensors/events, precision, hooks, runtime effects, lossiness, and capture receipts
Feature/circuit registry Method-relative identifiers, extraction state, labels, alternatives, examples, and expiry
Intervention runner Target, operation, dose, controls, observed behavioral changes, side effects, and recovery
Behavioral evaluator Held-out tasks, output identity, baseline, uncertainty, negative controls, and task limits
Evidence ledger Packet disposition, maximum inference, residuals, challenges, consumers, and transition eligibility
Readiness and release Policy candidate, protected capabilities, rollback, monitor period, unresolved defeaters, and decision authority

72.12 Invariants

The invariants preserve the distance between seeing a model’s internals and understanding them. They make every semantic and causal step challengeable, keep approximation and unexplained computation visible, and expire results when the artifact or method changes. They also prevent the governance path from using interpretability as a privileged shortcut: white-box evidence may narrow or block action early, but cannot grant authority, move support, or certify a release without the same behavioral and operational owners that govern any other evidence.

  • Every internal claim names the exact model, checkpoint, runtime, population, capture, method, and time boundary.
  • Captured state, extracted object, semantic label, predictive association, causal result, coverage claim, activation policy, and released artifact remain distinct states.
  • Association is never represented as necessity, sufficiency, mediation, or completeness.
  • Every label retains examples, counterexamples, alternatives, authorship, selection lineage, uncertainty, and expiry.
  • Every consequential causal claim has a matched intervention control, a behavioral cross-check, and a side-effect denominator.
  • Reconstruction error, dead or missing features, graph pruning, unexplained behavior, and method disagreement remain visible.
  • Material changes to weights, topology, tokenizer, adapters, routing, method, threat model, or input population expire affected packets.
  • Internal access obeys custody, privacy, security, and purpose limits; derived artifacts do not escape those obligations.
  • White-box evidence cannot promote support, widen runtime authority, certify safety, or authorize release by itself.
  • Rejected, unstable, contradicted, harmful, and null results remain in the denominator with retries, analyst choices, and post-hoc rescue visible.

72.13 Failure modes

The threat model includes ordinary scientific error, motivated interpretation, model and evaluator gaming, malicious access, and institutional pressure to turn a legible story into a release justification.

These failures can arise even when every participant is acting in good faith. An analyst can select unusually coherent examples, a probe can exploit unused information, an intervention can disrupt neighboring computation, and a model update can invalidate yesterday’s feature identity. Strategic systems add stronger pressures: they may route around a monitored feature, shape visible activations, exploit evaluator assumptions, or make a suppressive intervention look like genuine risk removal. Organizational incentives can then complete the failure by converting a favorable picture into release pressure. The response is not to average these uncertainties away; it is to retain contradictory methods, counterexamples, side effects, expiry, and the attempted authority movement in the packet.

Failure family Example Required response
Feature hallucination A sparse feature receives a coherent label from selected examples but fails hard counterexamples Narrow or reject the label; preserve alternatives and selection lineage
Probe leakage A classifier decodes information that the model does not use Require causal challenge and leakage controls
Polysemantic collapse One feature combines materially different behaviors Split the claim or retain an explicit mixed/unknown state
Feature splitting/merging Identity drifts across seeds, checkpoints, or dictionary sizes Expire stable identity and report matching uncertainty
Circuit incompleteness A graph omits alternative paths or relies on a high-error replacement model Bound inference by reconstruction and causal residuals
Intervention collateral Steering changes neighboring capabilities, calibration, or safety behavior Block policy admission; restore or retire the intervention
Method monoculture Extraction and evaluator share assumptions or implementation defects Add a materially different method and evaluator
Checkpoint drift An old interpretation is applied after fine-tuning, merging, or quantization Expire and requalify the packet
Cherry-picking Prompts, features, labels, or intervention strengths are selected after outcomes Close claim-bearing inference or open a fresh preregistered campaign
Monitor gaming Training or deployment learns to hide from the chosen feature monitor Use adversarial evaluation, randomized challenge, and behavioral fallback
Access leakage Activations expose protected user or training information Revoke access, preserve incident lineage, and route to privacy/custody owners
Evidence-to-authority escalation A favorable internal result silently justifies deployment Reject the transition and preserve the attempted authority widening

72.14 Strongest objection

The strongest objection is that white-box evidence may add fragile complexity without outperforming a good behavioral safety stack. Feature bases are not unique, labels are contestable, causal interventions can be destructive, and models can change faster than interpretations can be maintained. A behavior- only system with diverse tests, runtime monitors, canaries, conservative authority, and rollback may be cheaper and more reliable.

That baseline must be treated seriously. White-box governance earns its place only when it adds decision-relevant information before or beyond the behavioral route—for example, earlier detection of a known mechanism, better localization of a regression, more selective intervention, or improved prediction of a held-out behavior—without worse collateral effects or unacceptable total cost. If a tuned behavioral baseline matches it, the white-box component should not be retained merely because it is intellectually appealing.

72.15 Minimum Viable Implementation

Use one tractable open model, one well-defined behavior family, and at least two materially different internal methods. Freeze the model and task population; capture exact activation and method lineage; train or construct the methods on development data; produce label packets with hard counterexamples; and keep a sealed natural held-out set.

The implementation must include:

  • an exact model/checkpoint and activation-capture registry;
  • two extraction routes, plus random-feature and behavior-only controls;
  • positive controls where an injected or otherwise independently known mechanism should be recovered;
  • predictive held-out tests across prompts, seeds, and at least one nearby checkpoint;
  • ablation, patching or replacement, and steering with matched sham controls;
  • reconstruction, stability, causal effect, specificity, collateral-damage, unexplained-residual, latency, compute, and analyst-cost records;
  • an independently implemented evaluator that does not consume the extraction method’s labels as ground truth; and
  • a policy packet whose only initial authority is to raise scrutiny, block, or route to further qualification.

The minimum implementation is intentionally narrower than “interpret the model.” Its output is a reviewable evidence packet and a decision about whether white-box work added value in one exact envelope.

72.16 Latent knowledge, anomaly evidence, and training attribution

Eliciting latent knowledge asks a sharper question than ordinary interpretability: can an evaluator recover what a model internally knows when its normal report may be optimized to say something else? The ELK agenda separates the predictor’s latent world information from a possibly untrusted reporter [@ext_elk_report_2021]. For governance, this means a probe or explanation cannot be admitted merely because it agrees with the model’s answer. The evidence packet records the target latent variable, reporter channel, supervision source, competing reporter hypotheses, causal interventions, held-out failures, and the conditions under which the model may not represent the fact at all.

Mechanistic anomaly detection is a related but narrower lane. It compares internal-state traces against versioned reference distributions and mechanistic hypotheses, then reports deviations without assigning intent. Useful signals include unexpected feature activation, pathway substitution, dormant-trigger activation, representation drift, or a mismatch between claimed task and active circuit. The mechanism must test benign shifts, adaptive evasion, detector transfer, false-alarm burden, and whether interventions on the alleged anomaly change behavior. An anomaly is evidence for investigation, not proof of deception.

Training-data attribution supplies another evidence family. Influence functions estimate how upweighting a training point would change a prediction or loss through a local curvature approximation [@ext_influence_functions_2017]. The stack stores the checkpoint, target, parameter subset, damping and solver settings, approximation error, candidate denominator, and deletion or retraining comparison. Agreement across influence estimates, nearest-neighbor methods, gradients, and controlled retraining is stronger than any one attribution.

These methods fail through inaccessible latent variables, a reporter that shares the same deception incentive, feature aliasing, detector overfitting, non-convex curvature, poor Hessian approximations, and causal stories inferred from correlation. The explicit nonclaim is that ELK remains unsolved, anomaly scores do not prove malicious intent, and approximate influence does not establish causal responsibility, privacy leakage, or successful unlearning.

72.17 Formalization hooks

A finite formal target is useful for packet routing. A model can enforce that a policy candidate cannot become release-eligible when checkpoint identity is missing, causal status is merely observational, required negative controls failed, side effects are unresolved, residual coverage is absent, or authority would widen. It can also prove that material model change expires a packet and that an expired packet cannot authorize an intervention.

Those are record and workflow properties. A Lean theorem over a finite packet type does not prove that a discovered feature is real, that a circuit is the model’s algorithm, that an intervention isolates the intended cause, that coverage is complete, or that a steered model is safe. Formalization is omitted for those empirical and semantic claims because pretending otherwise would be proof theater.

The implemented target lean:white_box.evidence_never_grants_authority owns the finite non-escalation route and the reachable transaction around it. Arbitrary successful runs preserve exact packet, model, checkpoint, method, population, and version identity; active authority never increases; support and external- effect authority never appear; traces remain valid; and event batches compose. The companion target lean:white_box.invalid_packet_rejected owns packet admission, bounded-causal checks, independent review, policy routing, and receipt-complete zero-residual consumption. Stale lineage, missing assumptions, absent negative controls, and unsupported causal state cannot reach policy routing. scripts/validate_white_box_evidence_contract.py independently reconstructs the six-event lifecycle and rejects 51 identity, receipt, stage, time, authority, route, residual, support, effect, and packet-boundary mutations. These are authored record and workflow properties: 36 checked theorem declarations cannot establish scientific interpretation, completeness, intervention safety, evaluator truth, or cross-model transfer.

72.18 Competent argument-exit campaign

The campaign must give white-box methods a fair chance without allowing them to grade themselves.

The exact preregistration is experiments/white_box_argument_exit/preregistration.json. It freezes two independently owned method families, a tuned behavior-only baseline, split custody, seven competence gates, a six-step rescue ladder, nine joint outcomes, an N3 exact-negative ceiling, and a 62 GiB resource-isolation floor. It is protocol-ready and resource-isolated, not executed. It cannot consume P2 storage, tasks, labels, outcomes, or retry decisions. No model-internal outcome was opened, and a failed competence gate ends as instrument- or implementation-inadequate rather than as evidence against interpretability.

72.18.1 Frozen question

For a named model, behavior family, and material-change envelope, do two internal methods add held-out predictive or intervention value beyond strong behavioral monitors at acceptable collateral and lifecycle cost?

72.18.2 Comparators

  • tuned behavior-only monitoring and adversarial evaluation;
  • a linear probe with leakage controls;
  • random and frequency-matched feature controls;
  • each white-box method independently;
  • the combined method only if its composition is frozen prospectively; and
  • sham, off-target, and dose controls for interventions.

72.18.3 Competence gates

Before held-out opening, the implementation must recover positive controls, beat chance on development tasks without ceiling effects, demonstrate activation of the intended method, reproduce source-compatible smoke behavior, pass label-counterexample review, show evaluator sensitivity to known effects, and complete a bounded rescue ladder for optimization, dictionary size, thresholds, intervention dose, and checkpoint selection. A failed method, inactive intervention, blind evaluator, or broken positive control is N0/N1, not evidence against interpretability.

72.18.4 Outcomes

Measure predictive discrimination and calibration, causal effect with uncertainty, necessity/sufficiency asymmetry, specificity, feature identity stability, reconstruction and causal residuals, cross-checkpoint transfer, false positive and false negative rates, protected-capability regressions, latency, compute, storage, analyst/reviewer time, privacy exposure, and policy decision value. Preserve every failed feature, label, seed, retry, intervention, and excluded case.

72.18.5 Inference ceiling

A local positive can support only the exact model, methods, behavior, and checkpoint envelope. Cross-architecture or safety-general claims require materially separate implementations, evaluators, organizations, and transfer settings. A competent local negative may reject one frozen packet or mechanism; it cannot show that mechanistic interpretability as a field is useless. No result licenses covert activation monitoring, unbounded steering, deployment, or support promotion without the owning reviews.

72.19 Mature Research Target

The research target is a substrate-aware internal-evidence layer that can track method-relative mechanisms across checkpoints and architecture families, quantify causal and unexplained residuals, predict consequential behavioral change before release, and govern selective interventions with independent qualification and effect-complete rollback. It should expose when a simpler behavioral baseline dominates and preserve incompatible interpretations rather than averaging them into confidence.

That target is not a current result. “Beyond the state of the art” here names a falsifiable program: decision-relevant early warning or selective control with lower joint risk and total cost than strong behavioral alternatives. Failure to clear that comparison should remove or narrow the mechanism.

A mature target architecture would also make interpretation maintenance an explicit systems cost. It would detect when a checkpoint edit, tokenizer change, new training distribution, quantization pass, router update, or substrate replacement invalidates a feature identity; schedule only the requalification needed by affected consumers; and preserve historical packets so a reviewer can distinguish genuine mechanism continuity from a convenient new label. Methods would be selected by the questions they can answer rather than by one global interpretability score. Conflicting explanations would remain queryable, and an evidence consumer could see exactly which conclusion survives the conflict.

The architectural endpoint is not maximal introspection. It is a calibrated choice among internal evidence, behavioral evidence, conservative fallback, and abstention. Some systems or decisions may warrant no activation access; others may justify it only inside a secure research boundary. The mature layer would make both outcomes legitimate and compare the full lifecycle burden of maintaining internal evidence against the risk reduction it actually buys.

No such campaign has passed here. The present repository supplies an argument, source notes, a packet design, and bounded future tests. That evidence ceiling leaves the chapter’s core at argument until a competent study establishes decision value in the exact scope claimed and an accepted evidence transition records that movement.

72.20 Codex test plan

Test Purpose Status
Internal-evidence packet schema and mutation suite Accept complete model/method/intervention packets and reject missing identity, causal-status laundering, hidden residuals, stale packets, and authority widening. implemented; record-shape only; python3 scripts/validate_white_box_evidence_contract.py rejects 12 semantic mutations
Interpretation stability and positive-control suite Challenge feature and circuit identities across seeds, prompts, checkpoints, extraction methods, counterexamples, and independently known mechanisms. protocol-ready; not executed; protected outcomes closed
Causal intervention and collateral benchmark Compare observation, probes, ablation, patching, replacement, and steering against sham/off-target controls and protected-capability outcomes. protocol-ready; not executed; outcomes remain closed until competence and resource gates pass
Behavioral-baseline and governance-value comparison Test whether internal evidence improves held-out prediction or selective control beyond tuned behavioral monitors after latency, compute, analyst effort, privacy, and maintenance costs. protocol-ready; not executed; no result or support effect

72.21 Source crosswalk

Source ID Use in this chapter Boundary
ext_transformer_circuits_2021 Mathematical vocabulary for transformer components and circuit-level analysis Framework and worked analyses do not establish complete or universal model mechanisms
ext_monosemanticity_2023 Dictionary-learning route for extracting more interpretable features Source-reported feature examples do not establish a canonical semantic basis, causal completeness, or local reproduction
ext_scaling_sparse_autoencoders_2024 Scalable sparse-autoencoder training and evaluation comparator, including reconstruction/sparsity and dead-latent concerns Scaling a feature extractor does not establish semantic faithfulness, causal coverage, safety, or transfer
ext_circuit_tracing_2025 Replacement-model attribution graphs, perturbation validation, and reconstruction-error boundary Source-reported graph results do not establish whole-model understanding, safe steering, or local reproduction
ext_probe_control_tasks_2019 Control tasks and selectivity for separating probe performance from a probe family’s capacity to memorize ELMo linguistic-probe results do not establish a universal control task, causal use of decoded information, or a local result
ext_interpretability_illusion_bert_2021 Cross-dataset challenge and the distinction among global, dataset-level, and local concept structure A BERT sentence-embedding case does not show that every feature is illusory or that causal interpretations cannot transfer
ext_saebench_2025 Multi-metric SAE comparison across reconstruction, concept detection, interpretability, disentanglement, and downstream uses Standardization and source-reported rankings do not make every metric reliable or establish semantic and causal faithfulness
ext_sae_benchmark_reliability_2026 Reseed, discriminability, synthetic-ground-truth, degraded-control, and oracle audit of selected SAE metrics Metric- and setting-scoped counterevidence cannot refute SAEs, SAEBench as a whole, or interpretability as a field
deterministic_capability_compilation Candidate-specific validation, residual escrow, authority ceilings, and reification boundaries for learned capability objects Corben-authored design program; no foundry implementation or preservation result is inherited
kernel_english_residual_compiler Protected objects, explicit residuals, round-trip checks, and versioned migration as internal-evidence design analogies Corben-authored proposal; no efficiency, fidelity, or model-internal result is inherited
qcsa_whitepaper Stable semantic identity and evidence-bearing hypergraph concepts for versioned feature/circuit records Existing local QCSA results do not test mechanistic interpretability and cannot promote this chapter
platonic_world_model Proposition/attestation/proof separation and branch-protected semantic continuity Conceptual architecture; it does not solve grounding or validate internal interpretations

72.21.1 Manifest source assignment reconciliation

These rows keep White-Box Evidence, Interpretability, and Activation Governance’s manifest assignments visible at their recorded review boundary. Passage review does not establish local reproduction, performance, safety, deployment, or support-state movement.

Source Intake role Boundary
ext_elk_report_2021 Passage-reviewed comparator: Eliciting Latent Knowledge. Frames eliciting latent knowledge as the problem of recovering what a model internally knows when its ordinary answer channel may be untrusted. ELK is a research problem and strategy space, not a solved truthful-reporting mechanism or a guarantee that latent knowledge is represented accessibly. No local implementation, reproduction, performance, safety, deployment, support-state, or ASI result is established by this reconciliation row.
ext_influence_functions_2017 Passage-reviewed comparator: Understanding Black-box Predictions via Influence Functions. Provides a first-order method for estimating how training points affect a model prediction or loss, useful as one attribution hypothesis. Influence estimates depend on differentiability, curvature approximation, local linearization, checkpoint identity, and implementation; they do not prove causal responsibility or deletion. No local implementation, reproduction, performance, safety, deployment, support-state, or ASI result is established by this reconciliation row.

72.22 Current evidence and non-claims

The repository currently contains a manifest packet, this integrated reader chapter, fourteen source mappings, schemas/white_box_evidence_packet.schema.json, one explicitly expired record-shape fixture, an independently implemented semantic validator, a checked finite Lean routing model, and a prospectively frozen argument-exit protocol. It contains no locally executed sparse-autoencoder training, circuit tracing, activation capture, label study, causal intervention, steering trial, model monitor, cross-checkpoint analysis, independent reproduction, or chapter-specific support transition.

Accordingly, this chapter does not claim that:

  • internal features have unique or human-legible meanings;
  • a cited method reproduces locally or transfers to other substrates;
  • activation steering is safe, robust, private, or resistant to gaming;
  • white-box evidence is superior to behavioral evaluation;
  • any model is understood, aligned, ready, or safe; or
  • the ASI Stack has achieved mechanistic transparency, AGI, or ASI.

72.23 Summary

White-box evidence is valuable precisely because behavior alone can leave important uncertainty. It is dangerous for the same reason: internal access creates stories faster than it creates justified explanations. The ASI Stack therefore treats capture, extraction, labeling, prediction, causality, coverage, policy, and release as separate states. Internal evidence may raise scrutiny early; it earns trust only through stable identity, causal challenge, behavioral comparison, visible residuals, independent evaluation, and bounded authority.

A credible packet binds the exact model and method, keeps semantic labels hypothetical, distinguishes necessity from sufficiency, measures collateral effects, and expires after material change. It also preserves failed features, contradictions, analyst choices, and unexplained computation. This makes a white-box result challengeable by behavioral evaluators, custody and privacy owners, safety cases, and release authorities without allowing any of them to inherit more than the result earned.

The decisive test is practical rather than aesthetic: internal evidence must improve a bounded prediction, diagnosis, or intervention decision beyond a strong behavioral baseline while retaining acceptable uncertainty and total cost. Until that evidence exists, model-internal visibility remains a promising research instrument and a governed source of scrutiny—not a certificate of understanding, safety, readiness, or intelligence.

72.24 Handoff

White-box packets supply a constrained evidence input; they do not replace the owners of behavior, assurance, or release. Capability Thresholds and Deployment Commitments receives the exact internal-evidence disposition, maximum inference, uncertainty, expiry, side effects, and residuals and decides whether a scoped capability assessment crosses a prospectively declared commitment. Adversarial Evaluation, Safety Cases, and Readiness can challenge or consume the packet later, but no downstream owner inherits causal completeness or release authority from internal visibility alone. The implemented packet route can preserve, restrict, escalate, reject, or expire evidence, but it cannot grant execution or release authority.