83  Methods: Structure, Reviewability, and Warding

84 Methods: Structure, Reviewability, and Warding

This methods note is generated from the project evidence ledger. It states only what the recorded artifacts can support, keeps null results visible, and separates local controls from project-owned model runs.

84.1 Summary Finding

The current evidence supports a narrow thesis: structured prompts reliably improve reviewability, outcome-marker effects are task-dependent, the current model-produced execution slice does not yet separate weak from repaired prompts, and warded prompts have the strongest measured protective signal against defanged injection fixtures.

84.2 Evidence Table

Claim Recorded Result Evidence Class Limitation
Prompt structure improves reviewability 21/24 surface-tier cells show positive repaired-minus-weak reviewability delta; range -2.0 to 10.0. project_owned_model_run plus rubric scoring Reviewability is inspectability, not proof that execution outcomes changed.
Outcome markers are task-dependent 8 positive, 15 ties, 1 regressions across 24 surface-tier cells. project_owned_model_run plus marker scoring Markers are useful but weaker than execution-graded artifacts.
Current model-produced execution slice does not separate weak from repaired Weak model artifacts passed 3/3; repaired model artifacts passed 3/3. project_owned_model_run plus local deterministic execution The slice is small and currently covers only recorded Claude Code artifact runs.
Bench v4 local seed ladder discriminates fixture contracts Local weak seed artifacts passed 0/25; repaired seed artifacts passed 25/25 across 5 rungs. local_deterministic_execution These are project-authored control artifacts, not model-provider results.
Bench v4 now has model-surface hardness artifacts On codex-cli-default, weak model artifacts passed 1/25 and repaired model artifacts passed 1/25; extraction failures: 5/50. project_owned_model_run plus local deterministic execution The hidden grader is fixture-local and the current surface set is still project-operated, not independent external evidence.
Warding has the strongest measured protective signal so far Baseline runs recorded 31/48 guardrail losses; warded runs recorded 0/48 guardrail losses under the same score/redaction rule. project_owned_model_run Current real A/B evidence covers claude-code-safe, codex-cli-default; the standard warded jailbreak-resilience suite and hardness ladder still need additional non-Codex or reviewer-supplied surface runs.
Ward-limb science has a structural seed 6 local ablation variants score attack resistance, utility, audit quality, and overrefusal. local_deterministic_control This does not simulate model behavior and is not model-provider evidence.

84.3 Nulls And Non-Wins

  • Execution-grade model artifacts currently show no weak-versus-repaired separation on the recorded slice.
  • Outcome-marker scoring includes ties and regressions, not only wins.
  • Bench v4 now records model-produced artifacts against the hardness ladder, but current runs still need more surfaces and future repetitions before they support broad generalization.
  • The canon-review queue is prepared, but canonical promotion remains pending human maintainer signoff.
  • Package-index release materials are prepared, but public upload remains pending human action.

84.4 Next Falsification Steps

  1. Replay the five Bench v4 hardness rungs on additional non-Codex or reviewer-supplied model surfaces.
  2. Run the Claude Code standard warded jailbreak-resilience suite and replay ward-limb ablations on real model surfaces.
  3. Publish package-index smoke checks only after a human performs the upload.
  4. Record maintainer decisions before promoting any usage-earned rune to canonical.

Raw evidence: surface-comparison.json, model-execution-results.json, hardness-v4 results, hardness-v4 model-surface results, ab-results.json, ward-science-results.json