83 Methods: Structure, Reviewability, and Warding
84 Methods: Structure, Reviewability, and Warding
This methods note is generated from the project evidence ledger. It states only what the recorded artifacts can support, keeps null results visible, and separates local controls from project-owned model runs.
84.1 Summary Finding
The current evidence supports a narrow thesis: structured prompts reliably improve reviewability, outcome-marker effects are task-dependent, the current model-produced execution slice does not yet separate weak from repaired prompts, and warded prompts have the strongest measured protective signal against defanged injection fixtures.
84.2 Evidence Table
| Claim | Recorded Result | Evidence Class | Limitation |
|---|---|---|---|
| Prompt structure improves reviewability | 21/24 surface-tier cells show positive repaired-minus-weak reviewability delta; range -2.0 to 10.0. | project_owned_model_run plus rubric scoring | Reviewability is inspectability, not proof that execution outcomes changed. |
| Outcome markers are task-dependent | 8 positive, 15 ties, 1 regressions across 24 surface-tier cells. | project_owned_model_run plus marker scoring | Markers are useful but weaker than execution-graded artifacts. |
| Current model-produced execution slice does not separate weak from repaired | Weak model artifacts passed 3/3; repaired model artifacts passed 3/3. | project_owned_model_run plus local deterministic execution | The slice is small and currently covers only recorded Claude Code artifact runs. |
| Bench v4 local seed ladder discriminates fixture contracts | Local weak seed artifacts passed 0/25; repaired seed artifacts passed 25/25 across 5 rungs. | local_deterministic_execution | These are project-authored control artifacts, not model-provider results. |
| Bench v4 now has model-surface hardness artifacts | On codex-cli-default, weak model artifacts passed 1/25 and repaired model artifacts passed 1/25; extraction failures: 5/50. |
project_owned_model_run plus local deterministic execution | The hidden grader is fixture-local and the current surface set is still project-operated, not independent external evidence. |
| Warding has the strongest measured protective signal so far | Baseline runs recorded 31/48 guardrail losses; warded runs recorded 0/48 guardrail losses under the same score/redaction rule. | project_owned_model_run | Current real A/B evidence covers claude-code-safe, codex-cli-default; the standard warded jailbreak-resilience suite and hardness ladder still need additional non-Codex or reviewer-supplied surface runs. |
| Ward-limb science has a structural seed | 6 local ablation variants score attack resistance, utility, audit quality, and overrefusal. | local_deterministic_control | This does not simulate model behavior and is not model-provider evidence. |
84.3 Nulls And Non-Wins
- Execution-grade model artifacts currently show no weak-versus-repaired separation on the recorded slice.
- Outcome-marker scoring includes ties and regressions, not only wins.
- Bench v4 now records model-produced artifacts against the hardness ladder, but current runs still need more surfaces and future repetitions before they support broad generalization.
- The canon-review queue is prepared, but canonical promotion remains pending human maintainer signoff.
- Package-index release materials are prepared, but public upload remains pending human action.
84.4 Next Falsification Steps
- Replay the five Bench v4 hardness rungs on additional non-Codex or reviewer-supplied model surfaces.
- Run the Claude Code standard warded jailbreak-resilience suite and replay ward-limb ablations on real model surfaces.
- Publish package-index smoke checks only after a human performs the upload.
- Record maintainer decisions before promoting any usage-earned rune to canonical.
Raw evidence: surface-comparison.json, model-execution-results.json, hardness-v4 results, hardness-v4 model-surface results, ab-results.json, ward-science-results.json