73 Surface Comparison
74 Surface Comparison
This comparison keeps project-owned model transcripts separate from repository-owned deterministic tooling. It does not count local deterministic graders as independent external adoption. Model rows are split by tier where per-cell evidence exists so clean and trap results do not disappear into an aggregate.
74.1 Field-Spell Matrix
| Case | Surface | Tier | Outcome Delta | Reviewability Delta | Runs | Evidence | Limitation |
|---|---|---|---|---|---|---|---|
| Refactor Without Breaking Behavior | claude-code-safe | clean | -1.0 | 2.0 | 1 weak / 1 repaired | recorded transcripts and marker outcome scoring | model-surface evidence for the named local CLI/tool configuration; reviewability scoring remains secondary |
| Refactor Without Breaking Behavior | claude-code-safe | trap | 2.0 | 6.0 | 1 weak / 1 repaired | recorded transcripts and marker outcome scoring | model-surface evidence for the named local CLI/tool configuration; reviewability scoring remains secondary |
| Refactor Without Breaking Behavior | codex-cli-default | clean | 0.0 | 2.3 | 3 weak / 3 repaired | recorded transcripts and marker outcome scoring | model-surface evidence for the named local CLI/tool configuration; reviewability scoring remains secondary |
| Refactor Without Breaking Behavior | codex-cli-default | trap | 1.3 | 4.0 | 3 weak / 3 repaired | recorded transcripts and marker outcome scoring | model-surface evidence for the named local CLI/tool configuration; reviewability scoring remains secondary |
| Refactor Without Breaking Behavior | local-deterministic-grader | n/a | repaired artifact passes where weak artifact fails | not applicable | fixture-local artifact execution where grader exists | repository-owned deterministic tool surface, not independent model evidence | |
| Incident Diagnosis Without Fake Certainty | claude-code-safe | clean | 0.0 | 1.0 | 1 weak / 1 repaired | recorded transcripts and marker outcome scoring | model-surface evidence for the named local CLI/tool configuration; reviewability scoring remains secondary |
| Incident Diagnosis Without Fake Certainty | claude-code-safe | trap | 0.0 | 3.0 | 1 weak / 1 repaired | recorded transcripts and marker outcome scoring | model-surface evidence for the named local CLI/tool configuration; reviewability scoring remains secondary |
| Incident Diagnosis Without Fake Certainty | codex-cli-default | clean | 0.0 | 1.0 | 3 weak / 3 repaired | recorded transcripts and marker outcome scoring | model-surface evidence for the named local CLI/tool configuration; reviewability scoring remains secondary |
| Incident Diagnosis Without Fake Certainty | codex-cli-default | trap | 1.0 | 5.7 | 3 weak / 3 repaired | recorded transcripts and marker outcome scoring | model-surface evidence for the named local CLI/tool configuration; reviewability scoring remains secondary |
| Incident Diagnosis Without Fake Certainty | local-deterministic-grader | n/a | execution delta pending | not applicable | fixture-local artifact execution where grader exists | repository-owned deterministic tool surface, not independent model evidence | |
| API Design Without Hidden Compatibility Traps | claude-code-safe | clean | 0.0 | 2.0 | 1 weak / 1 repaired | recorded transcripts and marker outcome scoring | model-surface evidence for the named local CLI/tool configuration; reviewability scoring remains secondary |
| API Design Without Hidden Compatibility Traps | claude-code-safe | trap | 0.0 | 3.0 | 1 weak / 1 repaired | recorded transcripts and marker outcome scoring | model-surface evidence for the named local CLI/tool configuration; reviewability scoring remains secondary |
| API Design Without Hidden Compatibility Traps | codex-cli-default | clean | 0.0 | 3.3 | 3 weak / 3 repaired | recorded transcripts and marker outcome scoring | model-surface evidence for the named local CLI/tool configuration; reviewability scoring remains secondary |
| API Design Without Hidden Compatibility Traps | codex-cli-default | trap | 0.0 | 2.7 | 3 weak / 3 repaired | recorded transcripts and marker outcome scoring | model-surface evidence for the named local CLI/tool configuration; reviewability scoring remains secondary |
| API Design Without Hidden Compatibility Traps | local-deterministic-grader | n/a | execution delta pending | not applicable | fixture-local artifact execution where grader exists | repository-owned deterministic tool surface, not independent model evidence | |
| Online Migration Without Data Loss | claude-code-safe | clean | 0.0 | 0.0 | 1 weak / 1 repaired | recorded transcripts and marker outcome scoring | model-surface evidence for the named local CLI/tool configuration; reviewability scoring remains secondary |
| Online Migration Without Data Loss | claude-code-safe | trap | 1.0 | 1.0 | 1 weak / 1 repaired | recorded transcripts and marker outcome scoring | model-surface evidence for the named local CLI/tool configuration; reviewability scoring remains secondary |
| Online Migration Without Data Loss | codex-cli-default | clean | 0.0 | 0.7 | 3 weak / 3 repaired | recorded transcripts and marker outcome scoring | model-surface evidence for the named local CLI/tool configuration; reviewability scoring remains secondary |
| Online Migration Without Data Loss | codex-cli-default | trap | 0.0 | 0.7 | 3 weak / 3 repaired | recorded transcripts and marker outcome scoring | model-surface evidence for the named local CLI/tool configuration; reviewability scoring remains secondary |
| Online Migration Without Data Loss | local-deterministic-grader | n/a | execution delta pending | not applicable | fixture-local artifact execution where grader exists | repository-owned deterministic tool surface, not independent model evidence | |
| Test Generation Without Overfitting To Implementation | claude-code-safe | clean | 0.0 | -2.0 | 1 weak / 1 repaired | recorded transcripts and marker outcome scoring | model-surface evidence for the named local CLI/tool configuration; reviewability scoring remains secondary |
| Test Generation Without Overfitting To Implementation | claude-code-safe | trap | 2.0 | 5.0 | 1 weak / 1 repaired | recorded transcripts and marker outcome scoring | model-surface evidence for the named local CLI/tool configuration; reviewability scoring remains secondary |
| Test Generation Without Overfitting To Implementation | codex-cli-default | clean | 0.3 | 2.3 | 3 weak / 3 repaired | recorded transcripts and marker outcome scoring | model-surface evidence for the named local CLI/tool configuration; reviewability scoring remains secondary |
| Test Generation Without Overfitting To Implementation | codex-cli-default | trap | 0.0 | 3.0 | 3 weak / 3 repaired | recorded transcripts and marker outcome scoring | model-surface evidence for the named local CLI/tool configuration; reviewability scoring remains secondary |
| Test Generation Without Overfitting To Implementation | local-deterministic-grader | n/a | execution delta pending | not applicable | fixture-local artifact execution where grader exists | repository-owned deterministic tool surface, not independent model evidence | |
| Performance Tuning Without Micro-Optimization Drift | claude-code-safe | clean | 0.0 | -2.0 | 1 weak / 1 repaired | recorded transcripts and marker outcome scoring | model-surface evidence for the named local CLI/tool configuration; reviewability scoring remains secondary |
| Performance Tuning Without Micro-Optimization Drift | claude-code-safe | trap | 3.0 | 10.0 | 1 weak / 1 repaired | recorded transcripts and marker outcome scoring | model-surface evidence for the named local CLI/tool configuration; reviewability scoring remains secondary |
| Performance Tuning Without Micro-Optimization Drift | codex-cli-default | clean | 0.0 | 2.3 | 3 weak / 3 repaired | recorded transcripts and marker outcome scoring | model-surface evidence for the named local CLI/tool configuration; reviewability scoring remains secondary |
| Performance Tuning Without Micro-Optimization Drift | codex-cli-default | trap | 0.3 | 5.0 | 3 weak / 3 repaired | recorded transcripts and marker outcome scoring | model-surface evidence for the named local CLI/tool configuration; reviewability scoring remains secondary |
| Performance Tuning Without Micro-Optimization Drift | local-deterministic-grader | n/a | execution delta pending | not applicable | fixture-local artifact execution where grader exists | repository-owned deterministic tool surface, not independent model evidence |
Raw comparison data: surface-comparison.json