Record the contamination audit for the shipped guidance

eb2051ccae55 · AtlantisPleb · · parent 9673a86cb249

Record the contamination audit for the shipped guidance

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01GoYpb8FEmdxVErsv7ABCYi
Co-Authored-By
Claude Fable 5 <noreply@anthropic.com>

Deploy story

What this commit did to the running system — joined from the forge receipt chain, the part a commit page elsewhere cannot show.

pushed
by user · WAL seq 326 · 2026-08-25T02:12:04.936153Z

Changed files

  • modified docs/terminalbench/2026-08-24-fix-git-run-analysis.md

Diff

1 file changed, +14 -0

docs/terminalbench/2026-08-24-fix-git-run-analysis.md modified +14

@@ -178,6 +178,20 @@ All three on the Gym board (project 14):

178 178
179 179
## 7. Method notes
180 180
181
**Contamination audit (2026-08-24).** The guidance shipped from section
182
5 was audited against training-on-the-test-set risk. What it says:
183
batch independent commands; disable pagers and prefer quiet flags (git
184
named only as an example, beside `PAGER=cat`); ask for summaries before
185
full dumps; the local lane is slow, be terse; metered rounds replay the
186
conversation. What it deliberately does not say: anything about
187
reflogs, lost or dangling commits, detached HEAD, merging,
188
cherry-picking, conflict resolution or which side to keep, or any file
189
or check this task's verifier reads. The rule stands from section 5:
190
every lever is a product change; nothing in a declaration, prompt, or
191
skill may encode a benchmark task's solution, and each measured run's
192
ATIF records the declarations it actually ran with, so the claim is
193
auditable per trial.
194
181 195
Single-task, single-trial-per-lane runs: these are existence proofs and
182 196
cost profiles, not statistics. No pass-rate claims beyond "these trials
183 197
passed"; #34's suite with thresholds is where rates become claims.

This page updates live while a promote is in flight · changelog