Context tokens per correct answer
This divides a run's context tokens (Figure 1) by the number of questions it answered correctly, so an arm that reads less but also answers less is not rewarded.
Pooled models shows one column per arm, its models averaged with equal weight — the readable overview. Per model shows one column per model for every arm, for like-for-like reads at a fixed model. This choice applies to every figure on the site and is remembered.
Raw token counts are never pooled across vendors (tokenizers differ), so arms whose selected models span vendors are shown one column per model; ★ marks the best value at each model.
Data behind this figure — ZeroShot + MultiTurn · both batteries
| arm | role | column | context tokens per correct answer | runs | cells | min | max | sd | rank | runs in the repository |
|---|---|---|---|---|---|---|---|---|---|---|
| TrueArchitect | TrueArchitect | Haiku 4.5 ★ | 448k | 39 | 4 | 84k | 1.35M | 378k | — | Haiku 4.5 |
| Sonnet 5 ★ | 156k | 40 | 4 | 51k | 399k | 109k | Sonnet 5 | |||
| Opus 4.6 ★ | 126k | 40 | 4 | 27k | 434k | 115k | Opus 4.6 | |||
| Opus 4.8 ★ | 134k | 40 | 4 | 29k | 425k | 122k | Opus 4.8 | |||
| Opus 5 ★ | 140k | 40 | 4 | 46k | 311k | 78k | Opus 5 | |||
| Fable 5 ★ | 63k | 11 | 2 | 26k | 106k | 34k | Fable 5 | |||
| GPT 5.6 Luna ★ | 232k | 40 | 4 | 29k | 890k | 254k | GPT 5.6 Luna | |||
| GPT 5.6 Terra ★ | 148k | 40 | 4 | 22k | 509k | 154k | GPT 5.6 Terra | |||
| GPT 5.6 Sol | 131k | 40 | 4 | 22k | 455k | 129k | GPT 5.6 Sol | |||
| CC Bare | bare harness | 6 models | 454k | 201 | 22 | 112k | 2.17M | 382k | 3rd of 7 | Haiku 4.5 · Sonnet 5 · Opus 4.6 · Opus 4.8 · Opus 5 · Fable 5 |
| Codex Bare | bare harness | 3 models | 245k | 120 | 12 | 72k | 901k | 165k | 1st of 7 | GPT 5.6 Luna · GPT 5.6 Terra · GPT 5.6 Sol |
| Cursor Bare | bare harness | Haiku 4.5 | 712k | 40 | 4 | 80k | 1.49M | 396k | — | Haiku 4.5 |
| Sonnet 5 | 439k | 40 | 4 | 172k | 988k | 243k | Sonnet 5 | |||
| Opus 5 | 366k | 40 | 4 | 219k | 865k | 158k | Opus 5 | |||
| GPT 5.6 Luna | 331k | 40 | 4 | 157k | 823k | 161k | GPT 5.6 Luna | |||
| GPT 5.6 Terra | 167k | 40 | 4 | 70k | 438k | 94k | GPT 5.6 Terra | |||
| GPT 5.6 Sol ★ | 113k | 40 | 4 | 66k | 369k | 53k | GPT 5.6 Sol | |||
| Grok 4.6 ★ | 312k | 40 | 4 | 143k | 594k | 140k | Grok 4.6 | |||
| Composer 2.5 ★ | 432k | 40 | 4 | 179k | 843k | 205k | Composer 2.5 | |||
| Graphify | indexing tool on Claude Code | 6 models | 467k | 201 | 22 | 119k | 2.04M | 380k | 6th of 7 | Haiku 4.5 · Sonnet 5 · Opus 4.6 · Opus 4.8 · Opus 5 · Fable 5 |
| GitNexus | indexing tool on Claude Code | 6 models | 465k | 201 | 22 | 163k | 1.95M | 344k | 5th of 7 | Haiku 4.5 · Sonnet 5 · Opus 4.6 · Opus 4.8 · Opus 5 · Fable 5 |
| CodeGraph | indexing tool on Claude Code | 6 models | 359k | 203 | 22 | 97k | 1.97M | 288k | 2nd of 7 | Haiku 4.5 · Sonnet 5 · Opus 4.6 · Opus 4.8 · Opus 5 · Fable 5 |
| CodebaseMemory | indexing tool on Claude Code | 6 models | 457k | 201 | 22 | 136k | 1.76M | 330k | 4th of 7 | Haiku 4.5 · Sonnet 5 · Opus 4.6 · Opus 4.8 · Opus 5 · Fable 5 |
| Serena | indexing tool on Claude Code | 6 models | 502k | 207 | 23 | 138k | 2.09M | 417k | 7th of 7 | Haiku 4.5 · Sonnet 5 · Opus 4.6 · Opus 4.8 · Opus 5 · Fable 5 |
Figure 6. This divides a run's context tokens (Figure 1) by the number of questions it answered correctly, so an arm that reads less but also answers less is not rewarded. Each column is the mean over valid, scored runs at one model (or pooled over a single-vendor roster in the pooled-models view), lower is better, and TrueArchitect and the bare harnesses are the dashed references. The ordering of Figure 1 survives the normalisation: the token saving is a saving per correct answer, not a saving bought with wrong answers.
How this figure is computed
Definition. For a scored run with score(r) > 0, ctx_per_correct(r) = context(r) / score(r). Runs with no correct answer have no defined value and are excluded from the mean; their count is in the table.
Aggregation. As in Figure 1: cell means averaged with equal weight; raw token counts are never pooled across vendors.
Reading it. Lower is better. Cost in currency is deliberately not shown here: the package publishes cost only where a vendor reported it, and TrueArchitect's runs carry no vendor-reported figure. Tokens are the quantity a reader can recompute from the run records.
Provenance
Computed at build time from epochs/E004/summary/runs.json of the proof package (export tool tabench export public 0.1.25, scorer era tabench-1.1). Each row of the table above links to the run directories in the repository, where run.json, the answers, the verdicts and the transcript of every run can be read.