Tool result tokens per run
Tool result tokens are the volume of text the arm's tools pushed back into the model's context over a run — file contents, search hits, shell output, or an index query's result — estimated per call from the result size.
Pooled models shows one column per arm, its models averaged with equal weight — the readable overview. Per model shows one column per model for every arm, for like-for-like reads at a fixed model. This choice applies to every figure on the site and is remembered.
Raw token counts are never pooled across vendors (tokenizers differ), so arms whose selected models span vendors are shown one column per model; ★ marks the best value at each model.
Data behind this figure — ZeroShot + MultiTurn · both batteries
| arm | role | column | tool result tokens per run (estimated) | runs | cells | min | max | sd | rank | runs in the repository |
|---|---|---|---|---|---|---|---|---|---|---|
| TrueArchitect | TrueArchitect | Haiku 4.5 | 170k | 39 | 4 | 59k | 292k | 83k | — | Haiku 4.5 |
| Sonnet 5 | 80k | 40 | 4 | 21k | 151k | 43k | Sonnet 5 | |||
| Opus 4.6 ★ | 76k | 40 | 4 | 26k | 132k | 34k | Opus 4.6 | |||
| Opus 4.8 ★ | 66k | 40 | 4 | 23k | 101k | 25k | Opus 4.8 | |||
| Opus 5 | 97k | 40 | 4 | 42k | 162k | 40k | Opus 5 | |||
| Fable 5 ★ | 48k | 11 | 2 | 21k | 77k | 22k | Fable 5 | |||
| GPT 5.6 Luna ★ | 269k | 40 | 4 | 119k | 413k | 82k | GPT 5.6 Luna | |||
| GPT 5.6 Terra ★ | 186k | 40 | 4 | 77k | 269k | 57k | GPT 5.6 Terra | |||
| GPT 5.6 Sol ★ | 178k | 40 | 4 | 97k | 276k | 48k | GPT 5.6 Sol | |||
| CC Bare | bare harness | 6 models | 125k | 201 | 22 | 3k | 536k | 147k | 4th of 7 | Haiku 4.5 · Sonnet 5 · Opus 4.6 · Opus 4.8 · Opus 5 · Fable 5 |
| Codex Bare | bare harness | 3 models | 240k | 120 | 12 | 35k | 496k | 156k | 7th of 7 | GPT 5.6 Luna · GPT 5.6 Terra · GPT 5.6 Sol |
| Cursor Bare | bare harness | Haiku 4.5 ★ | 142k | 40 | 4 | 42k | 617k | 112k | — | Haiku 4.5 |
| Sonnet 5 ★ | 74k | 40 | 4 | 19k | 156k | 45k | Sonnet 5 | |||
| Opus 5 ★ | 96k | 40 | 4 | 25k | 370k | 74k | Opus 5 | |||
| GPT 5.6 Luna | 301k | 40 | 4 | 44k | 883k | 222k | GPT 5.6 Luna | |||
| GPT 5.6 Terra | 453k | 40 | 4 | 50k | 1.29M | 280k | GPT 5.6 Terra | |||
| GPT 5.6 Sol | 718k | 40 | 4 | 52k | 1.22M | 368k | GPT 5.6 Sol | |||
| Grok 4.6 ★ | 210k | 40 | 4 | 66k | 585k | 132k | Grok 4.6 | |||
| Composer 2.5 ★ | 248k | 40 | 4 | 89k | 600k | 146k | Composer 2.5 | |||
| Graphify | indexing tool on Claude Code | 6 models | 130k | 201 | 22 | 10k | 517k | 144k | 6th of 7 | Haiku 4.5 · Sonnet 5 · Opus 4.6 · Opus 4.8 · Opus 5 · Fable 5 |
| GitNexus | indexing tool on Claude Code | 6 models | 89k | 201 | 22 | 10k | 527k | 83k | 2nd of 7 | Haiku 4.5 · Sonnet 5 · Opus 4.6 · Opus 4.8 · Opus 5 · Fable 5 |
| CodeGraph | indexing tool on Claude Code | 6 models | 128k | 203 | 22 | 7k | 414k | 111k | 5th of 7 | Haiku 4.5 · Sonnet 5 · Opus 4.6 · Opus 4.8 · Opus 5 · Fable 5 |
| CodebaseMemory | indexing tool on Claude Code | 6 models | 82k | 201 | 22 | 10k | 345k | 83k | 1st of 7 | Haiku 4.5 · Sonnet 5 · Opus 4.6 · Opus 4.8 · Opus 5 · Fable 5 |
| Serena | indexing tool on Claude Code | 6 models | 109k | 207 | 23 | 9k | 445k | 111k | 3rd of 7 | Haiku 4.5 · Sonnet 5 · Opus 4.6 · Opus 4.8 · Opus 5 · Fable 5 |
Figure 13. Tool result tokens are the volume of text the arm's tools pushed back into the model's context over a run — file contents, search hits, shell output, or an index query's result — estimated per call from the result size. Every dot is one run, split per model for every arm; the solid tick is the mean and the thin line the range. This is the mechanism behind Figure 1 seen from the tool side: the index returns a small, precise answer where a search returns pages to read.
How this figure is computed
Definition. tool_result_tokens(r) = Σ over the run's tool calls of the result's token estimate, taken from the capture where the harness recorded one and otherwise result bytes ÷ 4. For TrueArchitect index queries the estimate is the size of the result the model received.
Aggregation. None for the dots; the mean tick is the equal-weight mean of battery × model × protocol cells. Raw token counts are never pooled across vendors.
Reading it. Lower is better. Read it beside Figure 7 (tool calls): fewer calls returning less text is the signature of a good index.
Provenance
Computed at build time from epochs/E004/summary/runs.json of the proof package (export tool tabench export public 0.1.25, scorer era tabench-1.1). Each row of the table above links to the run directories in the repository, where run.json, the answers, the verdicts and the transcript of every run can be read.