Context tokens per run
Context tokens are the tokens the model read to answer a whole run: the uncached input plus every cache read, summed over every API call in the run, all threads included.
Pooled models shows one column per arm, its models averaged with equal weight — the readable overview. Per model shows one column per model for every arm, for like-for-like reads at a fixed model. This choice applies to every figure on the site and is remembered.
Raw token counts are never pooled across vendors (tokenizers differ), so arms whose selected models span vendors are shown one column per model; ★ marks the best value at each model.
Data behind this figure — ZeroShot + MultiTurn · both batteries
| arm | role | column | context tokens per run (input + cache read) | runs | cells | min | max | sd | rank | runs in the repository |
|---|---|---|---|---|---|---|---|---|---|---|
| TrueArchitect | TrueArchitect | Haiku 4.5 ★ | 8.75M | 39 | 4 | 2.33M | 20.85M | 6.00M | — | Haiku 4.5 |
| Sonnet 5 ★ | 3.37M | 40 | 4 | 1.54M | 7.77M | 1.87M | Sonnet 5 | |||
| Opus 4.6 ★ | 2.63M | 40 | 4 | 786k | 7.38M | 1.96M | Opus 4.6 | |||
| Opus 4.8 ★ | 2.86M | 40 | 4 | 884k | 7.66M | 2.13M | Opus 4.8 | |||
| Opus 5 ★ | 3.24M | 40 | 4 | 1.34M | 5.90M | 1.52M | Opus 5 | |||
| Fable 5 ★ | 1.90M | 11 | 2 | 793k | 3.17M | 1.01M | Fable 5 | |||
| GPT 5.6 Luna ★ | 4.35M | 40 | 4 | 774k | 14.24M | 3.95M | GPT 5.6 Luna | |||
| GPT 5.6 Terra ★ | 2.85M | 40 | 4 | 603k | 8.14M | 2.44M | GPT 5.6 Terra | |||
| GPT 5.6 Sol | 2.58M | 40 | 4 | 664k | 7.91M | 2.10M | GPT 5.6 Sol | |||
| CC Bare | bare harness | 6 models | 8.59M | 201 | 22 | 3.36M | 29.61M | 4.69M | 3rd of 7 | Haiku 4.5 · Sonnet 5 · Opus 4.6 · Opus 4.8 · Opus 5 · Fable 5 |
| Codex Bare | bare harness | 3 models | 4.55M | 120 | 12 | 2.15M | 10.61M | 2.03M | 1st of 7 | GPT 5.6 Luna · GPT 5.6 Terra · GPT 5.6 Sol |
| Cursor Bare | bare harness | Haiku 4.5 | 13.83M | 40 | 4 | 2.32M | 23.84M | 5.11M | — | Haiku 4.5 |
| Sonnet 5 | 9.22M | 40 | 4 | 5.15M | 16.79M | 3.26M | Sonnet 5 | |||
| Opus 5 | 8.27M | 40 | 4 | 6.57M | 16.43M | 2.11M | Opus 5 | |||
| GPT 5.6 Luna | 6.18M | 40 | 4 | 4.18M | 10.70M | 1.46M | GPT 5.6 Luna | |||
| GPT 5.6 Terra | 3.30M | 40 | 4 | 1.23M | 5.89M | 1.26M | GPT 5.6 Terra | |||
| GPT 5.6 Sol ★ | 2.47M | 40 | 4 | 985k | 5.16M | 1.28M | GPT 5.6 Sol | |||
| Grok 4.6 ★ | 6.58M | 40 | 4 | 4.29M | 9.61M | 1.85M | Grok 4.6 | |||
| Composer 2.5 ★ | 9.38M | 40 | 4 | 5.02M | 14.81M | 3.53M | Composer 2.5 | |||
| Graphify | indexing tool on Claude Code | 6 models | 8.96M | 201 | 22 | 3.58M | 27.70M | 4.57M | 4th of 7 | Haiku 4.5 · Sonnet 5 · Opus 4.6 · Opus 4.8 · Opus 5 · Fable 5 |
| GitNexus | indexing tool on Claude Code | 6 models | 9.28M | 201 | 22 | 4.74M | 31.16M | 4.74M | 6th of 7 | Haiku 4.5 · Sonnet 5 · Opus 4.6 · Opus 4.8 · Opus 5 · Fable 5 |
| CodeGraph | indexing tool on Claude Code | 6 models | 6.92M | 203 | 22 | 2.75M | 29.62M | 4.32M | 2nd of 7 | Haiku 4.5 · Sonnet 5 · Opus 4.6 · Opus 4.8 · Opus 5 · Fable 5 |
| CodebaseMemory | indexing tool on Claude Code | 6 models | 9.11M | 201 | 22 | 4.07M | 25.47M | 4.84M | 5th of 7 | Haiku 4.5 · Sonnet 5 · Opus 4.6 · Opus 4.8 · Opus 5 · Fable 5 |
| Serena | indexing tool on Claude Code | 6 models | 9.64M | 207 | 23 | 4.13M | 30.98M | 5.29M | 7th of 7 | Haiku 4.5 · Sonnet 5 · Opus 4.6 · Opus 4.8 · Opus 5 · Fable 5 |
Figure 1. Context tokens are the tokens the model read to answer a whole run: the uncached input plus every cache read, summed over every API call in the run, all threads included. Each column is the mean over the valid, scored runs of one arm at one model, every arm split by model because token counts are model facts; the dashed lines mark TrueArchitect's and each bare harness's pooled value wherever pooling is licensed. At every model the TrueArchitect column sits well below the same model running bare, and the indexing tools cluster at or above bare Claude Code: the index replaces reading.
How this figure is computed
Definition. For run r, context(r) = tokens_input(r) + tokens_cache_read(r), where tokens_input is the uncached tail of every request and tokens_cache_read the cached prefix re-read on each call. Both are summed over all API calls of the run, including sub-agent threads, from the per-call record named by token_source in the run record (a gateway ledger for TrueArchitect; the harness session logs or the harness's own per-call usage for the comparison group). Cache writes and output tokens are excluded here and reported in the run records.
Aggregation. A column is the arithmetic mean of context(r) over the valid, scored runs of one arm at one model, for the selected exam and battery. When both batteries are selected, each battery forms its own cell and the column is the mean of the two cell means, so the longer battery does not dominate.
Why per model. Token counts are tokenizer facts: a Sonnet token and a GPT token are not the same unit. The site therefore never pools raw token counts across vendors. The pooled-models view (the default) pools an arm over its roster only where the roster is a single vendor; an arm spanning two vendors stays split into one column per model whatever the columns choice, which the table discloses.
Reading it. Lower is better. The like-for-like comparison is column to column at the same model: TrueArchitect at Sonnet 5 against bare Claude Code at Sonnet 5, TrueArchitect at GPT 5.6 Sol against Codex at GPT 5.6 Sol. The dashed reference lines are TrueArchitect's and the bare harnesses' pooled means, drawn only where a pooled value is licensed (a single-vendor roster), so the indexing tools have both TrueArchitect and their baseline in the same picture.
Disclosures
Exam protocols have different token denominators (a MultiTurn question replays its history; ZeroShot pays the fixed prefix per question). Switch protocols with the control above; the site never pools across them.
A run whose token facts were not captured is excluded from its column, never counted as zero; the table's n is the number of runs behind each column.
Provenance
Computed at build time from epochs/E004/summary/runs.json of the proof package (export tool tabench export public 0.1.25, scorer era tabench-1.1). Each row of the table above links to the run directories in the repository, where run.json, the answers, the verdicts and the transcript of every run can be read.