Accuracy by model
The same accuracy as Figure 2, split into one column per model for every arm, so the cross-tier comparison is readable directly.
This figure is always one column per model; the site-wide columns choice does not apply to it.
Data behind this figure — ZeroShot + MultiTurn · both batteries
| arm | role | column | accuracy (% of questions correct) | runs | cells | min | max | sd | rank | runs in the repository |
|---|---|---|---|---|---|---|---|---|---|---|
| TrueArchitect | TrueArchitect | Haiku 4.5 | 89.9% | 39 | 4 | 65.0% | 100.0% | 8.6% | 1st of 9 | Haiku 4.5 |
| Sonnet 5 | 96.9% | 40 | 4 | 80.0% | 100.0% | 4.8% | Sonnet 5 | |||
| Opus 4.6 | 94.8% | 40 | 4 | 80.0% | 100.0% | 5.8% | Opus 4.6 | |||
| Opus 4.8 | 95.1% | 40 | 4 | 80.0% | 100.0% | 5.5% | Opus 4.8 | |||
| Opus 5 | 98.8% | 40 | 4 | 95.0% | 100.0% | 2.1% | Opus 5 | |||
| Fable 5 | 100.0% | 11 | 2 | 100.0% | 100.0% | 0.0% | Fable 5 | |||
| GPT 5.6 Luna | 87.7% | 40 | 4 | 70.0% | 100.0% | 9.6% | GPT 5.6 Luna | |||
| GPT 5.6 Terra | 89.7% | 40 | 4 | 70.0% | 100.0% | 9.4% | GPT 5.6 Terra | |||
| GPT 5.6 Sol | 90.3% | 40 | 4 | 70.0% | 100.0% | 9.8% | GPT 5.6 Sol | |||
| CC Bare | bare harness | Haiku 4.5 | 79.1% | 36 | 4 | 40.0% | 96.7% | 16.2% | 7th of 9 | Haiku 4.5 |
| Sonnet 5 | 87.0% | 40 | 4 | 65.0% | 100.0% | 10.1% | Sonnet 5 | |||
| Opus 4.6 | 89.4% | 40 | 4 | 75.0% | 100.0% | 7.2% | Opus 4.6 | |||
| Opus 4.8 | 89.3% | 30 | 3 | 65.0% | 100.0% | 10.5% | Opus 4.8 | |||
| Opus 5 | 96.0% | 40 | 4 | 85.0% | 100.0% | 3.8% | Opus 5 | |||
| Fable 5 | 97.8% | 15 | 3 | 90.0% | 100.0% | 3.1% | Fable 5 | |||
| Codex Bare | bare harness | GPT 5.6 Luna | 85.2% | 40 | 4 | 55.0% | 100.0% | 14.1% | 9th of 9 | GPT 5.6 Luna |
| GPT 5.6 Terra | 87.5% | 40 | 4 | 55.0% | 100.0% | 12.2% | GPT 5.6 Terra | |||
| GPT 5.6 Sol | 85.0% | 40 | 4 | 60.0% | 100.0% | 14.4% | GPT 5.6 Sol | |||
| Cursor Bare | bare harness | Haiku 4.5 | 87.9% | 40 | 4 | 65.0% | 100.0% | 9.8% | 6th of 9 | Haiku 4.5 |
| Sonnet 5 | 94.1% | 40 | 4 | 80.0% | 100.0% | 6.7% | Sonnet 5 | |||
| Opus 5 | 97.2% | 40 | 4 | 90.0% | 100.0% | 3.7% | Opus 5 | |||
| GPT 5.6 Luna | 83.7% | 40 | 4 | 60.0% | 100.0% | 13.7% | GPT 5.6 Luna | |||
| GPT 5.6 Terra | 84.3% | 40 | 4 | 55.0% | 100.0% | 14.5% | GPT 5.6 Terra | |||
| GPT 5.6 Sol | 85.3% | 40 | 4 | 60.0% | 100.0% | 13.6% | GPT 5.6 Sol | |||
| Grok 4.6 | 91.8% | 40 | 4 | 65.0% | 100.0% | 9.3% | Grok 4.6 | |||
| Composer 2.5 | 92.7% | 40 | 4 | 70.0% | 100.0% | 7.9% | Composer 2.5 | |||
| Graphify | indexing tool on Claude Code | Haiku 4.5 | 78.1% | 36 | 4 | 40.0% | 100.0% | 17.4% | 5th of 9 | Haiku 4.5 |
| Sonnet 5 | 91.1% | 40 | 4 | 75.0% | 100.0% | 8.1% | Sonnet 5 | |||
| Opus 4.6 | 89.1% | 40 | 4 | 75.0% | 100.0% | 7.1% | Opus 4.6 | |||
| Opus 4.8 | 89.3% | 30 | 3 | 70.0% | 100.0% | 9.6% | Opus 4.8 | |||
| Opus 5 | 97.0% | 40 | 4 | 90.0% | 100.0% | 3.3% | Opus 5 | |||
| Fable 5 | 96.8% | 15 | 3 | 85.0% | 100.0% | 4.9% | Fable 5 | |||
| GitNexus | indexing tool on Claude Code | Haiku 4.5 | 83.4% | 36 | 4 | 55.0% | 100.0% | 11.5% | 2nd of 9 | Haiku 4.5 |
| Sonnet 5 | 90.3% | 40 | 4 | 75.0% | 100.0% | 7.4% | Sonnet 5 | |||
| Opus 4.6 | 88.4% | 40 | 4 | 65.0% | 100.0% | 9.2% | Opus 4.6 | |||
| Opus 4.8 | 90.3% | 30 | 3 | 75.0% | 100.0% | 7.9% | Opus 4.8 | |||
| Opus 5 | 96.5% | 40 | 4 | 85.0% | 100.0% | 4.2% | Opus 5 | |||
| Fable 5 | 95.9% | 15 | 3 | 85.0% | 100.0% | 5.3% | Fable 5 | |||
| CodeGraph | indexing tool on Claude Code | Haiku 4.5 | 74.5% | 38 | 4 | 40.0% | 96.7% | 17.9% | 8th of 9 | Haiku 4.5 |
| Sonnet 5 | 90.4% | 40 | 4 | 70.0% | 100.0% | 10.0% | Sonnet 5 | |||
| Opus 4.6 | 85.9% | 40 | 4 | 65.0% | 96.7% | 7.9% | Opus 4.6 | |||
| Opus 4.8 | 88.3% | 30 | 3 | 70.0% | 100.0% | 10.4% | Opus 4.8 | |||
| Opus 5 | 96.1% | 40 | 4 | 85.0% | 100.0% | 4.3% | Opus 5 | |||
| Fable 5 | 97.0% | 15 | 3 | 90.0% | 100.0% | 4.6% | Fable 5 | |||
| CodebaseMemory | indexing tool on Claude Code | Haiku 4.5 | 81.0% | 36 | 4 | 60.0% | 100.0% | 10.9% | 3rd of 9 | Haiku 4.5 |
| Sonnet 5 | 92.7% | 40 | 4 | 75.0% | 100.0% | 8.1% | Sonnet 5 | |||
| Opus 4.6 | 88.2% | 40 | 4 | 70.0% | 100.0% | 9.0% | Opus 4.6 | |||
| Opus 4.8 | 89.5% | 30 | 3 | 70.0% | 100.0% | 8.8% | Opus 4.8 | |||
| Opus 5 | 96.0% | 40 | 4 | 85.0% | 100.0% | 4.4% | Opus 5 | |||
| Fable 5 | 96.4% | 15 | 3 | 85.0% | 100.0% | 5.2% | Fable 5 | |||
| Serena | indexing tool on Claude Code | Haiku 4.5 | 80.8% | 38 | 4 | 50.0% | 100.0% | 14.2% | 4th of 9 | Haiku 4.5 |
| Sonnet 5 | 88.5% | 40 | 4 | 60.0% | 100.0% | 10.9% | Sonnet 5 | |||
| Opus 4.6 | 90.0% | 40 | 4 | 75.0% | 100.0% | 7.1% | Opus 4.6 | |||
| Opus 4.8 | 88.5% | 30 | 3 | 65.0% | 100.0% | 10.8% | Opus 4.8 | |||
| Opus 5 | 96.5% | 40 | 4 | 85.0% | 100.0% | 4.2% | Opus 5 | |||
| Fable 5 | 96.1% | 19 | 4 | 85.0% | 100.0% | 4.3% | Fable 5 |
Figure 9. The same accuracy as Figure 2, split into one column per model for every arm, so the cross-tier comparison is readable directly. Within an arm the columns run from the fastest, cheapest model to the frontier model of each vendor. The comparison to make is diagonal: TrueArchitect at a fast or everyday model against a bare harness or an indexing tool at a frontier model.
How this figure is computed
Definition and aggregation as in Figure 2, with cells restricted to one model per column.
Reading it. Higher is better. The vendor tiers are declared, not inferred: fast = Haiku 4.5 and GPT 5.6 Luna; everyday = Sonnet 5 and GPT 5.6 Terra; frontier = Opus 5, GPT 5.6 Sol and Composer 2.5.
Provenance
Computed at build time from epochs/E004/summary/runs.json of the proof package (export tool tabench export public 0.1.25, scorer era tabench-1.1). Each row of the table above links to the run directories in the repository, where run.json, the answers, the verdicts and the transcript of every run can be read.