Outcome composition per run
Every valid run becomes one thin column showing how its questions divided between correct (the arm's color), incorrect (dark) and unanswered or unjudged (black), so an arm's whole record is visible at once.
Data behind this figure — ZeroShot + MultiTurn · both batteries
| arm | role | column | questions correct · incorrect · unanswered, per run | runs | cells | min | max | sd | rank | runs in the repository |
|---|---|---|---|---|---|---|---|---|---|---|
| TrueArchitect | TrueArchitect | 9 models | 93.3% | 330 | 34 | 65.0% | 100.0% | 8.2% | 1st of 9 | Haiku 4.5 · Sonnet 5 · Opus 4.6 · Opus 4.8 · Opus 5 · Fable 5 · GPT 5.6 Luna · GPT 5.6 Terra · GPT 5.6 Sol |
| CC Bare | bare harness | 6 models | 89.4% | 201 | 22 | 40.0% | 100.0% | 11.5% | 7th of 9 | Haiku 4.5 · Sonnet 5 · Opus 4.6 · Opus 4.8 · Opus 5 · Fable 5 |
| Codex Bare | bare harness | 3 models | 85.9% | 120 | 12 | 55.0% | 100.0% | 13.6% | 9th of 9 | GPT 5.6 Luna · GPT 5.6 Terra · GPT 5.6 Sol |
| Cursor Bare | bare harness | 8 models | 89.6% | 320 | 32 | 55.0% | 100.0% | 11.4% | 6th of 9 | Haiku 4.5 · Sonnet 5 · Opus 5 · GPT 5.6 Luna · GPT 5.6 Terra · GPT 5.6 Sol · Grok 4.6 · Composer 2.5 |
| Graphify | indexing tool on Claude Code | 6 models | 90.0% | 201 | 22 | 40.0% | 100.0% | 11.6% | 5th of 9 | Haiku 4.5 · Sonnet 5 · Opus 4.6 · Opus 4.8 · Opus 5 · Fable 5 |
| GitNexus | indexing tool on Claude Code | 6 models | 90.6% | 201 | 22 | 55.0% | 100.0% | 9.2% | 2nd of 9 | Haiku 4.5 · Sonnet 5 · Opus 4.6 · Opus 4.8 · Opus 5 · Fable 5 |
| CodeGraph | indexing tool on Claude Code | 6 models | 88.4% | 203 | 22 | 40.0% | 100.0% | 12.9% | 8th of 9 | Haiku 4.5 · Sonnet 5 · Opus 4.6 · Opus 4.8 · Opus 5 · Fable 5 |
| CodebaseMemory | indexing tool on Claude Code | 6 models | 90.4% | 201 | 22 | 60.0% | 100.0% | 9.7% | 3rd of 9 | Haiku 4.5 · Sonnet 5 · Opus 4.6 · Opus 4.8 · Opus 5 · Fable 5 |
| Serena | indexing tool on Claude Code | 6 models | 90.1% | 207 | 23 | 50.0% | 100.0% | 10.9% | 4th of 9 | Haiku 4.5 · Sonnet 5 · Opus 4.6 · Opus 4.8 · Opus 5 · Fable 5 |
Figure 5. Every valid run becomes one thin column showing how its questions divided between correct (the arm's color), incorrect (dark) and unanswered or unjudged (black), so an arm's whole record is visible at once. Columns are ordered by model then repetition; the number above each group is the arm's mean accuracy and the number inside is its difference from TrueArchitect. The dark bands are the honesty layer: the failures are on the page, not in a footnote.
How this figure is computed
Definition. For a scored run, correct = score, incorrect = the count of effective verdicts of fail, and the residue denominator − correct − incorrect is the unanswered or unjudged portion (an adjudicate or infra verdict, or a question with no verdict). Each is drawn as a share of the denominator.
Aggregation. None: every scored run is one column. The header figure is the hierarchical mean accuracy of the group and the inner figure is that mean minus TrueArchitect's.
Reading it. Higher is better for the colored band. The width of a group is proportional to the number of runs the arm made in the selection, which is itself information: an arm with fewer columns has fewer runs behind its numbers.
Disclosures
Did-not-finish runs have no score and do not appear here; their count per arm is in the run records and in the Limitations page.
Provenance
Computed at build time from epochs/E004/summary/runs.json of the proof package (export tool tabench export public 0.1.25, scorer era tabench-1.1). Each row of the table above links to the run directories in the repository, where run.json, the answers, the verdicts and the transcript of every run can be read.