Effect size against bare Claude Code
Cohen's d expresses how far an arm's accuracy distribution sits from bare Claude Code's in units of their pooled standard deviation, so it is comparable across batteries and models.
Pooled models shows one column per arm, its models averaged with equal weight — the readable overview. Per model shows one column per model for every arm, for like-for-like reads at a fixed model. This choice applies to every figure on the site and is remembered.
Data behind this figure — ZeroShot + MultiTurn · both batteries
| arm | role | column | Cohen's d vs bare Claude Code, accuracy | runs | cells | min | max | sd | rank | runs in the repository |
|---|---|---|---|---|---|---|---|---|---|---|
| TrueArchitect | TrueArchitect | 9 models | 1.35 | 200 | 21 | 65.00 | 100.00 | 8.19 | 1st of 8 | Haiku 4.5 · Sonnet 5 · Opus 4.6 · Opus 4.8 · Opus 5 · Fable 5 · GPT 5.6 Luna · GPT 5.6 Terra · GPT 5.6 Sol |
| CC Bare | bare harness | 6 models | 0.00 | 201 | 22 | 40.00 | 100.00 | 11.53 | 7th of 8 | Haiku 4.5 · Sonnet 5 · Opus 4.6 · Opus 4.8 · Opus 5 · Fable 5 |
| Cursor Bare | bare harness | 8 models | 1.28 | 120 | 12 | 55.00 | 100.00 | 11.42 | 2nd of 8 | Haiku 4.5 · Sonnet 5 · Opus 5 · GPT 5.6 Luna · GPT 5.6 Terra · GPT 5.6 Sol · Grok 4.6 · Composer 2.5 |
| Graphify | indexing tool on Claude Code | 6 models | 0.17 | 201 | 22 | 40.00 | 100.00 | 11.63 | 5th of 8 | Haiku 4.5 · Sonnet 5 · Opus 4.6 · Opus 4.8 · Opus 5 · Fable 5 |
| GitNexus | indexing tool on Claude Code | 6 models | 0.14 | 201 | 22 | 55.00 | 100.00 | 9.21 | 6th of 8 | Haiku 4.5 · Sonnet 5 · Opus 4.6 · Opus 4.8 · Opus 5 · Fable 5 |
| CodeGraph | indexing tool on Claude Code | 6 models | -0.07 | 203 | 22 | 40.00 | 100.00 | 12.94 | 8th of 8 | Haiku 4.5 · Sonnet 5 · Opus 4.6 · Opus 4.8 · Opus 5 · Fable 5 |
| CodebaseMemory | indexing tool on Claude Code | 6 models | 0.24 | 201 | 22 | 60.00 | 100.00 | 9.67 | 3rd of 8 | Haiku 4.5 · Sonnet 5 · Opus 4.6 · Opus 4.8 · Opus 5 · Fable 5 |
| Serena | indexing tool on Claude Code | 6 models | 0.20 | 203 | 22 | 50.00 | 100.00 | 10.92 | 4th of 8 | Haiku 4.5 · Sonnet 5 · Opus 4.6 · Opus 4.8 · Opus 5 · Fable 5 |
Figure 4. Cohen's d expresses how far an arm's accuracy distribution sits from bare Claude Code's in units of their pooled standard deviation, so it is comparable across batteries and models. Each column is the mean d over cells that contain both the arm and bare Claude Code at the same model, battery and protocol; bare Claude Code is the zero line by definition, and arms with no shared Anthropic cell have no value. By convention d of 0.2 is small, 0.5 medium and 0.8 large; TrueArchitect is the only arm above the large threshold.
How this figure is computed
Definition. Within a cell shared by the arm (n₁ runs, accuracies with mean m₁ and standard deviation s₁) and bare Claude Code (n₂, m₂, s₂), d = (m₁ − m₂) / s_p with s_p = √(((n₁−1)s₁² + (n₂−1)s₂²) / (n₁+n₂−2)). Standard deviations use the n−1 denominator.
Aggregation. The column is the equal-weight mean of d over the cells the arm shares with bare Claude Code. A cell with a single run on either side contributes a degenerate deviation and is reported as such in the table. Codex and Cursor at non-Anthropic models have no bare Claude Code counterpart and are therefore absent, not zero.
Reading it. Higher is better; negative values would mean the arm is worse than bare Claude Code. The figure answers "is the difference large relative to run-to-run noise", the question a mean alone cannot answer.
Provenance
Computed at build time from epochs/E004/summary/runs.json of the proof package (export tool tabench export public 0.1.25, scorer era tabench-1.1). Each row of the table above links to the run directories in the repository, where run.json, the answers, the verdicts and the transcript of every run can be read.