Figure 12 · E004

Cost per run

Every dot is what one run cost (vendor-reported where available, otherwise estimated from the rate tables), so the spread of an arm's price is visible, not just its mean.

protocols
battery
model
columns

Pooled models shows one column per arm, its models averaged with equal weight — the readable overview. Per model shows one column per model for every arm, for like-for-like reads at a fixed model. This choice applies to every figure on the site and is remembered.

Cost per runlower is betterZeroShot + MultiTurn · both batteries
$0.00$5.0$10$15$20$25TrueArchitect $2.51CC Bare $6.18Codex Bare $1.80Cursor Bare $3.66$2.51$0.097$7.27TrueArchitectTrueArchitect2nd of 9$6.18$2.14$15.0CC Barebare harness6th of 9$1.80$0.110$4.08Codex Barebare harness1st of 9$3.66$0.141$12.9Cursor Barebare harness3rd of 9$7.14$1.94$24.1Graphifyindexing tool9th of 9$6.49$2.06$13.6GitNexusindexing tool7th of 9$5.81$1.16$14.4CodeGraphindexing tool5th of 9$5.70$2.10$12.0CodebaseMemoryindexing tool4th of 9$6.80$2.12$13.9Serenaindexing tool8th of 9
Data behind this figure ZeroShot + MultiTurn · both batteries
armrolecolumnUSD per runrunscellsminmaxsdrankruns in the repository
TrueArchitectTrueArchitect9 models$2.5133034$0.097$7.27$1.612nd of 9Haiku 4.5 · Sonnet 5 · Opus 4.6 · Opus 4.8 · Opus 5 · Fable 5 · GPT 5.6 Luna · GPT 5.6 Terra · GPT 5.6 Sol
CC Barebare harness6 models$6.1820122$2.14$15.0$2.666th of 9Haiku 4.5 · Sonnet 5 · Opus 4.6 · Opus 4.8 · Opus 5 · Fable 5
Codex Barebare harness3 models$1.8012012$0.110$4.08$1.281st of 9GPT 5.6 Luna · GPT 5.6 Terra · GPT 5.6 Sol
Cursor Barebare harness8 models$3.6632032$0.141$12.9$3.053rd of 9Haiku 4.5 · Sonnet 5 · Opus 5 · GPT 5.6 Luna · GPT 5.6 Terra · GPT 5.6 Sol · Grok 4.6 · Composer 2.5
Graphifyindexing tool on Claude Code6 models$7.1420122$1.94$24.1$3.899th of 9Haiku 4.5 · Sonnet 5 · Opus 4.6 · Opus 4.8 · Opus 5 · Fable 5
GitNexusindexing tool on Claude Code6 models$6.4920122$2.06$13.6$2.497th of 9Haiku 4.5 · Sonnet 5 · Opus 4.6 · Opus 4.8 · Opus 5 · Fable 5
CodeGraphindexing tool on Claude Code6 models$5.8120322$1.16$14.4$2.495th of 9Haiku 4.5 · Sonnet 5 · Opus 4.6 · Opus 4.8 · Opus 5 · Fable 5
CodebaseMemoryindexing tool on Claude Code6 models$5.7020122$2.10$12.0$2.024th of 9Haiku 4.5 · Sonnet 5 · Opus 4.6 · Opus 4.8 · Opus 5 · Fable 5
Serenaindexing tool on Claude Code6 models$6.8020723$2.12$13.9$2.828th of 9Haiku 4.5 · Sonnet 5 · Opus 4.6 · Opus 4.8 · Opus 5 · Fable 5

Figure 12. Every dot is what one run cost (vendor-reported where available, otherwise estimated from the rate tables), so the spread of an arm's price is visible, not just its mean. The solid tick is the pooled mean and the thin line the range; hover a dot for its model, repetition and run directory. A tight, low cloud is an arm whose price is predictable; a tall cloud is one whose price depends on the run.

How this figure is computed

Definition. cost(r) as in Figure 11, drawn per run rather than averaged.

Aggregation. None for the dots; the mean tick is the equal-weight mean of battery × model × protocol cells.

Reading it. Lower is better; read spread as well as position.

Provenance

Computed at build time from epochs/E004/summary/runs.json of the proof package (export tool tabench export public 0.1.25, scorer era tabench-1.1). Each row of the table above links to the run directories in the repository, where run.json, the answers, the verdicts and the transcript of every run can be read.