Figure 11 · E004

Cost per correct answer

Cost per correct answer divides what a run cost by the number of questions it answered correctly: the vendor-reported figure where the harness reported one, otherwise the benchmark's estimate from the published rate tables over the run's four token columns.

protocols
battery
model
columns

Pooled models shows one column per arm, its models averaged with equal weight — the readable overview. Per model shows one column per model for every arm, for like-for-like reads at a fixed model. This choice applies to every figure on the site and is remembered.

Cost per correct answerlower is betterZeroShot + MultiTurn · both batteries
$0.00$0.10$0.20$0.30TrueArchitect $0.115CC Bare $0.310Codex Bare $0.093Cursor Bare $0.172$0.1159 modelsTrueArchitectTrueArchitect2nd of 9$0.3106 modelsCC Barebare harness6th of 9$0.0933 modelsCodex Barebare harness1st of 9$0.1728 modelsCursor Barebare harness3rd of 9$0.3426 modelsGraphifyindexing tool9th of 9$0.3116 modelsGitNexusindexing tool7th of 9$0.2836 modelsCodeGraphindexing tool5th of 9$0.2746 modelsCodebaseMemoryindexing tool4th of 9$0.3386 modelsSerenaindexing tool8th of 9
Data behind this figure ZeroShot + MultiTurn · both batteries
armrolecolumnUSD per correct answerrunscellsminmaxsdrankruns in the repository
TrueArchitectTrueArchitect9 models$0.11533034$0.004$0.424$0.0862nd of 9Haiku 4.5 · Sonnet 5 · Opus 4.6 · Opus 4.8 · Opus 5 · Fable 5 · GPT 5.6 Luna · GPT 5.6 Terra · GPT 5.6 Sol
CC Barebare harness6 models$0.31020122$0.071$0.939$0.1866th of 9Haiku 4.5 · Sonnet 5 · Opus 4.6 · Opus 4.8 · Opus 5 · Fable 5
Codex Barebare harness3 models$0.09312012$0.004$0.240$0.0761st of 9GPT 5.6 Luna · GPT 5.6 Terra · GPT 5.6 Sol
Cursor Barebare harness8 models$0.17232032$0.005$0.652$0.1533rd of 9Haiku 4.5 · Sonnet 5 · Opus 5 · GPT 5.6 Luna · GPT 5.6 Terra · GPT 5.6 Sol · Grok 4.6 · Composer 2.5
Graphifyindexing tool on Claude Code6 models$0.34220122$0.065$0.901$0.1999th of 9Haiku 4.5 · Sonnet 5 · Opus 4.6 · Opus 4.8 · Opus 5 · Fable 5
GitNexusindexing tool on Claude Code6 models$0.31120122$0.070$0.677$0.1537th of 9Haiku 4.5 · Sonnet 5 · Opus 4.6 · Opus 4.8 · Opus 5 · Fable 5
CodeGraphindexing tool on Claude Code6 models$0.28320322$0.067$0.750$0.1445th of 9Haiku 4.5 · Sonnet 5 · Opus 4.6 · Opus 4.8 · Opus 5 · Fable 5
CodebaseMemoryindexing tool on Claude Code6 models$0.27420122$0.072$0.623$0.1314th of 9Haiku 4.5 · Sonnet 5 · Opus 4.6 · Opus 4.8 · Opus 5 · Fable 5
Serenaindexing tool on Claude Code6 models$0.33820723$0.071$0.898$0.1898th of 9Haiku 4.5 · Sonnet 5 · Opus 4.6 · Opus 4.8 · Opus 5 · Fable 5

Figure 11. Cost per correct answer divides what a run cost by the number of questions it answered correctly: the vendor-reported figure where the harness reported one, otherwise the benchmark's estimate from the published rate tables over the run's four token columns. Each column is the mean over valid, scored runs at one model (or pooled over the arm's roster in the pooled-models view); TrueArchitect and the bare harnesses are the dashed references, and the data table names each row's cost basis. The context saving of Figure 1 becomes a price saving at every model, and the ordering survives the normalisation by correct answers.

How this figure is computed

Definition. For a scored run with score(r) > 0, cost_per_correct(r) = cost(r) / score(r), where cost(r) is the vendor-reported cost when the harness reported one (Claude Code's own result envelope) and otherwise the estimate Σ tokens × rate for the four token columns, using the rate table that applies to the arm: Cursor's published prices for the Cursor arm, Augment's vendor-list-plus-40-percent for Auggie, the vendor list prices for everything else. The tables are published as `epochs/E00N/summary/rate-tables.json`; each run row carries `cost_basis` naming which one priced it.

Aggregation. Cell means averaged with equal weight over the selected batteries, models and protocols. A run whose model has no rate row has no cost and is excluded from the mean, never counted as zero.

Reading it. Lower is better. Because estimates ride list prices, an arm on a subscription plan may pay less in practice; the figure is an API-equivalent price, comparable across arms because it is computed the same way for all of them. Codex reports no cache-write split, so its estimate charges written tokens at the input rate: a disclosed floor of at most 25 percent of the input component.

Disclosures

Vendor-reported and estimated costs sit in the same column; the table beneath the figure carries the basis per row and the rate tables are published beside the run rows.

Provenance

Computed at build time from epochs/E004/summary/runs.json of the proof package (export tool tabench export public 0.1.25, scorer era tabench-1.1). Each row of the table above links to the run directories in the repository, where run.json, the answers, the verdicts and the transcript of every run can be read.