Figure 7 · E004

Tool calls per run

A tool call is one invocation of any tool the arm offered its model: file reads, searches, shell commands, sub-agents, an indexing tool's query surface, or the TrueArchitect index.

protocols
battery
model
columns

Pooled models shows one column per arm, its models averaged with equal weight — the readable overview. Per model shows one column per model for every arm, for like-for-like reads at a fixed model. This choice applies to every figure on the site and is remembered.

Tool calls per runlower is betterZeroShot + MultiTurn · both batteries
050100150200TrueArchitect 156.7CC Bare 210.2Codex Bare 89.0Cursor Bare 200.6156.79 modelsTrueArchitectTrueArchitect4th of 9210.26 modelsCC Barebare harness8th of 989.03 modelsCodex Barebare harness2nd of 9200.68 modelsCursor Barebare harness6th of 9211.16 modelsGraphifyindexing tool9th of 9149.76 modelsGitNexusindexing tool3rd of 981.56 modelsCodeGraphindexing tool1st of 9160.26 modelsCodebaseMemoryindexing tool5th of 9202.56 modelsSerenaindexing tool7th of 9
Data behind this figure ZeroShot + MultiTurn · both batteries
armrolecolumntool calls per runrunscellsminmaxsdrankruns in the repository
TrueArchitectTrueArchitect9 models156.73303450.0418.067.94th of 9Haiku 4.5 · Sonnet 5 · Opus 4.6 · Opus 4.8 · Opus 5 · Fable 5 · GPT 5.6 Luna · GPT 5.6 Terra · GPT 5.6 Sol
CC Barebare harness6 models210.2201226.0762.0183.88th of 9Haiku 4.5 · Sonnet 5 · Opus 4.6 · Opus 4.8 · Opus 5 · Fable 5
Codex Barebare harness3 models89.01201240.0179.037.72nd of 9GPT 5.6 Luna · GPT 5.6 Terra · GPT 5.6 Sol
Cursor Barebare harness8 models200.63203218.0543.0123.76th of 9Haiku 4.5 · Sonnet 5 · Opus 5 · GPT 5.6 Luna · GPT 5.6 Terra · GPT 5.6 Sol · Grok 4.6 · Composer 2.5
Graphifyindexing tool on Claude Code6 models211.12012247.0661.0169.69th of 9Haiku 4.5 · Sonnet 5 · Opus 4.6 · Opus 4.8 · Opus 5 · Fable 5
GitNexusindexing tool on Claude Code6 models149.72012248.0465.091.23rd of 9Haiku 4.5 · Sonnet 5 · Opus 4.6 · Opus 4.8 · Opus 5 · Fable 5
CodeGraphindexing tool on Claude Code6 models81.52032218.0226.032.81st of 9Haiku 4.5 · Sonnet 5 · Opus 4.6 · Opus 4.8 · Opus 5 · Fable 5
CodebaseMemoryindexing tool on Claude Code6 models160.22012245.0443.098.55th of 9Haiku 4.5 · Sonnet 5 · Opus 4.6 · Opus 4.8 · Opus 5 · Fable 5
Serenaindexing tool on Claude Code6 models202.52072346.0646.0150.37th of 9Haiku 4.5 · Sonnet 5 · Opus 4.6 · Opus 4.8 · Opus 5 · Fable 5

Figure 7. A tool call is one invocation of any tool the arm offered its model: file reads, searches, shell commands, sub-agents, an indexing tool's query surface, or the TrueArchitect index. Each column is the mean number of calls per run over valid, scored runs, split by model for TrueArchitect and the bare harnesses. Fewer calls with higher accuracy is the signature of a good index: the model asks its question once instead of searching for the answer many times.

How this figure is computed

Definition. tool_calls(r) is the count of tool invocations recorded for the run across all sessions and threads; the run record breaks the count down by tool and by class (built-in, sub-agent, skill, an installed tool's MCP surface, a native harness's own tools, TrueArchitect file tools, and TrueArchitect index queries).

Aggregation. Cell means averaged with equal weight over the selected batteries and models.

Reading it. Lower is better only in combination with Figures 2 and 3: a low count with low accuracy is an arm that gave up early. Read the three together.

Provenance

Computed at build time from epochs/E004/summary/runs.json of the proof package (export tool tabench export public 0.1.25, scorer era tabench-1.1). Each row of the table above links to the run directories in the repository, where run.json, the answers, the verdicts and the transcript of every run can be read.