Tool calls per run
A tool call is one invocation of any tool the arm offered its model: file reads, searches, shell commands, sub-agents, an indexing tool's query surface, or the TrueArchitect index.
Pooled models shows one column per arm, its models averaged with equal weight — the readable overview. Per model shows one column per model for every arm, for like-for-like reads at a fixed model. This choice applies to every figure on the site and is remembered.
Data behind this figure — ZeroShot + MultiTurn · both batteries
| arm | role | column | tool calls per run | runs | cells | min | max | sd | rank | runs in the repository |
|---|---|---|---|---|---|---|---|---|---|---|
| TrueArchitect | TrueArchitect | 9 models | 156.7 | 330 | 34 | 50.0 | 418.0 | 67.9 | 4th of 9 | Haiku 4.5 · Sonnet 5 · Opus 4.6 · Opus 4.8 · Opus 5 · Fable 5 · GPT 5.6 Luna · GPT 5.6 Terra · GPT 5.6 Sol |
| CC Bare | bare harness | 6 models | 210.2 | 201 | 22 | 6.0 | 762.0 | 183.8 | 8th of 9 | Haiku 4.5 · Sonnet 5 · Opus 4.6 · Opus 4.8 · Opus 5 · Fable 5 |
| Codex Bare | bare harness | 3 models | 89.0 | 120 | 12 | 40.0 | 179.0 | 37.7 | 2nd of 9 | GPT 5.6 Luna · GPT 5.6 Terra · GPT 5.6 Sol |
| Cursor Bare | bare harness | 8 models | 200.6 | 320 | 32 | 18.0 | 543.0 | 123.7 | 6th of 9 | Haiku 4.5 · Sonnet 5 · Opus 5 · GPT 5.6 Luna · GPT 5.6 Terra · GPT 5.6 Sol · Grok 4.6 · Composer 2.5 |
| Graphify | indexing tool on Claude Code | 6 models | 211.1 | 201 | 22 | 47.0 | 661.0 | 169.6 | 9th of 9 | Haiku 4.5 · Sonnet 5 · Opus 4.6 · Opus 4.8 · Opus 5 · Fable 5 |
| GitNexus | indexing tool on Claude Code | 6 models | 149.7 | 201 | 22 | 48.0 | 465.0 | 91.2 | 3rd of 9 | Haiku 4.5 · Sonnet 5 · Opus 4.6 · Opus 4.8 · Opus 5 · Fable 5 |
| CodeGraph | indexing tool on Claude Code | 6 models | 81.5 | 203 | 22 | 18.0 | 226.0 | 32.8 | 1st of 9 | Haiku 4.5 · Sonnet 5 · Opus 4.6 · Opus 4.8 · Opus 5 · Fable 5 |
| CodebaseMemory | indexing tool on Claude Code | 6 models | 160.2 | 201 | 22 | 45.0 | 443.0 | 98.5 | 5th of 9 | Haiku 4.5 · Sonnet 5 · Opus 4.6 · Opus 4.8 · Opus 5 · Fable 5 |
| Serena | indexing tool on Claude Code | 6 models | 202.5 | 207 | 23 | 46.0 | 646.0 | 150.3 | 7th of 9 | Haiku 4.5 · Sonnet 5 · Opus 4.6 · Opus 4.8 · Opus 5 · Fable 5 |
Figure 7. A tool call is one invocation of any tool the arm offered its model: file reads, searches, shell commands, sub-agents, an indexing tool's query surface, or the TrueArchitect index. Each column is the mean number of calls per run over valid, scored runs, split by model for TrueArchitect and the bare harnesses. Fewer calls with higher accuracy is the signature of a good index: the model asks its question once instead of searching for the answer many times.
How this figure is computed
Definition. tool_calls(r) is the count of tool invocations recorded for the run across all sessions and threads; the run record breaks the count down by tool and by class (built-in, sub-agent, skill, an installed tool's MCP surface, a native harness's own tools, TrueArchitect file tools, and TrueArchitect index queries).
Aggregation. Cell means averaged with equal weight over the selected batteries and models.
Reading it. Lower is better only in combination with Figures 2 and 3: a low count with low accuracy is an arm that gave up early. Read the three together.
Provenance
Computed at build time from epochs/E004/summary/runs.json of the proof package (export tool tabench export public 0.1.25, scorer era tabench-1.1). Each row of the table above links to the run directories in the repository, where run.json, the answers, the verdicts and the transcript of every run can be read.