Pass rate by question category and difficulty
Every question carries one or more categories and a difficulty; each cell is the pass rate of an arm on the questions in that category (or at that difficulty), pooled over the selected models — by default the Anthropic models, which every arm ran.
Each cell pools the arm's Anthropic models with equal weight per model, so every arm is compared over the same vendor's models; an arm that ran none of them shows a gap.
Data behind this figure — ZeroShot · both batteries · Anthropic models
| arm | column | pass rate | attempts | correct | models |
|---|---|---|---|---|---|
| ta-ask-fz-009 | absence-proof | 82.0% | 100 | 82 | 5 |
| ta-ask-fz-009 | aggregation | 84.0% | 100 | 84 | 5 |
| ta-ask-fz-009 | calls | 96.0% | 448 | 429 | 6 |
| ta-ask-fz-009 | collision | 92.0% | 100 | 92 | 5 |
| ta-ask-fz-009 | cross-stack | 98.5% | 448 | 441 | 6 |
| ta-ask-fz-009 | db | 99.2% | 448 | 444 | 6 |
| ta-ask-fz-009 | dispatch | 99.0% | 100 | 99 | 5 |
| ta-ask-fz-009 | endpoints | 99.8% | 504 | 503 | 6 |
| ta-ask-fz-009 | field-grain | 67.0% | 100 | 67 | 5 |
| ta-ask-fz-009 | impact | 98.1% | 336 | 329 | 6 |
| ta-ask-fz-009 | multi-hop | 91.0% | 100 | 91 | 5 |
| ta-ask-fz-009 | nameless-entry | 100.0% | 150 | 150 | 5 |
| ta-ask-fz-009 | reality-check | 97.2% | 250 | 243 | 5 |
| ta-ask-fz-009 | types | 98.0% | 280 | 274 | 6 |
| ta-ask-fz-009 | difficulty 1 | 100.0% | 560 | 560 | 6 |
| ta-ask-fz-009 | difficulty 2 | 98.0% | 560 | 548 | 6 |
| ta-ask-fz-009 | difficulty 3 | 97.5% | 560 | 545 | 6 |
| ta-ask-fz-009 | difficulty 4 | 90.8% | 1000 | 908 | 5 |
| cold | absence-proof | 76.0% | 100 | 76 | 5 |
| cold | aggregation | 60.0% | 100 | 60 | 5 |
| cold | calls | 89.3% | 360 | 318 | 5 |
| cold | collision | 56.0% | 100 | 56 | 5 |
| cold | cross-stack | 96.3% | 360 | 345 | 5 |
| cold | db | 96.3% | 360 | 345 | 5 |
| cold | dispatch | 84.0% | 100 | 84 | 5 |
| cold | endpoints | 93.8% | 405 | 377 | 5 |
| cold | field-grain | 63.0% | 100 | 63 | 5 |
| cold | impact | 81.7% | 270 | 216 | 5 |
| cold | multi-hop | 88.0% | 100 | 88 | 5 |
| cold | nameless-entry | 88.7% | 150 | 133 | 5 |
| cold | reality-check | 84.4% | 250 | 211 | 5 |
| cold | types | 94.4% | 225 | 211 | 5 |
| cold | difficulty 1 | 98.4% | 450 | 442 | 5 |
| cold | difficulty 2 | 94.0% | 450 | 420 | 5 |
| cold | difficulty 3 | 88.0% | 450 | 391 | 5 |
| cold | difficulty 4 | 77.1% | 1000 | 771 | 5 |
| cursor | absence-proof | 83.3% | 60 | 50 | 3 |
| cursor | aggregation | 70.0% | 60 | 42 | 3 |
| cursor | calls | 97.5% | 240 | 234 | 3 |
| cursor | collision | 61.7% | 60 | 37 | 3 |
| cursor | cross-stack | 98.8% | 240 | 237 | 3 |
| cursor | db | 100.0% | 240 | 240 | 3 |
| cursor | dispatch | 95.0% | 60 | 57 | 3 |
| cursor | endpoints | 99.6% | 270 | 269 | 3 |
| cursor | field-grain | 80.0% | 60 | 48 | 3 |
| cursor | impact | 97.2% | 180 | 175 | 3 |
| cursor | multi-hop | 98.3% | 60 | 59 | 3 |
| cursor | nameless-entry | 100.0% | 90 | 90 | 3 |
| cursor | reality-check | 92.7% | 150 | 139 | 3 |
| cursor | types | 98.0% | 150 | 147 | 3 |
| cursor | difficulty 1 | 100.0% | 300 | 300 | 3 |
| cursor | difficulty 2 | 98.0% | 300 | 294 | 3 |
| cursor | difficulty 3 | 98.0% | 300 | 294 | 3 |
| cursor | difficulty 4 | 87.0% | 600 | 522 | 3 |
| graphify | absence-proof | 75.0% | 100 | 75 | 5 |
| graphify | aggregation | 51.0% | 100 | 51 | 5 |
| graphify | calls | 88.8% | 360 | 315 | 5 |
| graphify | collision | 63.0% | 100 | 63 | 5 |
| graphify | cross-stack | 95.5% | 360 | 342 | 5 |
| graphify | db | 98.0% | 360 | 352 | 5 |
| graphify | dispatch | 90.0% | 100 | 90 | 5 |
| graphify | endpoints | 96.9% | 405 | 391 | 5 |
| graphify | field-grain | 66.0% | 100 | 66 | 5 |
| graphify | impact | 89.0% | 270 | 237 | 5 |
| graphify | multi-hop | 92.0% | 100 | 92 | 5 |
| graphify | nameless-entry | 92.0% | 150 | 138 | 5 |
| graphify | reality-check | 88.0% | 250 | 220 | 5 |
| graphify | types | 97.2% | 225 | 218 | 5 |
| graphify | difficulty 1 | 99.6% | 450 | 448 | 5 |
| graphify | difficulty 2 | 95.4% | 450 | 427 | 5 |
| graphify | difficulty 3 | 90.0% | 450 | 400 | 5 |
| graphify | difficulty 4 | 79.5% | 1000 | 795 | 5 |
| gitnexus | absence-proof | 84.0% | 100 | 84 | 5 |
| gitnexus | aggregation | 49.0% | 100 | 49 | 5 |
| gitnexus | calls | 89.5% | 360 | 318 | 5 |
| gitnexus | collision | 61.0% | 100 | 61 | 5 |
| gitnexus | cross-stack | 96.5% | 360 | 347 | 5 |
| gitnexus | db | 98.8% | 360 | 355 | 5 |
| gitnexus | dispatch | 98.0% | 100 | 98 | 5 |
| gitnexus | endpoints | 98.9% | 405 | 401 | 5 |
| gitnexus | field-grain | 71.0% | 100 | 71 | 5 |
| gitnexus | impact | 88.3% | 270 | 235 | 5 |
| gitnexus | multi-hop | 91.0% | 100 | 91 | 5 |
| gitnexus | nameless-entry | 96.0% | 150 | 144 | 5 |
| gitnexus | reality-check | 91.2% | 250 | 228 | 5 |
| gitnexus | types | 98.4% | 225 | 221 | 5 |
| gitnexus | difficulty 1 | 99.4% | 450 | 447 | 5 |
| gitnexus | difficulty 2 | 98.8% | 450 | 444 | 5 |
| gitnexus | difficulty 3 | 90.2% | 450 | 402 | 5 |
| gitnexus | difficulty 4 | 82.6% | 1000 | 826 | 5 |
| codegraph | absence-proof | 55.0% | 100 | 55 | 5 |
| codegraph | aggregation | 54.0% | 100 | 54 | 5 |
| codegraph | calls | 92.0% | 360 | 328 | 5 |
| codegraph | collision | 70.0% | 100 | 70 | 5 |
| codegraph | cross-stack | 93.3% | 360 | 333 | 5 |
| codegraph | db | 96.5% | 360 | 346 | 5 |
| codegraph | dispatch | 87.0% | 100 | 87 | 5 |
| codegraph | endpoints | 97.3% | 405 | 393 | 5 |
| codegraph | field-grain | 46.0% | 100 | 46 | 5 |
| codegraph | impact | 86.7% | 270 | 230 | 5 |
| codegraph | multi-hop | 92.0% | 100 | 92 | 5 |
| codegraph | nameless-entry | 100.0% | 150 | 150 | 5 |
| codegraph | reality-check | 83.6% | 250 | 209 | 5 |
| codegraph | types | 96.0% | 225 | 215 | 5 |
| codegraph | difficulty 1 | 99.8% | 450 | 449 | 5 |
| codegraph | difficulty 2 | 95.2% | 450 | 426 | 5 |
| codegraph | difficulty 3 | 90.8% | 450 | 404 | 5 |
| codegraph | difficulty 4 | 76.3% | 1000 | 763 | 5 |
| cbm | absence-proof | 80.0% | 100 | 80 | 5 |
| cbm | aggregation | 47.0% | 100 | 47 | 5 |
| cbm | calls | 90.8% | 360 | 323 | 5 |
| cbm | collision | 60.0% | 100 | 60 | 5 |
| cbm | cross-stack | 95.5% | 360 | 342 | 5 |
| cbm | db | 98.3% | 360 | 353 | 5 |
| cbm | dispatch | 90.0% | 100 | 90 | 5 |
| cbm | endpoints | 99.8% | 405 | 404 | 5 |
| cbm | field-grain | 70.0% | 100 | 70 | 5 |
| cbm | impact | 93.7% | 270 | 251 | 5 |
| cbm | multi-hop | 87.0% | 100 | 87 | 5 |
| cbm | nameless-entry | 100.0% | 150 | 150 | 5 |
| cbm | reality-check | 90.4% | 250 | 226 | 5 |
| cbm | types | 98.8% | 225 | 222 | 5 |
| cbm | difficulty 1 | 99.8% | 450 | 449 | 5 |
| cbm | difficulty 2 | 95.6% | 450 | 428 | 5 |
| cbm | difficulty 3 | 94.2% | 450 | 421 | 5 |
| cbm | difficulty 4 | 81.0% | 1000 | 810 | 5 |
| serena | absence-proof | 80.8% | 108 | 85 | 6 |
| serena | aggregation | 62.1% | 108 | 67 | 6 |
| serena | calls | 90.5% | 360 | 323 | 5 |
| serena | collision | 64.6% | 108 | 70 | 6 |
| serena | cross-stack | 95.3% | 360 | 341 | 5 |
| serena | db | 96.8% | 360 | 347 | 5 |
| serena | dispatch | 92.5% | 108 | 99 | 6 |
| serena | endpoints | 95.1% | 405 | 383 | 5 |
| serena | field-grain | 70.8% | 108 | 73 | 6 |
| serena | impact | 85.7% | 270 | 228 | 5 |
| serena | multi-hop | 87.1% | 108 | 94 | 6 |
| serena | nameless-entry | 91.1% | 162 | 146 | 6 |
| serena | reality-check | 84.8% | 270 | 226 | 6 |
| serena | types | 98.0% | 225 | 220 | 5 |
| serena | difficulty 1 | 99.0% | 450 | 445 | 5 |
| serena | difficulty 2 | 95.0% | 450 | 425 | 5 |
| serena | difficulty 3 | 89.8% | 450 | 400 | 5 |
| serena | difficulty 4 | 80.7% | 1080 | 860 | 6 |
Figure 10. Every question carries one or more categories and a difficulty; each cell is the pass rate of an arm on the questions in that category (or at that difficulty), pooled over the selected models — by default the Anthropic models, which every arm ran. The number in each cell is the pass rate; the color is that cell's difference from bare Claude Code in the same column — blue above, orange below, neutral at the baseline — so a row reads as where an arm beats or trails the null hypothesis. The categories where an index should matter most, structural questions such as callers, impact and cross-stack tracing, are where the separation is widest.
How this figure is computed
Definition. For an arm, model, exam, battery and question, the per-question record gives pass (attempts marked correct) and n (attempts). For a category c, the pass rate at one model is Σ pass / Σ n over the questions carrying c; the cell value is the equal-weight mean over the arm's models. Difficulty cells are formed the same way over the questions at that difficulty.
Aggregation. One question may carry several categories and then contributes to each; category rates are therefore not additive across categories. The model control decides which models a cell pools, with equal weight per model. The figure opens on the Anthropic roster because every arm ran those models, so the cells compare like with like; "OpenAI models" does the same for the GPT roster (TrueArchitect, Codex and Cursor); a single model is the sharpest read. With "all" a cell pools the arm's OWN roster: TrueArchitect ran GPT 5.6 Luna, Terra and Sol as well as the Claude models, while bare Claude Code and the indexing tools ran Claude models only, so a pooled TrueArchitect cell averages in three models its comparators never ran. The multi-hop column is the clearest case: two questions (h06, h07) on which TrueArchitect at every Claude model matches or beats bare Claude Code, while the GPT models score far lower and pull the all-models cell from 91 down to 69; field-grain flips the same way. The comparison pages carry this grid restricted to their own arms and models.
Reading it. Higher is better. Read a row to see an arm's profile and a column to see which arms handle a question class. Color encodes polarity, not magnitude: each cell is colored by its distance from bare Claude Code's pass rate in the same column (the column mean if that arm is absent from the selection), on a symmetric scale whose half-width is the largest difference in the figure and never less than ten points; the number is always the pass rate itself. A two-hue diverging scale with a neutral midpoint was chosen over a green-to-red one because the latter is unreadable under red–green color vision deficiency.
Provenance
Computed at build time from epochs/E004/summary/runs.json of the proof package (export tool tabench export public 0.1.25, scorer era tabench-1.1). Each row of the table above links to the run directories in the repository, where run.json, the answers, the verdicts and the transcript of every run can be read.