Figure 8 · E004

Wall time per run

Wall time is the sum of the active answering segments of a run: the session for HumanExam, each question's session for ZeroShot, each turn for MultiTurn.

protocols
battery
model
columns

Pooled models shows one column per arm, its models averaged with equal weight — the readable overview. Per model shows one column per model for every arm, for like-for-like reads at a fixed model. This choice applies to every figure on the site and is remembered.

Wall time per runlower is betterZeroShot + MultiTurn · both batteries
0s0s0s0s0s0s
Data behind this figure ZeroShot + MultiTurn · both batteries
armrolecolumnactive answering time per runrunscellsminmaxsdrankruns in the repository

Figure 8. Wall time is the sum of the active answering segments of a run: the session for HumanExam, each question's session for ZeroShot, each turn for MultiTurn. Each column is the mean over valid, scored runs; the benchmark's own dispatch and capture time between segments is excluded. Timing is environment-coupled and reported as indicative magnitude only; see the disclosures below before comparing across the host boundary.

How this figure is computed

Definition. wall(r) = Σ duration of the run's answering segments, as recorded by the lane driver. Provider latency, lane parallelism and retry back-offs all ride this figure.

Aggregation. Cell means averaged with equal weight. Runs whose duration was not captured are excluded, never counted as zero.

Reading it. Lower is better, with the caveat that this is not a controlled latency benchmark.

Disclosures

The TrueArchitect arm executed on the host through its gateway; the comparison group executed in containers on the same machine. Accuracy, tokens and tool counts are location-independent; wall time is not, and the figure is published for completeness rather than as a claim.

Runs executed in parallel lanes shared the host; per-lane contention is not modelled.

Provenance

Computed at build time from epochs/E004/summary/runs.json of the proof package (export tool tabench export public 0.1.25, scorer era tabench-1.1). Each row of the table above links to the run directories in the repository, where run.json, the answers, the verdicts and the transcript of every run can be read.