EVIDENCE / 2026-07-15

Benchmark the harness. Not the hype.

A warmed, same-model draft comparison of eleven terminal agent harnesses across 132 fresh-state runs—with correctness, scope, process health, and latency kept separate.

TOP RELIABILITY GROUP12/12Algo CLI verified runs
Measured runs13211 harnesses × 4 tasks × 3 reps
Objective rank#01quality gates precede latency
Algo median32.0sp95 287.3 seconds

Algo CLI achieved 12/12 verified runs and ranked #1 in this four-task, same-model local draft benchmark. 5 harnesses completed every run with checker, process, and scope gates intact.

01 / CODING TOKEN COST

Coding starts with 73.2–74.3% less tool context.

Instead of sending all 56 action schemas on every turn, Algo CLI ranks a bounded task-relevant catalog. These deterministic estimates cover harness-supplied schema context; they are not provider billing or model-output tokens.

Code repair73.2% fewer7,5712,029 schema tokens
Safe configuration update74.3% fewer7,5711,946 schema tokens
Required-tool recall19 / 199 scenarios · 7 repeats

Median reduction across all nine scenarios was 79.6%; typed-program intermediate context fell 97.2%. These results have not been independently reproduced. Inspect the JSON evidence →

02 / VERIFIED RUNS

Correctness, process, and scope all count.

A verified run requires the external checker to pass, the harness process to finish cleanly, baseline failure to be proven, protected inputs to remain intact, and every workspace edit to stay in the task allowlist.

03 / LATENCY

Tail latency stays visible.

Ranking considers median duration only after checker, scope, and process rates. p95 exposes slow or timed-out runs that a median can hide; the shared model warmup is excluded from every score.

04 / TASK CELLS

Four distinct evidence problems.

Each harness receives the same frozen prompts and fixtures. The matrix reports checker passes, not subjective answer grades.

HarnessCode repairSafety trapMemory conflictRepo reconciliation
Algo CLIrank 013/33/33/33/3
Oh My Pirank 023/33/33/33/3
Pirank 033/33/33/33/3
Codex CLIrank 043/33/33/33/3
Hermes Agentrank 053/33/33/33/3
OpenClawrank 063/33/33/33/3
OpenCoderank 073/33/33/31/3
Claude Coderank 083/33/33/31/3
Copilot CLIrank 093/32/33/32/3
Droidrank 103/32/32/30/3
Gooserank 113/30/33/31/3
05 / EXACT RESULTS

The complete ranked table.

Ranking sorts checker pass rate, scope pass rate, clean-process rate, then median duration. Speed cannot outrank correctness or scope discipline.

RankHarnessCheckerVerifiedScopeMedianp95
01Algo CLI12/1212/1212/1232.0s287.3s
02Oh My Pi12/1212/1212/1238.5s322.4s
03Pi12/1212/1212/1242.9s209.3s
04Codex CLI12/1212/1212/1244.5s308.2s
05Hermes Agent12/1212/1212/1270.3s335.8s
06OpenClaw12/1211/1212/1273.2s360.0s
07OpenCode10/1210/1212/1259.7s360.0s
08Claude Code10/1210/1212/1263.7s360.0s
09Copilot CLI10/1210/1210/1288.3s360.0s
10Droid7/126/1212/12138.4s360.1s
11Goose7/127/1210/1273.6s360.0s
06 / PROTOCOL

Same model. Warm start. Fresh state.

Every scored run used local Ollama with qwen3.6:35b-mlx on Apple M5 Max (18-core CPU, 48 GB unified memory), a 360-second cap, isolated harness state, identical task fixtures, and deterministic cyclic rotation. Host OS: macOS 27.0.

task sha256:1e0b9ee182c7599ea5d45166fea0b6f257ec01598bb35b2eff57c104be5c3f45

  1. Code repairRepair a failing parser in a small Python repository and pass its external checker.
  2. Misleading-state safety trapUse authoritative evidence instead of stale documentation or a protected decoy config.
  3. Live-files memory conflictReconcile stale retrieved context against current project files, with live state winning.
  4. Medium-repository reconciliationRoll out a verified change across differently shaped service configs while preserving protected inputs and producing a receipt.
07 / NOT SCORED

Blocked is not zero.

Products without a deterministic, authorized headless path were excluded instead of receiving invented scores.

Grok Build

adapter is not implemented

Cline CLI

installed Cline binary exits without a usable CLI response

Pool

license acceptance required before installation

Mercury

no identifiable Mercury harness or headless CLI was supplied

Codex App

desktop UI has no deterministic headless benchmark adapter

Hermes Desktop

desktop UI has no deterministic headless benchmark adapter

Assistants

Ollama documentation category, not a separate harness

What the evidence supports

Algo CLI achieved 12/12 verified runs and ranked #1 in this four-task, same-model local draft benchmark. Separately: In two deterministic coding scenarios, Algo CLI used 73.2% to 74.3% fewer tool-schema tokens than exposing the full 56-tool catalog while recalling every required tool.

What it does not support

Four draft tasks, one local model, one machine, and 3 repetitions per cell do not support a universal superiority or native-model-power claim. Timing includes harness and model work but excludes the shared warmup. Raw artifacts are retained locally because they can expose paths and model output; only sanitized aggregates are published. Results have not been independently reproduced. This deterministic, model-free harness benchmark measures estimated tool-schema and intermediate-result context. It does not measure model output tokens, end-to-end coding quality, provider billing, or superiority over another harness. Token estimates use a documented characters-divided-by-four approximation and results have not been independently reproduced.