Agent Harness Leaderboard — Edition 1
Judging criteria pre-registered & sha256-anchored before any score existed · Published 2026-10-12 · Task set: SWE-bench Verified (500 human-verified instances, resolved %)
🔒 Rules locked firstThe complete judging rules were frozen and hashed before any score was recorded. Rule edits change the hash — detectable by anyone.
♻️ ReproducibleEvery entry ships with the exact eval stack tags and commands. Re-derive any number yourself.
📋 Config disclosureTool surface, inference backend, max steps listed per harness. A score without its config is not a score.
The Board
| Harness | Tool surface | Inference backend | Max steps | Resolved % | Score source |
| lands 2026-10-12 | from custodian-verified registry | — | — | — | — |
Honest framing: Edition 1 is a compilation board — official published scores normalized under the published criteria. It is not our own re-runs. Same-rig paired runs begin in Edition 2. We state this up front because method honesty matters more than impressive numbers.
Criteria
Edition-1 judging criteria: sha256-16 = {DATA_CRITERIA_SHA16} · full text: {DATA_CRITERIA_URL}
Reproduction commands: {DATA_REPRO_CMD}
Your harness, judged under the same locked rules.
Board listing:
free · Deep report (repro pack + improvement notes):
$199 · Enterprise / custom eval:
from $999
SLA: L1 24h card · L2 48h report · L3 5 business days · ≤2 revision rounds · criteria sha pre-registered & verifiable
Wondering what a $199 deep report looks like?
Read the sample report (real data, harness under test = our own)
Submit for verification →