Evo-Bench Construction

Tasks chosen by harnesses, not by hand.

Evo-Bench keeps tasks that respond directionally to better agent harnesses, then builds aligned validation and sealed evaluation suites across three agent domains.

01 · Framework

Harness-guided benchmark construction

Auxiliary tasks produce diverse harnesses. Those harnesses then reveal which candidate benchmark items genuinely measure harness quality.

Two-stage construction. Auxiliary tasks are disjoint from every benchmark source; candidate tasks are retained according to harness sensitivity and split with difficulty stratification.
02 · Selection

Four construction principles

01

Auxiliary tasks, fully disjoint

320 tasks from MiroRL, RedSearcher, Auto-ClawEval and internal sets are corpus- and instance-level disjoint from the five source benchmarks.

02

12 representative harnesses

Four evolution runs produce 73 evaluated harness variants; deterministic diversity-aware selection keeps 12 spanning capability, orchestration and program structure.

03

Sensitivity, not variance

Per-task scores are correlated with leave-one-out harness quality. Correlation measures whether better harnesses win, rather than whether outcomes merely vary.

04

Aligned, disjoint splits

Non-positive-sensitivity tasks are removed. The survivors are stratified by difficulty before drawing aligned validation and evaluation suites.

03 · Item Map

Difficulty and harness sensitivity

Positive sensitivity is the selection gate, while stratification preserves meaningful coverage across the difficulty range.

Candidate item map. The selected region favours tasks on which stronger harnesses improve scores without collapsing difficulty diversity.
04 · Splits

From 2,329 candidates to 608 tasks

SourceCandidate poolValidationEvaluationMean sens.Perf. allPerf. selected
APEX-Agents42132640.3770.3890.224
BrowseComp768321280.2560.4440.271
Claw-Eval15732640.3860.8390.747
GDPval21532640.3680.7970.495
HLE768321280.3130.3330.251
Total2,329160448

Sens. is the Pearson correlation between a task's score and leave-one-out harness quality across the 12 auxiliary harnesses. Perf. is the task's mean performance across those harnesses: “All” summarizes the full candidate pool, while “Selected” summarizes the final validation and evaluation tasks. Task difficulty is defined as 1 − Perf., so a lower value indicates more performance headroom.