Auxiliary tasks, fully disjoint
320 tasks from MiroRL, RedSearcher, Auto-ClawEval and internal sets are corpus- and instance-level disjoint from the five source benchmarks.
Evo-Bench keeps tasks that respond directionally to better agent harnesses, then builds aligned validation and sealed evaluation suites across three agent domains.
Auxiliary tasks produce diverse harnesses. Those harnesses then reveal which candidate benchmark items genuinely measure harness quality.
320 tasks from MiroRL, RedSearcher, Auto-ClawEval and internal sets are corpus- and instance-level disjoint from the five source benchmarks.
Four evolution runs produce 73 evaluated harness variants; deterministic diversity-aware selection keeps 12 spanning capability, orchestration and program structure.
Per-task scores are correlated with leave-one-out harness quality. Correlation measures whether better harnesses win, rather than whether outcomes merely vary.
Non-positive-sensitivity tasks are removed. The survivors are stratified by difficulty before drawing aligned validation and evaluation suites.
Positive sensitivity is the selection gate, while stratification preserves meaningful coverage across the difficulty range.
| Source | Candidate pool | Validation | Evaluation | Mean sens. | Perf. all | Perf. selected |
|---|---|---|---|---|---|---|
| APEX-Agents | 421 | 32 | 64 | 0.377 | 0.389 | 0.224 |
| BrowseComp | 768 | 32 | 128 | 0.256 | 0.444 | 0.271 |
| Claw-Eval | 157 | 32 | 64 | 0.386 | 0.839 | 0.747 |
| GDPval | 215 | 32 | 64 | 0.368 | 0.797 | 0.495 |
| HLE | 768 | 32 | 128 | 0.313 | 0.333 | 0.251 |
| Total | 2,329 | 160 | 448 | — | — | — |
Sens. is the Pearson correlation between a task's score and leave-one-out harness quality across the 12 auxiliary harnesses. Perf. is the task's mean performance across those harnesses: “All” summarizes the full candidate pool, while “Selected” summarizes the final validation and evaluation tasks. Task difficulty is defined as 1 − Perf., so a lower value indicates more performance headroom.
Explore gains, failure trajectories, domain differences, cost and integrity checks.