Evo-Bench Ablation Studies

What changes when budget and policy change?

Two controlled studies isolate whether a larger research budget keeps paying and whether an evolved harness survives a change of policy model.

01 · Budget Scaling

More research budget continues to help

Qwen3.7-Max and GLM-5.2 are evaluated under three increasingly large ceilings.

Budget scaling. 24h / 10 iterations / 500 steps → 36h / 15 / 750 → 48h / 20 / 1,000. Both evolvers improve monotonically.
01

GLM-5.2 climbs steeply, then flattens

The largest benefit arrives between the small and medium settings; the full budget still helps, but with smaller marginal returns.

02

Qwen3.7-Max grows more linearly

Its improvements are steadier across the three ceilings, suggesting useful search remains available later in the run.

03

The default ceiling is not arbitrary

Neither curve reverses at 48 hours / 20 iterations / 1,000 steps, so the main setting captures additional capability rather than only extra compute.

02 · Cross-Policy Transfer

Evolved harnesses survive a policy swap

The same Qwen3.7-Max and GLM-5.2 evolvers are evaluated with three different policy models.

EvolverSearchOfficeGeneralOverallATV
Policy: Qwen3.6-35B-A3B
CodeAct starting point2.714.235.913.9
Qwen3.7-Max12.533.048.427.929.2
GLM-5.216.434.045.329.230.8
Policy: DeepSeek-V4-Flash - main setting
CodeAct starting point11.738.448.429.7
Qwen3.7-Max36.337.859.441.549.3
GLM-5.245.439.248.443.551.0
Policy: GLM-5.2
CodeAct starting point18.040.273.438.0
Qwen3.7-Max38.345.146.942.746.7
GLM-5.235.645.280.348.450.4

Cross-policy result. Gains hold under every policy swap. With GLM-5.2 as both evolver and policy, the harness rises from 38.0 to 48.4 - above the human-engineered 47.5 measured in the main setting. ATV = AnytimeVal.

03 · Takeaway

The improvement is structural, not policy-specific

Every swap preserves a positive gain

No tested policy model erases the benefit of the evolved harness.

Absolute scores remain policy-dependent

The GLM-5.2 policy is strongest on General tasks, while Qwen3.6-35B-A3B produces much lower Search scores.

Transfer argues against endpoint overfitting

The harness captures reusable tool use and reasoning structure rather than a single model's response style.

Budget remains a genuine experimental variable

Scaling trends show that comparisons should report the full research budget, not only the final harness score.