GLM-5.2 climbs steeply, then flattens
The largest benefit arrives between the small and medium settings; the full budget still helps, but with smaller marginal returns.
Two controlled studies isolate whether a larger research budget keeps paying and whether an evolved harness survives a change of policy model.
Qwen3.7-Max and GLM-5.2 are evaluated under three increasingly large ceilings.
The largest benefit arrives between the small and medium settings; the full budget still helps, but with smaller marginal returns.
Its improvements are steadier across the three ceilings, suggesting useful search remains available later in the run.
Neither curve reverses at 48 hours / 20 iterations / 1,000 steps, so the main setting captures additional capability rather than only extra compute.
The same Qwen3.7-Max and GLM-5.2 evolvers are evaluated with three different policy models.
| Evolver | Search | Office | General | Overall | ATV |
|---|---|---|---|---|---|
| Policy: Qwen3.6-35B-A3B | |||||
| CodeAct starting point | 2.7 | 14.2 | 35.9 | 13.9 | — |
| Qwen3.7-Max | 12.5 | 33.0 | 48.4 | 27.9 | 29.2 |
| GLM-5.2 | 16.4 | 34.0 | 45.3 | 29.2 | 30.8 |
| Policy: DeepSeek-V4-Flash - main setting | |||||
| CodeAct starting point | 11.7 | 38.4 | 48.4 | 29.7 | — |
| Qwen3.7-Max | 36.3 | 37.8 | 59.4 | 41.5 | 49.3 |
| GLM-5.2 | 45.4 | 39.2 | 48.4 | 43.5 | 51.0 |
| Policy: GLM-5.2 | |||||
| CodeAct starting point | 18.0 | 40.2 | 73.4 | 38.0 | — |
| Qwen3.7-Max | 38.3 | 45.1 | 46.9 | 42.7 | 46.7 |
| GLM-5.2 | 35.6 | 45.2 | 80.3 | 48.4 | 50.4 |
Cross-policy result. Gains hold under every policy swap. With GLM-5.2 as both evolver and policy, the harness rises from 38.0 to 48.4 - above the human-engineered 47.5 measured in the main setting. ATV = AnytimeVal.
No tested policy model erases the benefit of the evolved harness.
The GLM-5.2 policy is strongest on General tasks, while Qwen3.6-35B-A3B produces much lower Search scores.
The harness captures reusable tool use and reasoning structure rather than a single model's response style.
Scaling trends show that comparisons should report the full research budget, not only the final harness score.
Compare all nine evolvers on the held-out evaluation suite.