Every model improves the harness, but the human frontier still leads
All nine evolvers raise the Overall Score above the 29.7 CodeAct starting point, showing that harness evolution delivers gains across model families. Yet the strongest evolved result, GPT-5.6 Sol at 46.3, still trails the 47.5 human-engineered Artificial Harness.