Terminal Bench 2.0 — GPT-5.6 results
Wolfbench: GPT-5.6 Sol on max thinking beats GPT-5.5 on both score and cost
Wolfram's Wolfbench, run on CoreWeave, added GPT-5.6 Sol, Terra, and Luna to its Terminal Bench 2.0 leaderboard. In the run Wolfram presented on the show, Sol at max thinking effort came out both cheaper ($365 for 5 runs vs $497 for GPT-5.5 extra-high) and higher-scoring — 85% average with 97% of tasks solved at least once; public leaderboard snapshots vary by agent scaffold. All traces logged to Weights & Biases; the benchmark is fully open source.