WolfBench (Wolfram Ravenwolf)

3 releases covered on ThursdAI · wolfbench.ai ↗

July 2026

Wolfbench
Benchmarks & EvalsOpen weights

Terminal Bench 2.0 — GPT-5.6 results

Wolfbench: GPT-5.6 Sol on max thinking beats GPT-5.5 on both score and cost

Wolfram's Wolfbench, run on CoreWeave, added GPT-5.6 Sol, Terra, and Luna to its Terminal Bench 2.0 leaderboard. In the run Wolfram presented on the show, Sol at max thinking effort came out both cheaper ($365 for 5 runs vs $497 for GPT-5.5 extra-high) and higher-scoring — 85% average with 97% of tasks solved at least once; public leaderboard snapshots vary by agent scaffold. All traces logged to Weights & Biases; the benchmark is fully open source.

$365 vs $497 Sol max vs GPT-5.5 extra-high, per 5 runs85% / 97% average score / tasks solved at least once

June 2026

Major Features & Updates

WolfBench Token-Usage Visualization

WolfBench adds 3D token-depth bars to show model efficiency

Wolfram Ravenwolf shipped a WolfBench feature that visualizes token usage alongside benchmark score as 3D token-depth bars. Two models can look close on a leaderboard while one burns dramatically more tokens, which changes the real cost and latency story; Gemini 3.5 Flash and GPT 5.5 were compared as examples.

April 2026

Benchmarks & Evals

WolfBench

WolfBench results show Hermes Agent beating Claude Code and OpenClaw

Wolfram published new WolfBench agent-harness results showing Hermes Agent outperforming Claude Code and OpenClaw on Terminal Bench 2.0 across most model combinations. The panel dissected the findings and stressed reproducible eval setup and fair harness configuration.