Center for AI Safety & Scale AI Releases: Humanity's Last Exam (HLE)

safe.ai ↗

ThursdAI — the weekly AI news podcast hosted by Alex Volkov — has covered one Center for AI Safety & Scale AI release so far: Humanity's Last Exam (HLE) on Jan 23, 2025. Humanity's Last Exam: a deliberately unsaturated frontier benchmark. The entry below has primary-source links and the episode where we covered it live.

1 release1 episodecovered Jan 23, 2025

January 2025 1

Benchmarks & Evals

Humanity's Last Exam (HLE)

Humanity's Last Exam: a deliberately unsaturated frontier benchmark

Humanity's Last Exam (HLE) launched as a new, very hard benchmark designed to stay unsaturated as models max out MMLU and math evals. It crowdsourced expert-level questions to measure frontier model capability where existing benchmarks are at 98-99% saturation.