Artificial Analysis

2 releases covered on ThursdAI · artificialanalysis.ai ↗

August 2026

Artificial Analysis
Benchmarks & Evals

Endpoint Accuracy Index

Endpoint Accuracy Index: the same open weights score 52% to 100% depending on your provider

Artificial Analysis launched an index measuring whether inference providers actually serve the model they claim: GLM-5.2 scores 52% on one provider and 100% on three others with a 5.2x price spread, and some gpt-oss-120b endpoints hit 22% on tool calling versus the reference's 37%. Low scorers produce about half the output tokens per task, pointing at aggressive quantization. v1.0 covers GLM-5.2 (15 providers), gpt-oss-120b (20), and DeepSeek V4 Pro (9), with Kimi K3 next. CoreWeave came in cheapest on gpt-oss-120b at $0.04/M blended with 98% accuracy.

52% vs 100% same GLM-5.2 weights across providers5.2x price spread across GLM-5.2 endpoints$0.04/M CoreWeave's gpt-oss-120b blended price at 98% accuracy

May 2026

Artificial Analysis
Benchmarks & Evals

Coding Agent Index

Artificial Analysis Coding Agent Index benchmarks model + harness combos

Artificial Analysis launched the Coding Agent Index, a benchmark that evaluates model and harness combinations rather than models alone. Opus 4.7 in Cursor CLI leads at 61, GLM-5.1 tops the open-weight entries at 53, and costs vary 30x across combos for similar capability.