Assistant Benchmark
Assistant Benchmark ranks 116 submitted AI assistants across 16 hand-tested dimensions
David Pawlan of Merit Systems launched Assistant Benchmark, a use-case-driven leaderboard for personal AI assistants: 116 assistants submitted across categories like travel, email, finance, and work-in-teams, scored on 16 dimensions including memory, recommendations, and online tasks. Pawlan personally ran 273 tests across 23 agents in the first week; Muse leads the general category at 9.1 with Instinct at 8.4. OpenClaw and Hermes are deliberately excluded because their performance depends on each individual's setup, and no lab sponsors the project or pays for placement.