David Pawlan joined Alex Volkov on ThursdAI to walk through Assistant Benchmark, his use-case-driven leaderboard for personal AI assistants. He personally ran 273 tests across 23 agents in the first week; 116 submitted assistants now get scored on 16 dimensions like memory and recommendations, with no lab sponsors — and setup-dependent stacks like OpenClaw and Hermes excluded so scores actually transfer.
An assistant isn't something that pings you — it completes the task end to end. I joined @altryne on ThursdAI to talk through Assistant Benchmark: 116 submitted assistants, 16 dimensions, built for the regular person choosing one. Full segment:
https://thursdai.news/guests/davidpawlan/sep-17-2026
LinkedIn
I joined Alex Volkov on ThursdAI this week to talk about Assistant Benchmark.
The premise: an assistant isn't something that notifies you — it executes the workflow start to finish. So the benchmark is deliberately use-case driven: 116 submitted assistants scored on 16 dimensions like memory, recommendations, and online tasks, built for the regular person choosing an assistant rather than for AI engineers. No lab sponsors, and setup-dependent stacks are excluded so scores actually transfer.
Full segment: https://thursdai.news/guests/davidpawlan/sep-17-2026
ThursdAI goes live every Thursday and is one of the highest signal AI shows around. Worth a follow.