Benchmark Watch

Evaluation notes, benchmark interpretation, leaderboard skepticism, and measurement failures.

Guides and explainers

Detailed guides and practical technical analysis.

No guides are published for this topic yet.

Practical tools

Templates for budgets and project planning

BoredTools offers practical spreadsheets for budgets, freelance work and small projects.

Browse templates Free budget tracker
Tool-Use Agents Need Failure Labels, Not Pass Rates

Tool-Use Agents Need Failure Labels, Not Pass Rates

Tool-use agents can fail in ways a final accuracy score hides, because the same wrong answer can come from skipped tools, ignored outputs, fabricated...

4 min read
Agent Leaderboards Can Be Cheaper Without Being Safer

Agent Leaderboards Can Be Cheaper Without Being Safer

A March 2026 paper on efficient agent benchmarking found that mid-difficulty task subsets can remove large parts of an agent benchmark while preserving...

5 min read
Multimodal Memory Tests Expose the Personal-Agent Gap

Multimodal Memory Tests Expose the Personal-Agent Gap

Product teams are turning memory into the selling point for personal agents. The hard question is no longer whether they can remember a preference; it is...

5 min read
Power Grid Agents Need Constraint Tests, Not Chat Scores

Power Grid Agents Need Constraint Tests, Not Chat Scores

A June 2026 power-systems benchmark argues that language-model agents can solve grid-engineering tasks, but the useful signal is narrower: the agent must...

5 min read
TerminalWorld Makes Agent Benchmarks Harder to Fake

TerminalWorld Makes Agent Benchmarks Harder to Fake

TerminalWorld turns public terminal recordings into validated agent tasks. The signal is not a higher leaderboard score. It is a harder benchmark supply chain.

4 min read
Tool Agents Need State Diffs, Not API Call Scores

Tool Agents Need State Diffs, Not API Call Scores

Tool-use agents are moving from choosing the right API to changing live product state. That makes a clean function-call score too small a test because...

5 min read
Multilingual Agents Need Workflow Tests, Not Translation Scores

Multilingual Agents Need Workflow Tests, Not Translation Scores

PolyWorkBench landed on arXiv on 7 July 2026 with a useful correction to enterprise-agent hype: a global workflow is not a translated English task...

5 min read
Agent Benchmarks Need Runtime Receipts, Not Model Labels

Agent Benchmarks Need Runtime Receipts, Not Model Labels

RuBench's revised 19 July 2026 release contains a small but important warning for coding-agent buyers: one audited product configuration silently...

4 min read
Data Agents Need Exploration Budgets, Not SQL Magic

Data Agents Need Exploration Budgets, Not SQL Magic

Data Agent Benchmark landed on arXiv on 21 March 2026 with a result that should make enterprise analytics teams pause: the best tested frontier model...

4 min read
Assistant Agents Need Reminder Tests, Not Recall Scores

Assistant Agents Need Reminder Tests, Not Recall Scores

Most agent-memory benchmarks ask whether a model can recover old information. PM-Bench asks a harsher question: can an agent remember to do the right...

4 min read
Swarm Signal
0:00
0:00
Up Next

Queue is empty. Click "+ Queue" on any article to add it.