Benchmark Watch

Evaluation notes, benchmark interpretation, leaderboard skepticism, and measurement failures.

Guides and explainers

Detailed guides and practical technical analysis.

No guides are published for this topic yet.

Practical tools

Templates for budgets and project planning

BoredTools offers practical spreadsheets for budgets, freelance work and small projects.

Browse templates Free budget tracker
Tool Agents Need Taint Boundaries, Not Bigger Scopes

Tool Agents Need Taint Boundaries, Not Bigger Scopes

APPA landed on arXiv on 27 July 2026 with a blunt result for tool-agent security: permission prompts are too late once poisoned or confidential data has...

4 min read
Chip-Design Agents Need Token ROI, Not Demo Flows

Chip-Design Agents Need Token ROI, Not Demo Flows

FluxBench landed on arXiv on 20 July 2026 with a useful warning for chip-design teams: two agents can start from the same foundation model and still...

5 min read
Terminal Agents Need Progress Curves, Not Victory Screens

Terminal Agents Need Progress Curves, Not Victory Screens

Long-Horizon-Terminal-Bench landed on arXiv in July 2026 with an awkward result for terminal-agent buyers: even the strongest tested model still failed...

4 min read
Million-Token Context Still Fails the Workload Test

Million-Token Context Still Fails the Workload Test

Anthropic reported on February 5, 2026 that Claude Opus 4.6 scored 76% on the 8-needle 1M-token MRCR v2 test while Claude Sonnet 4.5 scored 18.5% on the...

7 min read
Agent Evals Need Harnesses, Not More Scoreboards

Agent Evals Need Harnesses, Not More Scoreboards

AgentCompass first landed on arXiv on 15 July 2026 with a practical complaint: agent evaluation is fragmented, tightly coupled, and hard to reproduce...

3 min read
Coding Agent Benchmarks Hit the Generalization Wall

Coding Agent Benchmarks Hit the Generalization Wall

Scale's SWE-Bench Pro public leaderboard reports that top models scoring above 70% on SWE-Bench Verified fall to 23.3% for OpenAI GPT-5 and 23.1% for...

6 min read
Agent Test-Time Scaling Needs Reuse, Not More Rollouts

Agent Test-Time Scaling Needs Reuse, Not More Rollouts

General AgentBench reports that running agents for more interaction steps or more sampled trajectories did not reliably improve ten leading agents,...

3 min read
Multi-Agent Finance Workflows Need Cost Curves, Not More Agents

Multi-Agent Finance Workflows Need Cost Curves, Not More Agents

A March 2026 benchmark on financial-document processing makes the uncomfortable point: the most accurate multi-agent architecture was not the obvious...

4 min read
Self-Improving Agents Have an Evaluator Problem

Self-Improving Agents Have an Evaluator Problem

Anthropic's June 2026 update on recursive self-improvement is not a distant sci-fi warning. The company says its engineers now ship 8x as much code per...

3 min read
Coding Agents Need Trajectory Reviews, Not Pass Bits

Coding Agents Need Trajectory Reviews, Not Pass Bits

Most coding-agent benchmarks still compress a whole run into one bit: did the task pass? AgentLens argues that users experience the whole trajectory...

4 min read
Swarm Signal
0:00
0:00
Up Next

Queue is empty. Click "+ Queue" on any article to add it.