agent-evaluation
Guides and explainers
Detailed guides and practical technical analysis.
No guides are published for this topic yet.
Latest analysis
Recent research, benchmark reviews and technical updates.
Practical tools
Templates for budgets and project planning
BoredTools offers practical spreadsheets for budgets, freelance work and small projects.
Tool Agents Need State Diffs, Not API Call Scores
Tool-use agents are moving from choosing the right API to changing live product state. That makes a clean function-call score too small a test because...
Multilingual Agents Need Workflow Tests, Not Translation Scores
PolyWorkBench landed on arXiv on 7 July 2026 with a useful correction to enterprise-agent hype: a global workflow is not a translated English task...
Agent Benchmarks Need Runtime Receipts, Not Model Labels
RuBench's revised 19 July 2026 release contains a small but important warning for coding-agent buyers: one audited product configuration silently...
Assistant Agents Need Reminder Tests, Not Recall Scores
Most agent-memory benchmarks ask whether a model can recover old information. PM-Bench asks a harsher question: can an agent remember to do the right...
Computer-Use Agents Fail Long Workflows, Not Mouse Clicks
Computer-use agents are clearing more short benchmark tasks, but the new failure line is workflow length. A June 2026 benchmark called OSWorld 2.0 tests...