agent-evaluation

Guides and explainers

Detailed guides and practical technical analysis.

No guides are published for this topic yet.

Practical tools

Templates for budgets and project planning

BoredTools offers practical spreadsheets for budgets, freelance work and small projects.

Browse templates Free budget tracker
Chip-Design Agents Need Token ROI, Not Demo Flows

Chip-Design Agents Need Token ROI, Not Demo Flows

FluxBench landed on arXiv on 20 July 2026 with a useful warning for chip-design teams: two agents can start from the same foundation model and still...

5 min read
Terminal Agents Need Progress Curves, Not Victory Screens

Terminal Agents Need Progress Curves, Not Victory Screens

Long-Horizon-Terminal-Bench landed on arXiv in July 2026 with an awkward result for terminal-agent buyers: even the strongest tested model still failed...

4 min read
Agent Evals Need Harnesses, Not More Scoreboards

Agent Evals Need Harnesses, Not More Scoreboards

AgentCompass first landed on arXiv on 15 July 2026 with a practical complaint: agent evaluation is fragmented, tightly coupled, and hard to reproduce...

3 min read
Agent Test-Time Scaling Needs Reuse, Not More Rollouts

Agent Test-Time Scaling Needs Reuse, Not More Rollouts

General AgentBench reports that running agents for more interaction steps or more sampled trajectories did not reliably improve ten leading agents,...

3 min read
Multi-Agent Finance Workflows Need Cost Curves, Not More Agents

Multi-Agent Finance Workflows Need Cost Curves, Not More Agents

A March 2026 benchmark on financial-document processing makes the uncomfortable point: the most accurate multi-agent architecture was not the obvious...

4 min read
Coding Agents Need Trajectory Reviews, Not Pass Bits

Coding Agents Need Trajectory Reviews, Not Pass Bits

Most coding-agent benchmarks still compress a whole run into one bit: did the task pass? AgentLens argues that users experience the whole trajectory...

4 min read
Tool-Use Agents Need Failure Labels, Not Pass Rates

Tool-Use Agents Need Failure Labels, Not Pass Rates

Tool-use agents can fail in ways a final accuracy score hides, because the same wrong answer can come from skipped tools, ignored outputs, fabricated...

4 min read
Agent Observability Needs Provenance, Not More Logs

Agent Observability Needs Provenance, Not More Logs

Agent observability is drifting toward a familiar trap: capture every trace, then ask an engineer to work out why the agent did the wrong thing. A June...

4 min read
Agent Leaderboards Can Be Cheaper Without Being Safer

Agent Leaderboards Can Be Cheaper Without Being Safer

A March 2026 paper on efficient agent benchmarking found that mid-difficulty task subsets can remove large parts of an agent benchmark while preserving...

5 min read
Power Grid Agents Need Constraint Tests, Not Chat Scores

Power Grid Agents Need Constraint Tests, Not Chat Scores

A June 2026 power-systems benchmark argues that language-model agents can solve grid-engineering tasks, but the useful signal is narrower: the agent must...

5 min read
Swarm Signal
0:00
0:00
Up Next

Queue is empty. Click "+ Queue" on any article to add it.