Signals
Production-oriented research signals and interpretation for AI systems builders.
Guides and explainers
Detailed guides and practical technical analysis.
Latest analysis
Recent research, benchmark reviews and technical updates.
Practical tools
Templates for budgets and project planning
BoredTools offers practical spreadsheets for budgets, freelance work and small projects.
Agent Memory Fails on Relationships, Not Recall
New June 2026 memory benchmarks show why long-running agents fail when facts conflict, evolve, or depend on hidden relationships.
SMAC-Talk Shows Agent Chat Is Not Coordination
SMAC-Talk adds natural-language communication and deception to StarCraft-style multi-agent evaluation. The result is a useful warning: agent chat can expose coordination failure as easily as it fixes
Agent Benchmarking Doesn't Need Every Task
Efficient agent benchmarking points to a cheaper way to compare agents: run the tasks that still separate systems, not every task in the suite.
Tool Agents Need State Diffs, Not API Call Scores
Tool-use agents are moving from choosing the right API to changing live product state. That makes a clean function-call score too small a test because...
Healthcare AI Agents Move Beyond Drug Discovery
Healthcare AI agents are moving into admin, triage and prior-authorisation workflows. The real gate is safety, evidence and accountable handoff.
Multilingual Agents Need Workflow Tests, Not Translation Scores
PolyWorkBench landed on arXiv on 7 July 2026 with a useful correction to enterprise-agent hype: a global workflow is not a translated English task...
Self-Improving Agents Need Hard Boundaries
Self-improving agents can rewrite code, prompts and memory. Production teams need rollback, approval gates and evaluator change control.
Agent Benchmarks Need Runtime Receipts, Not Model Labels
RuBench's revised 19 July 2026 release contains a small but important warning for coding-agent buyers: one audited product configuration silently...
Multimodal Agents Are Still Missing the Workflow
Multimodal agents can see and act in interfaces, but production value still depends on workflow grounding, reliable UI actions and verification.
Agent Accountability Is Becoming Runtime Infrastructure
Agent accountability is becoming runtime infrastructure: identity, delegated authority, trace logs, approvals and incident reconstruction.