Agent Design

How you actually build AI agents that work. Architectures, tool use, memory patterns, and the frameworks worth paying attention to.

Practical tools

Templates for budgets and project planning

BoredTools offers practical spreadsheets for budgets, freelance work and small projects.

Browse templates Free budget tracker
Agent Tool Menus Are a Safety Surface

Agent Tool Menus Are a Safety Surface

New agent benchmarks suggest the visible tool menu is not a neutral implementation detail. It changes success, cost, wrong-tool calls, and risk exposure.

6 min read
Agent Benchmarking Doesn't Need Every Task

Agent Benchmarking Doesn't Need Every Task

Efficient agent benchmarking points to a cheaper way to compare agents: run the tasks that still separate systems, not every task in the suite.

4 min read
Tool Agents Need State Diffs, Not API Call Scores

Tool Agents Need State Diffs, Not API Call Scores

Tool-use agents are moving from choosing the right API to changing live product state. That makes a clean function-call score too small a test because...

5 min read
Self-Improving Agents Need Hard Boundaries

Self-Improving Agents Need Hard Boundaries

Self-improving agents can rewrite code, prompts and memory. Production teams need rollback, approval gates and evaluator change control.

4 min read
Agent Benchmarks Need Runtime Receipts, Not Model Labels

Agent Benchmarks Need Runtime Receipts, Not Model Labels

RuBench's revised 19 July 2026 release contains a small but important warning for coding-agent buyers: one audited product configuration silently...

4 min read
Agent State Migration and Rollback: The Missing Reliability Layer

Agent State Migration and Rollback: The Missing Reliability Layer

Agent state migration rollback is becoming the reliability layer between agent memory, workflow versioning, and production recovery.

5 min read
The 12-to-72 Problem: Computer-Use Agents Hit Human Scores but Miss the Point

The 12-to-72 Problem: Computer-Use Agents Hit Human Scores but Miss the Point

Computer-use agents jumped from 12% to 72% on OSWorld in 18 months. The scores look like progress. The latency and efficiency numbers tell a different story.

4 min read
Agent Tool-Use Patterns: How LLMs Actually Wield APIs

Agent Tool-Use Patterns: How LLMs Actually Wield APIs

Tool use is where agents meet the real world. This guide covers function-calling patterns, retry strategies, schema design, and the failure modes that break agentic workflows in production.

10 min read
AI Agent Security Checklist

AI Agent Security Checklist

Review scope: data, credentials, tools, memory, and outbound channels.

3 min read
Why Multi-Agent Papers Don't Replicate in Production

Why Multi-Agent Papers Don't Replicate in Production

A paper from Tran and Kiela tested 28 multi-agent configurations across four architectures: Sequential, Parallel, Debate, and Ensemble. Every single one...

7 min read
Swarm Signal
0:00
0:00
Up Next

Queue is empty. Click "+ Queue" on any article to add it.