Failure Briefs
Postmortem-style analysis of AI system failures, fragility, and production risk.
Guides and explainers
Detailed guides and practical technical analysis.
Latest analysis
Recent research, benchmark reviews and technical updates.
Practical tools
Templates for budgets and project planning
BoredTools offers practical spreadsheets for budgets, freelance work and small projects.
ToolPrivacyBench Exposes Purpose Drift
ToolPrivacyBench turns agent privacy into a tool-call audit: when a workflow succeeds, did each tool receive only the private facts it needed? The June...
Agent Security Needs Owners, Not More Threat Lists
Security and Privacy in Agentic AI landed on arXiv on 7 July 2026 with a useful warning: agentic risk is now too operational for taxonomy work alone...
Agent Marketplaces Need Abuse Screens, Not Escrow
Agent marketplaces are no longer only payment demos; they are becoming tool surfaces where software can hire people. A February 2026 empirical study of...
Agent Sandboxes Need Egress Budgets, Not Trust Prompts
The live risk in agent security is shifting from "will the model say something unsafe?" to "what can the harness actually touch after the model decides?"...
RAG Cost Attacks Turn Retrieval Into a Budget Risk
A June 2026 paper on retrieval-augmented inference cost attacks reports a failure mode that many RAG teams are not testing: poisoned external documents...
Agent Data Injection Needs Trust Boundaries, Not Prompt Filters
Agent Data Injection argues that attackers can disguise malicious payloads as data the agent treats as trusted metadata, tool output, resource...
Agent Accountability Breaks When the Audit Trail Is Just a Trace
The EU AI Act's Article 12 now says high-risk AI systems must automatically record events across the system lifetime. Microsoft, in parallel, is migrating...
The Accountability Gap When AI Agents Act
When an AI agent causes harm, who pays? Current law can't answer that clearly.
Agent Browsers Need Traffic Policy, Not Bot Blocks
Agentic browser traffic is no longer a rounding error in website operations. HUMAN Security's 2026 benchmark report says traffic from AI agents and...
Agent Memory Needs Quarantine, Not Recall
Persistent memory is moving from chat convenience into personal-agent infrastructure. The failure mode is not just forgetting. It is remembering the wrong...