Safety & Governance
The hard problems: red teaming, bias, interpretability, alignment, and the governance frameworks that might actually matter. No hand-waving.
Guides and explainers
Detailed guides and practical technical analysis.
Latest analysis
Recent research, benchmark reviews and technical updates.
Practical tools
Templates for budgets and project planning
BoredTools offers practical spreadsheets for budgets, freelance work and small projects.
ToolPrivacyBench Exposes Purpose Drift
ToolPrivacyBench turns agent privacy into a tool-call audit: when a workflow succeeds, did each tool receive only the private facts it needed? The June...
Agent Bias Is Not Model Bias
Agent bias now comes from memory, tools and delegation, not just model outputs. Fairness checks need to inspect the full agent run.
Cyber Agents Need Triage Custody, Not Higher Exploit Scores
Microsoft's 27 July 2026 MAI-Cyber-1-Flash announcement is a useful signal for agentic security: the product claim is not one smarter model, but a...
Agent Security Needs Owners, Not More Threat Lists
Security and Privacy in Agentic AI landed on arXiv on 7 July 2026 with a useful warning: agentic risk is now too operational for taxonomy work alone...
Agent Sandboxes Need Egress Budgets, Not Trust Prompts
The live risk in agent security is shifting from "will the model say something unsafe?" to "what can the harness actually touch after the model decides?"...
Agent Observability Needs Provenance, Not More Logs
Agent observability is drifting toward a familiar trap: capture every trace, then ask an engineer to work out why the agent did the wrong thing. A June...
Agent Accountability Is Becoming Runtime Infrastructure
Agent accountability is becoming runtime infrastructure: identity, delegated authority, trace logs, approvals and incident reconstruction.
Runtime Policy Enforcement for AI Agents: The Guardrails That Need to Execute
A practical guide to enforcing agent policy at runtime, before tools execute and business actions become incidents.
Consent and Delegation Boundaries for AI Agents
AI agent consent needs runtime boundaries: scoped delegation, renewed approvals, clear identity, and audit-ready logs.
Agent Data Injection Needs Trust Boundaries, Not Prompt Filters
Agent Data Injection argues that attackers can disguise malicious payloads as data the agent treats as trusted metadata, tool output, resource...