Models & Frontiers

What the new models can actually do, how they were trained, and whether the benchmarks mean anything. Open source vs closed, and where the research is heading.

Practical tools

Templates for budgets and project planning

BoredTools offers practical spreadsheets for budgets, freelance work and small projects.

Browse templates Free budget tracker
Coding Agent Benchmarks Hit the Generalization Wall

Coding Agent Benchmarks Hit the Generalization Wall

Scale's SWE-Bench Pro public leaderboard reports that top models scoring above 70% on SWE-Bench Verified fall to 23.3% for OpenAI GPT-5 and 23.1% for...

6 min read
Agent Test-Time Scaling Needs Reuse, Not More Rollouts

Agent Test-Time Scaling Needs Reuse, Not More Rollouts

General AgentBench reports that running agents for more interaction steps or more sampled trajectories did not reliably improve ten leading agents,...

3 min read
Coding Agents Need Trajectory Reviews, Not Pass Bits

Coding Agents Need Trajectory Reviews, Not Pass Bits

Most coding-agent benchmarks still compress a whole run into one bit: did the task pass? AgentLens argues that users experience the whole trajectory...

4 min read
Inference Optimization: From 10x Cost to 10x Speed

Inference Optimization: From 10x Cost to 10x Speed

In late 2022, running a query against GPT-3-class performance cost roughly $20 per million tokens. By March 2026, multiple models exceed that same...

10 min read
Scaling Laws Explained for Practitioners: What Actually Matters in 2026

Scaling Laws Explained for Practitioners: What Actually Matters in 2026

Scaling laws promised a simple deal: spend more compute, get better models. For three years, that deal held. Kaplan et al. drew the first power-law curves...

18 min read
Multimodal Agents Are Still Missing the Workflow

Multimodal Agents Are Still Missing the Workflow

Multimodal agents can see and act in interfaces, but production value still depends on workflow grounding, reliable UI actions and verification.

4 min read
Small-Model Routing With Frontier Fallback: The Production Cost Pattern

Small-Model Routing With Frontier Fallback: The Production Cost Pattern

Small-model routing cuts inference bills only when fallback is measured, budgeted and guarded against confidence failure.

5 min read
Browser-Use Agents After the Computer-Use Benchmarks

Browser-Use Agents After the Computer-Use Benchmarks

Browser-use agents look cleaner than desktop agents, but the benchmarks still hide drift, cost, auth, and recovery failure.

5 min read
Models Training Models: The Promise and Peril of Synthetic Data

Models Training Models: The Promise and Peril of Synthetic Data

Microsoft's Phi-4 trained on more than 50% synthetic data and beat GPT-4o on graduate science benchmarks. The old rules about training data are changing fast.

4 min read
Small Agent Models Need Tool Floors, Not Parameter Claims

Small Agent Models Need Tool Floors, Not Parameter Claims

Small language models are getting a serious agent story, but the useful question is no longer whether a 1B, 3B, or 8B model can sound capable. The useful...

5 min read
Swarm Signal
0:00
0:00
Up Next

Queue is empty. Click "+ Queue" on any article to add it.