September 21, 2026
In the Stage 2 report we shipped 20 feature tickets on Fizzy and promised to explore benchmarking the agents all on max-effort. Now we have run it: every model on...
September 9, 2026
Today we’re expanding Agents on Rails with Stage 2, a new set of benchmark tasks that exercise a model’s ability to deliver real-world features, more in line with how a...
September 2, 2026
Claude Fable 5.1 dropped yesterday. See how it handles real Rails tasks, and where it landed on the leaderboard today. We also finally have a name for the stealth model...
August 24, 2026
Another week, another step for Agents on Rails. This one is a big one: lemans, the harness behind every number we’ve published, is now open source. We also ran four...
August 17, 2026
Last week we launched Agents on Rails and published the first benchmark report. The response was immediate: suggestions, questions, model requests, and more than a few “but have you tried…”...
August 13, 2026
TL;DR We ran 8 models against 21 atomic Rails tasks, 3 runs each. Every task runs against Writebook: a bug report, a security finding, a feature request, each written the...
August 12, 2026
Today we’re sharing the first results of Agents on Rails, a new, ongoing initiative to measure how well today’s leading agentic coding tools (both frontier and open-weight) actually perform on...