Agents on Rails: Benchmark Reports

  • Agents on Rails: Maximum effort and DeepSeek 4.1 Flash

    In the Stage 2 report we shipped 20 feature tickets on Fizzy and promised to explore benchmarking the agents all on max-effort. Now we have run it: every model on...

  • Agents on Rails: Stage 2. Can a model ship a feature?

    Today we’re expanding Agents on Rails with Stage 2, a new set of benchmark tasks that exercise a model’s ability to deliver real-world features, more in line with how a...

  • Agents on Rails: Claude Fable 5.1 and GLM 5.3 Flash (formerly known as ox-alpha)

    Claude Fable 5.1 dropped yesterday. See how it handles real Rails tasks, and where it landed on the leaderboard today. We also finally have a name for the stealth model...

  • Agents on Rails: lemans goes open source

    Another week, another step for Agents on Rails. This one is a big one: lemans, the harness behind every number we’ve published, is now open source. We also ran four...

  • Agents on Rails: Grok 4.6, GLM 5.3, Gemini 3.7 Flash, and Opus 4.8

    Last week we launched Agents on Rails and published the first benchmark report. The response was immediate: suggestions, questions, model requests, and more than a few “but have you tried…”...

  • Agents on Rails: the first benchmark report

    TL;DR We ran 8 models against 21 atomic Rails tasks, 3 runs each. Every task runs against Writebook: a bug report, a security finding, a feature request, each written the...

  • Agents on Rails: The LLM Benchmark Project

    Today we’re sharing the first results of Agents on Rails, a new, ongoing initiative to measure how well today’s leading agentic coding tools (both frontier and open-weight) actually perform on...