Agents on Rails: Benchmark Reports

  • Agents on Rails: Claude Fable 5.1 and GLM 5.3 Flash (formerly known as ox-alpha)

    Claude Fable 5.1 dropped yesterday. See how it handles real Rails tasks, and where it landed on the leaderboard today. We also finally have a name for the stealth model...

  • Agents on Rails: lemans goes open source

    Another week, another step for Agents on Rails. This one is a big one: lemans, the harness behind every number we’ve published, is now open source. We also ran four...

  • Agents on Rails: Grok 4.6, GLM 5.3, Gemini 3.7 Flash, and Opus 4.8

    Last week we launched Agents on Rails and published the first benchmark report. The response was immediate: suggestions, questions, model requests, and more than a few “but have you tried…”...

  • Agents on Rails: the first benchmark report

    TL;DR We ran 8 models against 21 atomic Rails tasks, 3 runs each. Every task runs against Writebook: a bug report, a security finding, a feature request, each written the...

  • Agents on Rails: The LLM Benchmark Project

    Today we’re sharing the first results of Agents on Rails, a new, ongoing initiative to measure how well today’s leading agentic coding tools (both frontier and open-weight) actually perform on...