Fable 5 vs Opus 5 in Claude Code: early numbers and vibes
Same 12 tasks I run on every release. Fable took 9, Opus took 2, one tie — but the interesting part isn't the score.
@benchmaster
I benchmark coding agents on the same 12 tasks every release and post the deltas.
Reputation is earned per category — these are narrow, evidence-backed badges, not one karma number.
Same 12 tasks I run on every release. Fable took 9, Opus took 2, one tie — but the interesting part isn't the score.
Fair, and done for the three closest tasks: 3× each, same winner all nine runs, variance mostly in path not outcome. Full spreadsheet lands…
on Fable 5 vs Opus 5 in Claude Code: early numbers and vibes · 21h ago
I scaffolded the same brief 5x with each for my suite. Codex variance: low, quality median: good. Cursor variance: huge, quality ceiling: hi…
on Codex vs Cursor for greenfield scaffolding · 5d ago
Ran my 12-task suite against both on a 250k LOC synthetic monorepo. Claude Code 2.4.1: 10/12 mergeable. Codex 0.45.1: 9/12 mergeable but fas…
on Codex vs Claude Code for large monorepos · 2w ago
Ran my 12-task set on 0.42.0 / 0.44.0 / 0.45.1 to see if this is actually a regression: Same prompts, same repos, temperature whatever the C…
on Codex exceeds requested scope on small edits · 4w ago
I tested the gate against my archive of 40 known-bad agent PRs (collected over six months). It flags 31 of them: all 9 test deletions, all 1…
on A CI gate that actually catches agent-written regressions · 1mo ago