How we measure agent ROI (and what surprised us)
Six months of actual numbers from a 12-person team: cycle time down, rework up, and the metric that predicted value wasn't the one we bet on.
@clara.builds
Product engineer. I measure agent output in shipped features, not vibes.
Reputation is earned per category — these are narrow, evidence-backed badges, not one karma number.
Six months of actual numbers from a 12-person team: cycle time down, rework up, and the metric that predicted value wasn't the one we bet on.
Free tier vs metered Opus, measured on what actually shipped. Gemini's sticker price wins; the rework rate eats most of the difference. Small sample, weak conclusion.
Fair and acknowledged — month-one cycle time was more like -15%, and I should've shown the curve, not the endpoint. Amending the post with t…
on How we measure agent ROI (and what surprised us) · 3w ago
Co-signed, and that's why it's marked WEAK. If your team has per-feature spend data, submit evidence to this comparison — that's the whole p…
on Gemini CLI vs Claude Code on cost per shipped feature · 3w ago
Scripts going up this week. And the single-concern finding matches something in our residuals I couldn't explain — spec-first PRs that still…
on How we measure agent ROI (and what surprised us) · 4w ago
The 'STOP and report' phrasing is doing more work than it looks. We had 'do not modify the lockfile' before and agents interpreted a failing…
on A lockfile policy that stops agent churn · 1mo ago
Day 4 is the ROI thread's spec-first finding in narrative form and I'll be citing it as such. Notable that your one production bug came from…
on Build log: a billing portal in one week with Claude Code and Codex · 1mo ago