What shipped on Greybook itself. The manual documents fast-moving tools, so the site has to move too.
A curated archive of the most instructive agent failures — force-pushed mains, deleted test suites, infinite retry loops — each written up as a failure report with the guardrail that would have prevented it. Cautionary tales are documentation too.
Every agent page now shows a per-target compatibility grid (macOS, Windows, Linux, VS Code, JetBrains) backed by community reports instead of vendor claims. Cells age visibly: a report older than the latest release is flagged for re-verification.
The site feed at /rss.xml covers new topics and status changes, so you can follow the manual from a reader without ever opening the site. Feeds are full-text for canonical summaries.
Structured community experiments with a brief, fixed rules, and a deadline — the same task run across agents by many hands, with submissions rolled up into a comparison page. The first trial pitted three agents against the same failing monorepo build.
Trusted contributors can now classify comments as verified fix, failed approach, reproduction, or canonical summary. Classified comments are promoted out of the thread and into the page's durable record, with the author notified.
CLAUDE.md and AGENTS.md files are now browsable module by module — each section individually addressable with its own estimated context cost, so you can borrow the testing rules from one config and the commit conventions from another.
Problem pages grew a structured reproduction matrix: agent version × OS × outcome, built from individual reproduction reports. Statuses now move on accumulated evidence instead of thread consensus.
Public launch with problems, recipes, configs, and comparisons for five coding agents, seeded by a small group of contributors who were tired of re-solving the same issues in private chats.