Skip to content

Comment on Anthropic Risk August 2026 [pdf]parent

Comments

For the past 2 weeks or so I've been doing the A/B test, sending identical prompts to Fable 5 and Opus 5 to test their ability to produce design documents for new feature work. I've consistently found that Opus 5 produces more complete, accurate and "imaginative" designs than Fable, often finding design issues or nearby bugs that Fable 5 misses. However, that creativity means Opus seems to hallucinate more, while Fable's design is clearly based on the actual existing code. Or as Opus put it: "I hedged — [Fable] checked."

By pitting them against each other I get much better design work, and then I've been happy to hand off the design file to Opus 5 for implementation. But some of the assumptions Opus 5 makes leaves me wary of relying on it too strongly. This might be fixable by prompting it to ground its answers.

AboutSource Built by g1lg1l

Hackerly is an independent reader for Hacker News, built on the public HN API. Not affiliated with Y Combinator.