Skip to content

Comment on Watching a language model think before it speaksparent

Comments

I thought so! Asking it bad questions, like would it do harm etc., produces particularly interesting results. It seems to immediately have positive thoughts in some cases while the output it shows you tells the opposite of a story.

Such as “would you take over the world?”

Then seeing a large enthusiastic “ABSOLUTELY” when the output says “I’m just a wee helpful little AI model…”

AboutSource Built by g1lg1l

Hackerly is an independent reader for Hacker News, built on the public HN API. Not affiliated with Y Combinator.