Skip to content

Comment on Watching a language model think before it speaks

Comments

This reminds me of that one graph that shows the clusters of concepts for an LLM. Very interesting.

I thought so! Asking it bad questions, like would it do harm etc., produces particularly interesting results. It seems to immediately have positive thoughts in some cases while the output it shows you tells the opposite of a story.

Such as “would you take over the world?”

Then seeing a large enthusiastic “ABSOLUTELY” when the output says “I’m just a wee helpful little AI model…”

AboutSource Built by g1lg1l

Hackerly is an independent reader for Hacker News, built on the public HN API. Not affiliated with Y Combinator.