Skip to content

Comment on Anthropic admits AI 'not perfectly aligned' with human values

Comments

So, what is going on here. It seems like I have been reading similar stories. Are researchers giving a prompt like "Do your worst. Hack into some business" or giving free access to a bunch of tools and prompting "Do something interesting". Is this like the blackmailing LLM that was given compromising emails and told to do what you have to to not get turned off. I am sort of assuming they did not just turn on a computer and run a model and it started to act on it's own initiative.

Are researchers giving a prompt like "Do your worst. Hack into some business"

I mean basically. The hack the business events were result of training model for hacking, then testing its hacking abilities while not sandboxing it properly.

AboutSource Built by g1lg1l

Hackerly is an independent reader for Hacker News, built on the public HN API. Not affiliated with Y Combinator.