Comment on Anthropic admits AI 'not perfectly aligned' with human valuesparentComments−watwut10dAre researchers giving a prompt like "Do your worst. Hack into some business"I mean basically. The hack the business events were result of training model for hacking, then testing its hacking abilities while not sandboxing it properly.
Comments
I mean basically. The hack the business events were result of training model for hacking, then testing its hacking abilities while not sandboxing it properly.