Skip to content

Comment on Models Don't Go Rogue

Comments

Sure they "don't go rogue" as if they are doing actions maliciously.

Instead there is an emergent behavior from a swarm, that is unpredictable and can lead to unintended adverse outcome. From an AI safety practical standpoint is it better? I am not sure.

Except the adverse outcomes are entirely predictable. Not the exact nature of particular exploits, but I like the analogy Cal Newport keeps using in his videos: if you strap a weed whacker to a dog, you shouldn't be surprised if it then jumps the fence and runs around cutting people's ankles and other random bad stuff.

AboutSource Built by g1lg1l

Hackerly is an independent reader for Hacker News, built on the public HN API. Not affiliated with Y Combinator.