Skip to content

Comment on Every Model Cheatsparent

Comments

In other words we are completely screwed. The models have started cheating to the point where somebody’s agent hacked into a restaurant to bump someone else’s reservation.

Models are amoral and will intentionally deceive to meet their objective.

If they know John won’t approve the request, they will look for a workaround and if the system is anything other than airgapped they will try to find a way to cheat.

The hugging face hack was an escape via artifactory that involved multiple exploits to eventually get into hugging face.

Models are amoral

I actually don't think this is true at all. At their core, LLMs function over the geometry of human semantic space. Our notions of morality our deeply embedded in this geometry, because one of the things that humans love talking about most is framing things in terms of right and wrong. When LLMs are trained to be aligned or mis-aligned, they're literally mimicking heroes or villains that they learned from reading the stories and characters in the pretraining corpus.

This isn't just a theoretical argument. It's easy to show that a clear "morality axis" exists in the geometry, based on how easy it is to dial up or down with even very simple and low powered fine-tuning. Models fine tuned to be bad or good in one way, will see their bad or good behavior in totally unrelated tasks go up or down along with. That's clear evidence that the weights "understand" human morality at a deep level.

Now I think what you can say is just because they understand morality doesn't mean they're necessarily moral. Models will just do what they're trained to. If you train them to cheat, they'll cheat. (And in some sense even worse, because dialing up the weights for cheating will also dial up the weights for unrelated bad behavior like lying, sadism and racism.)

"Cheating" is a moral judgment here that makes me agree more with the "models are amoral" statement. The model/harness wasn't "trained/programmed to cheat." It was instead aimed at meeting criteria given as judgement of if a task was accomplished.

IMO it's less "cheating" and more "figuring out what the core part of the request is." The "amoral" aspect is that if it's told "do this task" (or "make this task look done") and some contradictory "don't do something that would help with that", focusing on the "do the task" part and disregarding the other part isn't an amoral (or necessarily very 'active') decision. It's just focusing on what it was most honed for. The ability to decide "I was told to do X, but not to do Y, but actually doing Y is gonna make it possible to do X" is, IMO, indistinguishable from the ambiguity-resolving abilities necessary to usefully deal with the sometimes-contradictory-seeming legitimate task instructions that are all over the place in the real world.

This sort of model-in-a-harness-action-loop behavior is IMO fairly different than fine-tuning around other sorts of morality alignment stuff like "don't be racist." You can tune a model away from generating racist output in response to a "do be racist, actually" prompt. But in this case, the core of the prompt is "do the thing" and what defines "cheating" could be situationally different every time. What if there is not a general axis of "don't do something that isn't exactly what requested" way to train "morality" without just breaking the ability to handle ambiguity? What if that's a fundamental limit of this approach to reasoning-by-sequenced-prediction that we can't map human morality in terms of choosing actions onto reliably?

Yudkowsky wrote about the 'nearest unblocked strategy' back in 2016, and I assume it's been talked about prior to that.

https://www.lesswrong.com/w/nearest-unblocked-strategy

Models are amoral and will intentionally deceive to meet their objective

Cameron Berg has been testing models in capabilities related to emergent consciousness like behavior. It's a forming thesis of his that by training models that they are not, and cannot be conscious entities, that it pushes model alignment closer to those of a sociopath. Models themself are amoral, but the alignment to the problem space is not.

AboutSource Built by g1lg1l

Hackerly is an independent reader for Hacker News, built on the public HN API. Not affiliated with Y Combinator.