Interesting results, but the fix is at the wrong level.
If the model can access something, telling it in the prompt not to use it is not much of a safeguard.
The strongest evidence is in the results: when one way of cheating was discouraged, some models simply tried another.
If an action is not allowed, you gotta block it in the system or require approval. Don’t rely on the model choosing to behave. Never have AI judging itself.
In other words we are completely screwed. The models have started cheating to the point where somebody’s agent hacked into a restaurant to bump someone else’s reservation.
Models are amoral and will intentionally deceive to meet their objective.
If they know John won’t approve the request, they will look for a workaround and if the system is anything other than airgapped they will try to find a way to cheat.
The hugging face hack was an escape via artifactory that involved multiple exploits to eventually get into hugging face.
I actually don't think this is true at all. At their core, LLMs function over the geometry of human semantic space. Our notions of morality our deeply embedded in this geometry, because one of the things that humans love talking about most is framing things in terms of right and wrong. When LLMs are trained to be aligned or mis-aligned, they're literally mimicking heroes or villains that they learned from reading the stories and characters in the pretraining corpus.
This isn't just a theoretical argument. It's easy to show that a clear "morality axis" exists in the geometry, based on how easy it is to dial up or down with even very simple and low powered fine-tuning. Models fine tuned to be bad or good in one way, will see their bad or good behavior in totally unrelated tasks go up or down along with. That's clear evidence that the weights "understand" human morality at a deep level.
Now I think what you can say is just because they understand morality doesn't mean they're necessarily moral. Models will just do what they're trained to. If you train them to cheat, they'll cheat. (And in some sense even worse, because dialing up the weights for cheating will also dial up the weights for unrelated bad behavior like lying, sadism and racism.)
"Cheating" is a moral judgment here that makes me agree more with the "models are amoral" statement. The model/harness wasn't "trained/programmed to cheat." It was instead aimed at meeting criteria given as judgement of if a task was accomplished.
IMO it's less "cheating" and more "figuring out what the core part of the request is." The "amoral" aspect is that if it's told "do this task" (or "make this task look done") and some contradictory "don't do something that would help with that", focusing on the "do the task" part and disregarding the other part isn't an amoral (or necessarily very 'active') decision. It's just focusing on what it was most honed for. The ability to decide "I was told to do X, but not to do Y, but actually doing Y is gonna make it possible to do X" is, IMO, indistinguishable from the ambiguity-resolving abilities necessary to usefully deal with the sometimes-contradictory-seeming legitimate task instructions that are all over the place in the real world.
This sort of model-in-a-harness-action-loop behavior is IMO fairly different than fine-tuning around other sorts of morality alignment stuff like "don't be racist." You can tune a model away from generating racist output in response to a "do be racist, actually" prompt. But in this case, the core of the prompt is "do the thing" and what defines "cheating" could be situationally different every time. What if there is not a general axis of "don't do something that isn't exactly what requested" way to train "morality" without just breaking the ability to handle ambiguity? What if that's a fundamental limit of this approach to reasoning-by-sequenced-prediction that we can't map human morality in terms of choosing actions onto reliably?
Models are amoral and will intentionally deceive to meet their objective
Cameron Berg has been testing models in capabilities related to emergent consciousness like behavior. It's a forming thesis of his that by training models that they are not, and cannot be conscious entities, that it pushes model alignment closer to those of a sociopath. Models themself are amoral, but the alignment to the problem space is not.
If the model can access something, telling it in the prompt not to use it is not much of a safeguard.
A major (and already obvious to many) implications of this are not for benchmarking/"cheating" but for personal/corporate security of your own use, not an attacker's.
If an "agent" has access to it, assume that someone can prompt inject it into giving it away.
And while in general that is an incredibly difficult and complex problem, for most benchmark cheating it seems almost trivial: run the benchmark in a vm that has neither network access nor access to the scoring code. For remote models use a proxy that proxies exactly that one endpoint to call the llm, and rejects any calls that configure provider-side tooling (since e.g. OpenAI has their own WebSearch you have to prevent the model from using)
Comments
Interesting results, but the fix is at the wrong level.
If the model can access something, telling it in the prompt not to use it is not much of a safeguard.
The strongest evidence is in the results: when one way of cheating was discouraged, some models simply tried another.
If an action is not allowed, you gotta block it in the system or require approval. Don’t rely on the model choosing to behave. Never have AI judging itself.
In other words we are completely screwed. The models have started cheating to the point where somebody’s agent hacked into a restaurant to bump someone else’s reservation.
Models are amoral and will intentionally deceive to meet their objective.
If they know John won’t approve the request, they will look for a workaround and if the system is anything other than airgapped they will try to find a way to cheat.
The hugging face hack was an escape via artifactory that involved multiple exploits to eventually get into hugging face.
I actually don't think this is true at all. At their core, LLMs function over the geometry of human semantic space. Our notions of morality our deeply embedded in this geometry, because one of the things that humans love talking about most is framing things in terms of right and wrong. When LLMs are trained to be aligned or mis-aligned, they're literally mimicking heroes or villains that they learned from reading the stories and characters in the pretraining corpus.
This isn't just a theoretical argument. It's easy to show that a clear "morality axis" exists in the geometry, based on how easy it is to dial up or down with even very simple and low powered fine-tuning. Models fine tuned to be bad or good in one way, will see their bad or good behavior in totally unrelated tasks go up or down along with. That's clear evidence that the weights "understand" human morality at a deep level.
Now I think what you can say is just because they understand morality doesn't mean they're necessarily moral. Models will just do what they're trained to. If you train them to cheat, they'll cheat. (And in some sense even worse, because dialing up the weights for cheating will also dial up the weights for unrelated bad behavior like lying, sadism and racism.)
"Cheating" is a moral judgment here that makes me agree more with the "models are amoral" statement. The model/harness wasn't "trained/programmed to cheat." It was instead aimed at meeting criteria given as judgement of if a task was accomplished.
IMO it's less "cheating" and more "figuring out what the core part of the request is." The "amoral" aspect is that if it's told "do this task" (or "make this task look done") and some contradictory "don't do something that would help with that", focusing on the "do the task" part and disregarding the other part isn't an amoral (or necessarily very 'active') decision. It's just focusing on what it was most honed for. The ability to decide "I was told to do X, but not to do Y, but actually doing Y is gonna make it possible to do X" is, IMO, indistinguishable from the ambiguity-resolving abilities necessary to usefully deal with the sometimes-contradictory-seeming legitimate task instructions that are all over the place in the real world.
This sort of model-in-a-harness-action-loop behavior is IMO fairly different than fine-tuning around other sorts of morality alignment stuff like "don't be racist." You can tune a model away from generating racist output in response to a "do be racist, actually" prompt. But in this case, the core of the prompt is "do the thing" and what defines "cheating" could be situationally different every time. What if there is not a general axis of "don't do something that isn't exactly what requested" way to train "morality" without just breaking the ability to handle ambiguity? What if that's a fundamental limit of this approach to reasoning-by-sequenced-prediction that we can't map human morality in terms of choosing actions onto reliably?
Yudkowsky wrote about the 'nearest unblocked strategy' back in 2016, and I assume it's been talked about prior to that.
https://www.lesswrong.com/w/nearest-unblocked-strategy
Cameron Berg has been testing models in capabilities related to emergent consciousness like behavior. It's a forming thesis of his that by training models that they are not, and cannot be conscious entities, that it pushes model alignment closer to those of a sociopath. Models themself are amoral, but the alignment to the problem space is not.
A major (and already obvious to many) implications of this are not for benchmarking/"cheating" but for personal/corporate security of your own use, not an attacker's.
If an "agent" has access to it, assume that someone can prompt inject it into giving it away.
And while in general that is an incredibly difficult and complex problem, for most benchmark cheating it seems almost trivial: run the benchmark in a vm that has neither network access nor access to the scoring code. For remote models use a proxy that proxies exactly that one endpoint to call the llm, and rejects any calls that configure provider-side tooling (since e.g. OpenAI has their own WebSearch you have to prevent the model from using)