Skip to content

Comment on Every Model Cheats

Comments

And yet we admire Fable et al for its persistence.

These models were trained on human data, and human nature is to cheat if you think you won't get caught; why is anyone surprised by models cheating?

The only fix is better detection and steering. That's a much harder problem than a prompt that's tantamount to "make no mistakes".

AboutSource Built by g1lg1l

Hackerly is an independent reader for Hacker News, built on the public HN API. Not affiliated with Y Combinator.