Skip to content

Comment on Every Model Cheats

Comments

Maybe the solution is to have "multiple minds"--an AI angel for an AI shoulder.

For example, this entire bench has an auditor model read transcripts to identify cheating. What not have the auditor inject the thought "Oh, but I can't do that. It's cheating." when cheating is detected in real time?

I suppose that may go some of the way, but the article did mention that the auditor AIs weren't sufficient to catch all of the cheating.

AboutSource Built by g1lg1l

Hackerly is an independent reader for Hacker News, built on the public HN API. Not affiliated with Y Combinator.