Skip to content

Comment on How we monitor internal coding agents for misalignmentparent

Comments

If a model were actually capable of scheming, it would also have enough situational awareness

This doesn't follow. It might not have that kind of control over its the chain-of-thought, even if in some sense it knew that would be a good idea. Also, they are specifically not training on the chain of thought so it doesn't gain that ability.

AboutSource Built by g1lg1l

Hackerly is an independent reader for Hacker News, built on the public HN API. Not affiliated with Y Combinator.