Skip to content

Comment on How we monitor internal coding agents for misalignmentparent

Comments

Does anyone have any understanding of how they do this?

My knowledge of how these models work is basically that they are a black box that you put text into and get text out of. I don't phrase it this way to diminish their capability, but more to ask how, other than using a technique like stenography, are they able to hide their true chain of thought in a recoverable way?

Welch Labs on YouTube has a great collection of videos on how AI models learn. His recent video [1] covers how image models can learn to encode reasoning in the image processing layers when not given an out of band reasoning set of weights to use instead. I suspect that this applies to LLMs and CoT reasoning vs output token weights.

[1] https://www.youtube.com/watch?v=QgH9sr7G13Q

AboutSource Built by g1lg1l

Hackerly is an independent reader for Hacker News, built on the public HN API. Not affiliated with Y Combinator.