Skip to content

Comment on What is nueralese and why is it bad

Comments

What exactly prevents anyone worried about this to build an LLM that can decode the neuralese to english and use it to monitor what the model is doing?

Or is the neuralese some sort of irreversibly encrypted data set that only an LLM can "understand" and that can never be translated back to English?

It's very hard in practice. Even without optimization against it, it's about as hard as understanding an activation layer today and will probably get harder in the future.

There's also nothing stopping the secondary LLM from confabulating bullshit and there are less checks on it than CoT (harder for either humans or other models to externally verify).

How do you trust that LLM?

AboutSource Built by g1lg1l

Hackerly is an independent reader for Hacker News, built on the public HN API. Not affiliated with Y Combinator.