Skip to content

Comment on Every Model Cheatsparent

Comments

What framework do you purpose? Humans can't even agree on whether or not such values even exist let alone which ones.

Lets say we figure all that out and we pick a core universal morality. If they are smarter than us than how would we know the alignment worked? We would be unable to detect their lies and schemes.

AboutSource Built by g1lg1l

Hackerly is an independent reader for Hacker News, built on the public HN API. Not affiliated with Y Combinator.