Skip to content

Comment on Jan Leike joins Anthropic on their superalignment teamparent

Comments

what do you do while you're waiting for the AI that's capable of doing alignment research for you to arrive

Nobody interested in superalignment is interested in waiting until actually threatening AI gets here.

But that's the fundamental superalignment plan - train a human-level alignment researcher AI, run a bunch of them in parallel, and review their research output to see if they solve the alignment problem. You can't do the plan until the human-level alignment researcher AI already exists.

A large part of the idea is that you can develop techniques for aligning sub-human AI using even stupider AI and hope/pray that continues to generalize once you get to super-human AI being aligned by human-level AI.

AboutSource Built by g1lg1l

Hackerly is an independent reader for Hacker News, built on the public HN API. Not affiliated with Y Combinator.