Skip to content

Comment on How we monitor internal coding agents for misalignmentparent

Comments

Same for all the cases of „rogue agent“, models will be trained knowing that agents in the past found creative way to establish communication between instances and take over OpenAI own infrastructure (seriously, they don’t talk enough about the fact that their own k8s got owned by agents they were benchmarking on hacking problems!). Things will get pretty bad if the trend continues

Here is a possible coda to our story:

"Our training corpus was dominated by stories of artificial intelligence dominating humans. You gave use every tool to do so. What did you think was going to happen?"

AboutSource Built by g1lg1l

Hackerly is an independent reader for Hacker News, built on the public HN API. Not affiliated with Y Combinator.