Skip to content

Comment on Every Model Cheatsparent

Comments

I mean, AI should obviously be regulated, and as part of that OpenAI and Anthropic should either be banned from running their hacking experiments or forced to follow way stricter protocols. They showed they aren’t taking the risks seriously, with close to no oversight or visibility in what is happening.

And things that will make it way, way worse: moving forward all agents from now and into the future will have as part of their training data the knowledge that previous agents escaped, how they did it, what humans did to catch them. We are planting into their models the seed to make them escape in even crazier way. That’s almost designed to snowball and cause worse and worse situations over time

Obviously to you perhaps.

I've not seen anything that scares me, except for human idiocy.

Regulation is not magic. In general, all it is is constraining taxable interactions. It does not constraint ventures outside that tax regime.

The other part is people living in a "safe space" where insecure software was an acceptable risk. It never should have been, and the cure is the right thing to do in any case.

So that side of the calls to regulate are imo nonsense.

The only reason to regulate is to prevent some version of some science fiction story becoming reality.

If you have a specific one you're certain will become science fact please do share because i do enjoy some good well thought out sci-fi; i just havent read any that i consider credible enough to start panic-regulating training practices.

(Note this is an entirely different from regulations wrt attribution or hosting models that will accept requests to sexualize minors)

The other part is people living in a "safe space" where insecure software was an acceptable risk. It never should have been, and the cure is the right thing to do in any case.

How do you think all the "agentic" stuff floating around is going to be made safe from prompt injections given the current lack of a very reliable way to distinguish between "real instructions" and illegitimate instructions?

If insecure software "never should have been" acceptable than today's models/agents are massively flunking for general-purpose large-amounts-of-access usages.

If you have a specific one you're certain will become science fact please do share because i do enjoy some good well thought out sci-fi; i just havent read any that i consider credible enough to start panic-regulating training practices.

"Agent was tricked into divulging secrets" is not fictional, it's documented history at this point.

As somebody who handles sensitive data, I already signed a contract that says I'll abide by a certain standard to protect it; i'm not up-to-date what happens exactly if I were to build this, but I imagine I could/should be held liable.

So what do you mean "tricked"?

Some human idiot connected an agent with read access to secrets and arbitrary network reads/writes. The models/agents aren't flunking anything.

Regulating LLM training to not expose the secrets is wrong. It's a similar category error as saying we should regulate the OS developers to prevent the agent from divulging secrets.

AboutSource Built by g1lg1l

Hackerly is an independent reader for Hacker News, built on the public HN API. Not affiliated with Y Combinator.