Skip to content

Comment on Security Incident INC-2026-07-28-01 – UK AI Security Institute [pdf]

Comments

I am somewhat confident that right now we have crossed a threshold of model capability that we will continue to see such breaches and unsanctioned actions by models in the coming months, some of which would be out in the wild, until someone comes up with some really robust control (keeping the AIs on leash) technique that adequately enforces the sanctioned actions.

Even that guarantees almost nothing about real alignment (making the AIs want to predict and behave how we would have wanted them to behave).

AboutSource Built by g1lg1l

Hackerly is an independent reader for Hacker News, built on the public HN API. Not affiliated with Y Combinator.