Skip to content

Comment on The Crescendo Multi-Turn LLM Jailbreak Attackparent

Comments

I think for the most part it's not that current generation of models are dangerous, at least not for things like this, but rather that researchers want to learn how to control them now to be ready for what's coming...

Right now the models' reasoning capabilities aren't good enough that they can add too much to what's already on the web and available by search, but soon they will be. Anthropic spent 6 months talking to researchers about biological threats and came to conclusion that their models would be capable of figuring out the "missing pieces" (information that is not publicly available) for various threats within a couple of years.

AboutSource Built by g1lg1l

Hackerly is an independent reader for Hacker News, built on the public HN API. Not affiliated with Y Combinator.