In addition to naming one, I'd also be interesting in whether they actually do rigorous work.
The safety research that tends to get headlines is often extremely misleading, usually with directed prompting, or unreported additions to the system prompt specifying model roleplay behavior.
Comments
Are there any examples of successful startups doing this?
In addition to naming one, I'd also be interesting in whether they actually do rigorous work.
The safety research that tends to get headlines is often extremely misleading, usually with directed prompting, or unreported additions to the system prompt specifying model roleplay behavior.