Skip to content

Comment on Why eval startups fail (2025)

Comments

I believe eval startups can work when they're targeting safety benchmarks specifically.

Are there any examples of successful startups doing this?

In addition to naming one, I'd also be interesting in whether they actually do rigorous work.

The safety research that tends to get headlines is often extremely misleading, usually with directed prompting, or unreported additions to the system prompt specifying model roleplay behavior.

AboutSource Built by g1lg1l

Hackerly is an independent reader for Hacker News, built on the public HN API. Not affiliated with Y Combinator.