Skip to content

Comment on Evals are not all you needparent

Comments

Exactly. Benchmarks != evals.

Task specific evals are how you determine if an AI system works. Expanding your evals as you test, find new issues, and change techniques is the only way to ensure each change doesn’t regress a bunch of your past work. If you don’t build up evals, you end up stuck on some 3 year old model with mid performance, unable to update because of the risk. Without evals, your assessments are just “vibes”.

Random plug but it seems relevant here: I launched an eval toolkit earlier today. It includes synthetic eval data gen, automated evals, great UI so everyone on the team improve quality (QA, PM, not just data scientists), adversarial red-teaming, and human preference correlations. https://docs.getkiln.ai/docs/evaluations

AboutSource Built by g1lg1l

Hackerly is an independent reader for Hacker News, built on the public HN API. Not affiliated with Y Combinator.