Task specific evals are how you determine if an AI system works. Expanding your evals as you test, find new issues, and change techniques is the only way to ensure each change doesn’t regress a bunch of your past work. If you don’t build up evals, you end up stuck on some 3 year old model with mid performance, unable to update because of the risk. Without evals, your assessments are just “vibes”.
Random plug but it seems relevant here: I launched an eval toolkit earlier today. It includes synthetic eval data gen, automated evals, great UI so everyone on the team improve quality (QA, PM, not just data scientists), adversarial red-teaming, and human preference correlations. https://docs.getkiln.ai/docs/evaluations
Comments
Exactly. Benchmarks != evals.
Task specific evals are how you determine if an AI system works. Expanding your evals as you test, find new issues, and change techniques is the only way to ensure each change doesn’t regress a bunch of your past work. If you don’t build up evals, you end up stuck on some 3 year old model with mid performance, unable to update because of the risk. Without evals, your assessments are just “vibes”.
Random plug but it seems relevant here: I launched an eval toolkit earlier today. It includes synthetic eval data gen, automated evals, great UI so everyone on the team improve quality (QA, PM, not just data scientists), adversarial red-teaming, and human preference correlations. https://docs.getkiln.ai/docs/evaluations