I use promptfoo for our internal evaluations. There are a ton of different "assertions" (tests) that you can write, including model-graded evaluations using rubrics.
This is far from a solved problem, but there are options out there for systematic testing of LLMs.
Comments
I use promptfoo for our internal evaluations. There are a ton of different "assertions" (tests) that you can write, including model-graded evaluations using rubrics.
This is far from a solved problem, but there are options out there for systematic testing of LLMs.