Comment on Measuring What Matters: Construct Validity in Large Language Model BenchmarksComments−ammaox10moA very large review of AI benchmarks that reveals a worrying trend in their effectiveness and scientific rigor
Comments
A very large review of AI benchmarks that reveals a worrying trend in their effectiveness and scientific rigor