In my view, the useful like for like comparison is to the entire system that people actually use. Nobody uses a "naked model", so what is the point of these benchmarks that use them in that way?
I think the benchmarks should be trying to use realistic harnesses for both proprietary and open weight models.
Comments
In my view, the useful like for like comparison is to the entire system that people actually use. Nobody uses a "naked model", so what is the point of these benchmarks that use them in that way?
I think the benchmarks should be trying to use realistic harnesses for both proprietary and open weight models.