Yes, that's true. There are many tasks you cannot do this on, but it's far beyond reading other people's broad benchmarks IMHO. e.g. there are tasks where Gemma is great, but on my camera watching, Qwen 3.6's MoE far exceeds either the Gemma 4 Dense or MoE. I suppose the point I meant to make is that evals for a large category of tasks are so accessible that other people's benchmarks are not useful at all.
Using LLMs, whether it's for coding, internal, or external uses, without actually validating what you're doing somehow (and evals I think are the best option right now), is a bit like buying shoes solely based on the text description on the box.
Comments
Yes, that's true. There are many tasks you cannot do this on, but it's far beyond reading other people's broad benchmarks IMHO. e.g. there are tasks where Gemma is great, but on my camera watching, Qwen 3.6's MoE far exceeds either the Gemma 4 Dense or MoE. I suppose the point I meant to make is that evals for a large category of tasks are so accessible that other people's benchmarks are not useful at all.
Using LLMs, whether it's for coding, internal, or external uses, without actually validating what you're doing somehow (and evals I think are the best option right now), is a bit like buying shoes solely based on the text description on the box.