Skip to content

Comment on Choosing an AI model: one prompt, 11 models, different resultsparent

Comments

Yes, that's true. There are many tasks you cannot do this on, but it's far beyond reading other people's broad benchmarks IMHO. e.g. there are tasks where Gemma is great, but on my camera watching, Qwen 3.6's MoE far exceeds either the Gemma 4 Dense or MoE. I suppose the point I meant to make is that evals for a large category of tasks are so accessible that other people's benchmarks are not useful at all.

Using LLMs, whether it's for coding, internal, or external uses, without actually validating what you're doing somehow (and evals I think are the best option right now), is a bit like buying shoes solely based on the text description on the box.

AboutSource Built by g1lg1l

Hackerly is an independent reader for Hacker News, built on the public HN API. Not affiliated with Y Combinator.