Skip to content

Comment on Choosing an AI model: one prompt, 11 models, different resultsparent

Comments

You touch a point I quickly skimmed in another comment.

Yes, the most valuable benchmarks and evaluations you can write are those that resemble your work.

The evaluations are extremely hard to write and test.

And yes, virtually all benchmarks are E2E one shots, they do not reflect multi turn processes or how most people interact with LLMs.

Which is why every Opus after 4.6 looks better on benchmarks, but is hard to work with interactively.

AboutSource Built by g1lg1l

Hackerly is an independent reader for Hacker News, built on the public HN API. Not affiliated with Y Combinator.