Skip to content

Comment on Top AI models fail at >96% of tasks

Comments

Kinda sus that least known model did best and none of the more recent models were tested. Capabilities grow very fast. So things that now routinely succeed rarely ever succeeded even half a year ago.

I mean performance is so bad across the board that this is likely essentially random. Monkeys accidentally doing a bit of Shakespeare.

That's wildly overestimating what monkeys can do on a typewriter.

It takes a lot to just be mediocre. Which, don't get me wrong, I'll agree current ML is, it's just that "mediocre" is an incomprehensibly huge step up from "random".

AboutSource Built by g1lg1l

Hackerly is an independent reader for Hacker News, built on the public HN API. Not affiliated with Y Combinator.