Skip to content

Comment on Some critical issues with the SWE-bench datasetparent

Comments

Yeah, that's true in many fields with these AI agents. They demo well, but when you put them to actual work they fall right on their face. Even worse, the harder the task you set for them the more they lie to you. It's like hiring a junior dev from one of those highly regimented societies where it's more important to save face than to get the job done.

It's almost as if they're not trying to market to the people actually using the products, but trying to convince investors of features that don't exist

Yep it's "full self driving in 1 year" all over again.

Its the good old Elon musk playbook spread out across the industry.

Someone should coin a term for this very new phenomenon. Maybe “vaporware”?

Your last sentence feels kind of spot on. The lack of transparency around confidence in the answer makes it hard to use (and I know it would not be simple to add such a thing)

sounds like a skill issue to be honest. you could probably tell the assistant to just ask you questions when information is missing instead

Have you actually tried this? What happens is it will very often ask you questions at irrelevant times so you start ignoring the questions and it becomes wasted space.

Even OpenAI hasn't figured it out, because their Deep Research always asks questions before starting the search.

Programming is easy. Asking the right question is hard.

People don't know what questions to ask.

But it doesn't know when information is missing

AboutSource Built by g1lg1l

Hackerly is an independent reader for Hacker News, built on the public HN API. Not affiliated with Y Combinator.