Skip to content

Comment on Has the hallucination problem in AI been solved?

Comments

I think its solved, with the right setup and model. I couldn't tell u the last time my agent hallucinated. For me I consider it solved. However some dude feeding massive docs into gpt chat and long conversations, It is not solved in this context.

Man there really isn’t a criticism that won’t lead to a booster saying you’re holding it wrong

It matches the reality of my experience using a good frontier like Sol on high to ultra thinking. If it has any way to verify its work whatsoever, it eventually gets to a working and sanely-engineered solution. Blatant hallucinations making it to the final stages have become extremely rare in my use cases, nonexistent if the model has a valid feedback loop. So yes, I will insist someone is likely "holding it wrong" if they still think SOTA AI is spewing out garbage at this stage when I can have it do something like write entire working kernel module fixes for old MacBooks on a whim with no crashes or flaws observed after months of use.

I think this is the issue others are having, when you say:

when I can have it do something like write entire working kernel module fixes for old MacBooks on a whim

They’re not saying it can’t do that, and that’s not proof it doesn’t hallucinate. In fact, having used 6-8 agents at a time for a year plus while writing AI tooling for an AI startup, I can definitely surely tell you that they’re almost inversely correlated as in models that hallucinate a lot sometimes also put out the best most impressive solutions.

I’m definitely not anti AI and I definitely have found a way to make it work very well and I’m content with the work I get out of it (again maxing out several max 20x subs), but I have had sol definitely hallucinate this week and I’m a bit shocked you’re trying to say otherwise.

Listen I know it’s going to be I’m holding it wrong too, but I’ve been reading white papers and research on LLMs for a long time and was definitely at the cutting edge of context engineering, implementing features in our tooling harness a year before they were in codex or Claude.

maybe I am holding it wrong still but but like at some point if I’m holding it wrong who else will be holding it right? Dozens of people? At some point, the technology has to be approachable enough for everyone to have your point of view automatically.

What do we mean by hallucinate? I'm not counting it making a mistake that it fixes on its own without intervention.

I'm no expert on the inner workings/harnesses/etc beyond a basic understanding of the architecture. Maybe I've just developed a good sense for effective prompts? I could share some recent sessions.

Would you trust it to make life and death decisions? Because it is being used in that context. Drones with AI, armed, and with discretion to choose a target and kill it.

The types of people using it that way are not concerned with the best interests of anyone but themselves, and arguably not even that except within a very immediate time horizon.

Perhaps you should be concerned about it. I certainly am. This is deployed on the battlefield and has been for months.

What makes you think I'm not concerned about it? I've been positively mortified at the seeming cavalier way in which all caution is thrown to the winds. I've honestly had to fundamentally reassess my understanding of the average character of humanity given the last decade.

Fact is, those with the foresight to see how this can go badly are also seemingly the types of people who don't end up in a position to prevent harms by it. Furthermore, it seems inevitable in a sense, because greed for power seems to necessitate development of automated weapons in ASAP in spite of the risks, and the fact it basically renders traditional warfare pointless.

It appears we'll have to learn the hard lessons, same as our forebearers with. I just hope we can avoid having to regress back to sticks and stones on account of fucking ourselves by overdoing our capability to destroy on account of not being willing0able to peacefully coexist.

Thanks for the response. I see it pretty much the same way.

No, because that's Terminator territory (ULAWS). Since drones' inception, US military required an officer to approve lethal weapon release until at least 2013 and it may or may not still be required. The official policy is DODD 3000.09, which uses the vague word "appropriate" 22 times, but doesn't allow unsupervised lethal autonomous (ULAWS) yet. Maybe someone knows what the criteria are used around in the world's militaries these days beyond Political Declaration on Responsible Military Use of Artificial Intelligence and Autonomy (2023-2024) (which Ukraine signed) (There's a competing REAIM 2023 Call to Action that Ukraine didn't sign with a note.[0])

In other news, New Orleans 911 is using AI to triage localized incident duplication calls from unique emergencies. I'm not saying that it's good or their only practical choice, but it's happening.

The biggest dangers I see are the outsourcing of supervisory control, appeal to authority (when used to summarize content or answer a question), and hallucinated mistakes.

0. (PDF) https://docs-library.unoda.org/General_Assembly_First_Commit...

1. PDRMUAIA https://www.state.gov/bureau-of-arms-control-deterrence-and-...

2. REAIM 2023 Call to Action https://www.government.nl/documents/2023/02/16/reaim-2023-ca...

3. REAIM 2023 Endorsing Countries and Territories https://www.government.nl/documents/2023/02/16/reaim-2023-en...

I'm not talking about a hypothetical. I'm talking about Ukraine assymetrically fucking up the Russians with AI drones.

Completely different type of AI

AboutSource Built by g1lg1l

Hackerly is an independent reader for Hacker News, built on the public HN API. Not affiliated with Y Combinator.