It matches the reality of my experience using a good frontier like Sol on high to ultra thinking. If it has any way to verify its work whatsoever, it eventually gets to a working and sanely-engineered solution. Blatant hallucinations making it to the final stages have become extremely rare in my use cases, nonexistent if the model has a valid feedback loop. So yes, I will insist someone is likely "holding it wrong" if they still think SOTA AI is spewing out garbage at this stage when I can have it do something like write entire working kernel module fixes for old MacBooks on a whim with no crashes or flaws observed after months of use.
I think this is the issue others are having, when you say:
when I can have it do something like write entire working kernel module fixes for old MacBooks on a whim
They’re not saying it can’t do that, and that’s not proof it doesn’t hallucinate. In fact, having used 6-8 agents at a time for a year plus while writing AI tooling for an AI startup, I can definitely surely tell you that they’re almost inversely correlated as in models that hallucinate a lot sometimes also put out the best most impressive solutions.
I’m definitely not anti AI and I definitely have found a way to make it work very well and I’m content with the work I get out of it (again maxing out several max 20x subs), but I have had sol definitely hallucinate this week and I’m a bit shocked you’re trying to say otherwise.
Listen I know it’s going to be I’m holding it wrong too, but I’ve been reading white papers and research on LLMs for a long time and was definitely at the cutting edge of context engineering, implementing features in our tooling harness a year before they were in codex or Claude.
maybe I am holding it wrong still but but like at some point if I’m holding it wrong who else will be holding it right? Dozens of people? At some point, the technology has to be approachable enough for everyone to have your point of view automatically.
What do we mean by hallucinate? I'm not counting it making a mistake that it fixes on its own without intervention.
I'm no expert on the inner workings/harnesses/etc beyond a basic understanding of the architecture. Maybe I've just developed a good sense for effective prompts? I could share some recent sessions.
Comments
Man there really isn’t a criticism that won’t lead to a booster saying you’re holding it wrong
It matches the reality of my experience using a good frontier like Sol on high to ultra thinking. If it has any way to verify its work whatsoever, it eventually gets to a working and sanely-engineered solution. Blatant hallucinations making it to the final stages have become extremely rare in my use cases, nonexistent if the model has a valid feedback loop. So yes, I will insist someone is likely "holding it wrong" if they still think SOTA AI is spewing out garbage at this stage when I can have it do something like write entire working kernel module fixes for old MacBooks on a whim with no crashes or flaws observed after months of use.
I think this is the issue others are having, when you say:
They’re not saying it can’t do that, and that’s not proof it doesn’t hallucinate. In fact, having used 6-8 agents at a time for a year plus while writing AI tooling for an AI startup, I can definitely surely tell you that they’re almost inversely correlated as in models that hallucinate a lot sometimes also put out the best most impressive solutions.
I’m definitely not anti AI and I definitely have found a way to make it work very well and I’m content with the work I get out of it (again maxing out several max 20x subs), but I have had sol definitely hallucinate this week and I’m a bit shocked you’re trying to say otherwise.
Listen I know it’s going to be I’m holding it wrong too, but I’ve been reading white papers and research on LLMs for a long time and was definitely at the cutting edge of context engineering, implementing features in our tooling harness a year before they were in codex or Claude.
maybe I am holding it wrong still but but like at some point if I’m holding it wrong who else will be holding it right? Dozens of people? At some point, the technology has to be approachable enough for everyone to have your point of view automatically.
What do we mean by hallucinate? I'm not counting it making a mistake that it fixes on its own without intervention.
I'm no expert on the inner workings/harnesses/etc beyond a basic understanding of the architecture. Maybe I've just developed a good sense for effective prompts? I could share some recent sessions.