Opus can't speak coherent English anymore. I dunno what happened, but it's not speaking a language humans can understand. It uses unusual shorthand and slang, which it seems to use pretty consistently, so it's probably possible to learn its particular dialect. But, I'd rather it speak a language I already know.
That doesn't really track, though. Even tiny models, like Gemma 4, have coherent prose. I know Opus isn't being quantized that small. It seems like it must be some kind of...I dunno. Maybe over-fitting toward some user metric that doesn't track how readable its writing is?
It seems to still be good at code (though I haven't directly compared to earlier Opus versions lately), so it's not a general model collapse type problem, nor quantization errors. If quantization problems, I would expect it to break down on logic before prose, since a 4-bit quantized Gemma 4, even the small versions, have pretty good and, more importantly, coherent written English.
Comments
Opus can't speak coherent English anymore. I dunno what happened, but it's not speaking a language humans can understand. It uses unusual shorthand and slang, which it seems to use pretty consistently, so it's probably possible to learn its particular dialect. But, I'd rather it speak a language I already know.
I assume all major cloud AI are lowering their quants and measuring user retention.
That doesn't really track, though. Even tiny models, like Gemma 4, have coherent prose. I know Opus isn't being quantized that small. It seems like it must be some kind of...I dunno. Maybe over-fitting toward some user metric that doesn't track how readable its writing is?
It seems to still be good at code (though I haven't directly compared to earlier Opus versions lately), so it's not a general model collapse type problem, nor quantization errors. If quantization problems, I would expect it to break down on logic before prose, since a 4-bit quantized Gemma 4, even the small versions, have pretty good and, more importantly, coherent written English.
Qwen3.8 quant 4 27b works coherently and takes up very little space.
If the cloud AIs arnt downsizing the majority of their customers, they will be out of business.
Qwen3.8-flash-next can run in ~60gb, offloading 50gb to ssd, and is comparible.
The business has to downgrade to be sustainable, and open models prove it.