That doesn't really track, though. Even tiny models, like Gemma 4, have coherent prose. I know Opus isn't being quantized that small. It seems like it must be some kind of...I dunno. Maybe over-fitting toward some user metric that doesn't track how readable its writing is?
It seems to still be good at code (though I haven't directly compared to earlier Opus versions lately), so it's not a general model collapse type problem, nor quantization errors. If quantization problems, I would expect it to break down on logic before prose, since a 4-bit quantized Gemma 4, even the small versions, have pretty good and, more importantly, coherent written English.
Comments
That doesn't really track, though. Even tiny models, like Gemma 4, have coherent prose. I know Opus isn't being quantized that small. It seems like it must be some kind of...I dunno. Maybe over-fitting toward some user metric that doesn't track how readable its writing is?
It seems to still be good at code (though I haven't directly compared to earlier Opus versions lately), so it's not a general model collapse type problem, nor quantization errors. If quantization problems, I would expect it to break down on logic before prose, since a 4-bit quantized Gemma 4, even the small versions, have pretty good and, more importantly, coherent written English.
Qwen3.8 quant 4 27b works coherently and takes up very little space.
If the cloud AIs arnt downsizing the majority of their customers, they will be out of business.
Qwen3.8-flash-next can run in ~60gb, offloading 50gb to ssd, and is comparible.
The business has to downgrade to be sustainable, and open models prove it.