I don’t see how you could run Qwen3.8 27B on 16GB of memory that’s shared with Linux. Are people running models at 2bit quants? Are they even worth bothering with? I had assumed you go down to 4bit and if you need to go smaller you have to lose parameters.
You can, it's kind of cool to have this capability on something gaming at such a low price tag. Its not optimal for your time but beats nothing by a LOT. And that model is pretty reliable.
Comments
I don’t see how you could run Qwen3.8 27B on 16GB of memory that’s shared with Linux. Are people running models at 2bit quants? Are they even worth bothering with? I had assumed you go down to 4bit and if you need to go smaller you have to lose parameters.
You can, it's kind of cool to have this capability on something gaming at such a low price tag. Its not optimal for your time but beats nothing by a LOT. And that model is pretty reliable.
https://unsloth.ai/docs/basics/dynamic-3.0-ggufs
You can run a 4bit Qwen3.8 27B on a 8GB GPU, though on my RX 6650 XT it will only reach about 3 tokens/s.
If you can give the GPU units enough RAM to load the whole thing, it may be a little faster.
I'm not sure why would you want to run something at 3-4 tokens/s though. At least try a MoE model. I get about 30 tokens/s from those.