Comment on MLC-LLM: GPT/Llama on consumer-class GPUs and phonesparentComments−int_19h3yWhen people tried 3-bit quantization for 7B models before, it did not exactly go well in terms of detrimental side effects. Are you using some new quantization techniques that mitigate that?
Comments
When people tried 3-bit quantization for 7B models before, it did not exactly go well in terms of detrimental side effects. Are you using some new quantization techniques that mitigate that?