Comment on Advanced Quantization Algorithm for LLMsparentComments−rhdunn4moMy experience is that at Q5 and lower you start to see noticeable degredation in performance/quality. It's especially noticeable at Q4 where models will easily get trapped in repeating token loops. I generally use Q6.[1] https://medium.com/@paul.ilvez/demystifying-llm-quantization...−awestroke4moIs your experience with this new quantization approach from Intel? Otherwise your comment is a bit offtopic at best, misleading at worst.
Comments
My experience is that at Q5 and lower you start to see noticeable degredation in performance/quality. It's especially noticeable at Q4 where models will easily get trapped in repeating token loops. I generally use Q6.
[1] https://medium.com/@paul.ilvez/demystifying-llm-quantization...
Is your experience with this new quantization approach from Intel? Otherwise your comment is a bit offtopic at best, misleading at worst.