Skip to content

Comment on Ollama for Linux – Run LLMs on Linux with GPU Accelerationparent

Comments

exLLAMAv2 and GPTQ are different implementations, and currently exLLAMA is definitely the best for discrete Nvidia (and AMD?) GPUs. Its faster and uses less VRAM for the same perplexity than pretty much anything else.

But the EX2 quantization is very new, and you will have to quantize many models yourself.

But its missing some killer features of llama.cpp, like grammar based sampling.

AboutSource Built by g1lg1l

Hackerly is an independent reader for Hacker News, built on the public HN API. Not affiliated with Y Combinator.