Comment on Ollama for Linux – Run LLMs on Linux with GPU AccelerationparentComments−qeternity2yIt is hardware and use case dependent but I would say roughly that ExLlama is 10-20% faster than llama.cpp and ExLlama v2 is 10-20% faster than ExLlama (my experiences at 4 bit quantization).
Comments
It is hardware and use case dependent but I would say roughly that ExLlama is 10-20% faster than llama.cpp and ExLlama v2 is 10-20% faster than ExLlama (my experiences at 4 bit quantization).