Comment on Ollama for Linux – Run LLMs on Linux with GPU AccelerationparentComments−lhl2yThe numbers are always changing, but from my testing, they're close enough that it doesn't really matter. My most recent benchmarks: https://docs.google.com/spreadsheets/d/1kT4or6b0Fedd-W_jMwYp...I'd say that you should pick the backend that has the quantized models or other features (sampler, context window, API compatibility, etc) that suits you best.
Comments
The numbers are always changing, but from my testing, they're close enough that it doesn't really matter. My most recent benchmarks: https://docs.google.com/spreadsheets/d/1kT4or6b0Fedd-W_jMwYp...
I'd say that you should pick the backend that has the quantized models or other features (sampler, context window, API compatibility, etc) that suits you best.