Skip to content

Comment on Ollama for Linux – Run LLMs on Linux with GPU Accelerationparent

Comments

Hamel Husain hasn't done testing vs llama.cpp, but this still might be of interest (includes mlc which is roughly in line w/ llama.cpp batch=1 perf): https://hamel.dev/notes/llm/inference/03_inference.html

He has benchmarks on an A6000 which should be roughly in line w/ a 3090 if you want to compare to my numbers (I test mlc as well, although my 3090 results are slower since I'm testing a llama2-7b @ 4K context and mlc currently slows down significantly w/ longer context): https://docs.google.com/spreadsheets/d/1kT4or6b0Fedd-W_jMwYp...

AboutSource Built by g1lg1l

Hackerly is an independent reader for Hacker News, built on the public HN API. Not affiliated with Y Combinator.