Comment on Ollama for Linux – Run LLMs on Linux with GPU AccelerationparentComments−UncleOxidant2yWhere does one find this TVM implementation you mention?−capableweb2yMaybe they're talking about https://github.com/mlc-ai/mlc-llm which is used for web-llm (https://github.com/mlc-ai/web-llm)? Seems to be using TVM.−brucethemoose22yYep.Its very fast on Vulkan, and from what I understand fast on metal, but its not as feature packed as llama.cpp yet.
Comments
Where does one find this TVM implementation you mention?
Maybe they're talking about https://github.com/mlc-ai/mlc-llm which is used for web-llm (https://github.com/mlc-ai/web-llm)? Seems to be using TVM.
Yep.
Its very fast on Vulkan, and from what I understand fast on metal, but its not as feature packed as llama.cpp yet.