Comment on State-of-the-Art Chatbot, Vicuna-7B, now runs on MacBook with GPU accelerationparentComments−superkuh3yNah, I don't use huggingface transformers to run inference with the vicuna model. I use llama.cpp. But I do appreciate the tip.edit: Oh, I was completely wrong. That's in the training not the inference so it applies to all the weights.
Comments
Nah, I don't use huggingface transformers to run inference with the vicuna model. I use llama.cpp. But I do appreciate the tip.
edit: Oh, I was completely wrong. That's in the training not the inference so it applies to all the weights.