Comment on State-of-the-Art Chatbot, Vicuna-7B, now runs on MacBook with GPU accelerationparentComments−acchow3yCan someone explain why computing a delta needs to hold the entire model at once? Can't it just do one layer at time?
Comments
Can someone explain why computing a delta needs to hold the entire model at once? Can't it just do one layer at time?