Comment on MLC-LLM: GPT/Llama on consumer-class GPUs and phonesparentComments−trifurcate3yThe slowdown here is going to be more about the increased context length than throttling.−david-gpu3yGood point! Didn't think of that.
Comments
The slowdown here is going to be more about the increased context length than throttling.
Good point! Didn't think of that.