Comment on Efficient Memory Management for Large Language Model Serving with PagedAttentionparentComments−fredliu2yI might be wrong, but looks like this could help with speculative decoding which can already vastly improves the inference speed?
Comments
I might be wrong, but looks like this could help with speculative decoding which can already vastly improves the inference speed?