Skipping 90% of KV dequant work speeds up LLM decode by 22%github.com/TheTom 1 pointpidtom5 months agodiscussSaveHideCopy link On HNComments No comments yet.
Comments
No comments yet.