Comment on TransMLA: Multi-head latent attention is all you needComments−olq_plo1yVery cool idea. Can't wait for converted models on HF.−MichaelMoser1231ydeepseek-v2,v3,r1 are all using multi-headed attention.
Comments
Very cool idea. Can't wait for converted models on HF.
deepseek-v2,v3,r1 are all using multi-headed attention.