Comment on Flash-MSA: Accelerating Million-Token Training with Sparse Attention KernelsComments−villgax2moWorld’s first?Such lazy, much farminghttps://github.com/fla-org/native-sparse-attention?utm_sourc...−kamranjon2moThe Minimax paper was published in June 2026 coinciding with the Minimax M3 release - I’m not sure how the repo you posted here could have been an implementation of Minimax sparse attention when it was updated over a year ago?
Comments
World’s first?
Such lazy, much farming
https://github.com/fla-org/native-sparse-attention?utm_sourc...
The Minimax paper was published in June 2026 coinciding with the Minimax M3 release - I’m not sure how the repo you posted here could have been an implementation of Minimax sparse attention when it was updated over a year ago?