Comment on Flash-MSA: Accelerating Million-Token Training with Sparse Attention KernelsparentComments−kamranjon2moThe Minimax paper was published in June 2026 coinciding with the Minimax M3 release - I’m not sure how the repo you posted here could have been an implementation of Minimax sparse attention when it was updated over a year ago?
Comments
The Minimax paper was published in June 2026 coinciding with the Minimax M3 release - I’m not sure how the repo you posted here could have been an implementation of Minimax sparse attention when it was updated over a year ago?