Knowing Enough About MoE to Explain Dropped Tokens in GPT-4152334h.github.io 3 points152334H3 years ago1 commentSaveHideCopy link On HNComments−turtleyacht3yIn AI/ML, Mixture of Experts (MoE)."GPT-4 uses a simple top-2 Token Choice router for MLP MoE layers. It does not use MoE for attention."GPT won't fix, since "tokens being dropped are generally good for the performance of MoE models."https://152334h.github.io/blog/knowing-enough-about-moe/#con...
Comments
In AI/ML, Mixture of Experts (MoE).
"GPT-4 uses a simple top-2 Token Choice router for MLP MoE layers. It does not use MoE for attention."
GPT won't fix, since "tokens being dropped are generally good for the performance of MoE models."
https://152334h.github.io/blog/knowing-enough-about-moe/#con...