Comment on Show HN: Phase Router – capacity-aware routing for MoEComments−TSltdOP4moReducing dropped tokens could also improve model training by reducing gradient noise
Comments
Reducing dropped tokens could also improve model training by reducing gradient noise