Show HN: Phase Router – capacity-aware routing for MoEgithub.com/TSltd 5 pointsTSltd4 months ago1 commentSaveHideCopy link On HNComments−TSltdOP4moReducing dropped tokens could also improve model training by reducing gradient noise
Comments
Reducing dropped tokens could also improve model training by reducing gradient noise