Predicting the Order of Upcoming Tokens Improves Language Modelingarxiv.org 7 pointswavelander1 year ago2 commentsSaveHideCopy link On HNComments−NitpickLawyer1yAre any of these methods doable on pre-trained models? Like freeze the model and only train these add-ons? Having to redo the training runs with these optimisations doesn't sound too practical, in the great scheme of things.−impossiblefork1yIt's obviously practical for the next model you train from scratch. The point of research is obviously not to improve existing commercial products.
Comments
Are any of these methods doable on pre-trained models? Like freeze the model and only train these add-ons? Having to redo the training runs with these optimisations doesn't sound too practical, in the great scheme of things.
It's obviously practical for the next model you train from scratch. The point of research is obviously not to improve existing commercial products.