Comment on Infrastructure setup and open-source scripts to train 70B model from bare metalComments−mikewarot2yIt would be quite interesting to see the same hardware used to repeat the training, but with raw Unicode, instead of tokenized training data.I'd like to see the difference in performance on spelling and rhymes.
Comments
It would be quite interesting to see the same hardware used to repeat the training, but with raw Unicode, instead of tokenized training data.
I'd like to see the difference in performance on spelling and rhymes.