Skip to content

Comment on GigaToken: ~1000x faster Language model tokenization

Comments

Cool stuff. From my understanding, this is less valuable at inference time and more useful when running offline pre-training data prep.

When tokenizing terabytes of text for your training corpus, the speedup here is probably doing real work in saving you time (and money?). You get a faster iteration cycle when figuring out and adjusting your datasets.

Also for embeddings model

AboutSource Built by g1lg1l

Hackerly is an independent reader for Hacker News, built on the public HN API. Not affiliated with Y Combinator.