Skip to content

Comment on GigaToken: ~1000x faster Language model tokenization

Comments

Can I say this seems to be fantastic work. I cloned your repo earlier today after seeing it on the tokenization discord. I know everyone in the tokenization community wants to absorb the lessons of how you got such a speedup. The caching and replacing the regex for pretokenization seem like generally useful ideas.

And screw all the 0.1% haters on here, this is great stuff.

Thanks for the kind words, Craig! I'm planning to do a technical writeup+paper and a presentation video on the project in the near future. Will make sure to share it with the Discord!

That is my reaction too. It looks like great work!

Valuable not only for inference, but for training too (think proprietary datasets).

I would add, a single individual did this.

One person can make a difference :-)

tokenization discord

How can I join this? Sounds interesting

Send me an email (my address is in my profile)

also subscribing to this!

Send me an email (my address is in my profile)

the tokenization discord

Could I join this?

Send me an email (my address is in my profile)

AboutSource Built by g1lg1l

Hackerly is an independent reader for Hacker News, built on the public HN API. Not affiliated with Y Combinator.