Comment on GigaToken: ~1000x faster Language model tokenizationComments−zX41ZdbW1moThis is exactly what we need! Will try: https://github.com/ClickHouse/ClickHouse/issues/108247It will be nicer if the README focuses more on per-core performance.About the actual algorithm - will something like matching in a perfect hash table help?
Comments
This is exactly what we need! Will try: https://github.com/ClickHouse/ClickHouse/issues/108247
It will be nicer if the README focuses more on per-core performance.
About the actual algorithm - will something like matching in a perfect hash table help?