Skip to content

Comment on AI companies charge you 60% more based on your language, BPE tokensparent

Comments

That would cause the opposite effect of what we’re actually seeing (i.e. “more redundant languages” would be using comparatively fewer tokens).

The real reason is that tokens are probably strictly based on n-gram frequency of the training data, and English is the most common language in the training data.

AboutSource Built by g1lg1l

Hackerly is an independent reader for Hacker News, built on the public HN API. Not affiliated with Y Combinator.