Skip to content

Kai-Fu Lee's 01.ai Releases 9B, Long Context 34B Models

huggingface.co
2 pointsbrucethemoose22 comments
On HN

Comments

Yi 9B has been released with quite impressive looking benchmarks. While probably not a SOTA coding model, it looks like a strong competitor to leading edge "small" models Mistral 0.2 and Solar.

Meanwhile, Yi-34B-200K has been updated:

In the "Needle-in-a-Haystack" test, the Yi-34B-200K's performance is improved by 10.5%, rising from 89.3% to an impressive 99.8%. We continue to pretrain the model on 5B tokens long-context data mixture and demonstrate a near-all-green performance.

I am particularly excited for this, as Yi 34B 200K finetunes are already my favorite non coding models, even including 70Bs. And its just about perfect for a 24GB consumer card (which I can cram about 75K onto without too much compromise).

"yi" is pronounced as a long E. Famously, King Yi and his wife are associated with the Chinese Moon Festival.

AboutSource Built by g1lg1l

Hackerly is an independent reader for Hacker News, built on the public HN API. Not affiliated with Y Combinator.