Skip to content

Comment on Kai-Fu Lee's 01.ai Releases 9B, Long Context 34B Models

Comments

Yi 9B has been released with quite impressive looking benchmarks. While probably not a SOTA coding model, it looks like a strong competitor to leading edge "small" models Mistral 0.2 and Solar.

Meanwhile, Yi-34B-200K has been updated:

In the "Needle-in-a-Haystack" test, the Yi-34B-200K's performance is improved by 10.5%, rising from 89.3% to an impressive 99.8%. We continue to pretrain the model on 5B tokens long-context data mixture and demonstrate a near-all-green performance.

I am particularly excited for this, as Yi 34B 200K finetunes are already my favorite non coding models, even including 70Bs. And its just about perfect for a 24GB consumer card (which I can cram about 75K onto without too much compromise).

AboutSource Built by g1lg1l

Hackerly is an independent reader for Hacker News, built on the public HN API. Not affiliated with Y Combinator.