Skip to content

Comment on How has DeepSeek improved the Transformer architecture?parent

Comments

They likely continue to train dense models because they are far easier to fine tune and this is a huge use case for the Llama models

It probably also has to do with their internal infra. If it were just about dense models being easier for the OSS community to use & build on, they should probably be training MoEs and then distilling to dense.

AboutSource Built by g1lg1l

Hackerly is an independent reader for Hacker News, built on the public HN API. Not affiliated with Y Combinator.