Skip to content

Outrageously Small Neural Networks: Emergent Basic Reasoning at 6,616 tok/SEC [pdf]

huggingface.co
6 pointsAnon843 comments
On HN

Comments

Data is doing more of the work than it used to. Every source in our mixture is a curated artifact built with large models

Training a model this small on them is distillation

When models of this size were last studied seriously such corpora did not exist

blog: https://gregdiamos.com/2026/09/07/outrageously-small-neural-...

X discussion: https://x.com/GregoryDiamos/status/2096873745420075020?s=20

I added some of the main points to the thread so they are easier to read.

The loss does not saturate. Across a 4.91B-token run, smoothed training loss falls monotonically within each curriculum phase and is still descending at the end

AboutSource Built by g1lg1l

Hackerly is an independent reader for Hacker News, built on the public HN API. Not affiliated with Y Combinator.