Skip to content

Comment on Megatron-Turing NLG 530B, the World’s Largest Generative Language Modelparent

Comments

All we need to worry about is the weights and the dynamics of the network as a whole. How much of a simplification is that?

A lot. Parallel optimization is an art form. These models are trained on static datasets, they can't intervene in the environment to infer causal relations, so they need legs and hands.

AboutSource Built by g1lg1l

Hackerly is an independent reader for Hacker News, built on the public HN API. Not affiliated with Y Combinator.