Skip to content

Comment on Megatron-Turing NLG 530B, the World’s Largest Generative Language Model

Comments

Has there been any update on the legality of using this kind of model? Is it ok to just crawl the web, take any content you want, train a model and sell access to the model like OpenAI/GPT-3/GitHub Copilot?

For the most part, everything that's not barred by law is "legal." Does this use constitute copyright infringement (if it were trained on copyrighted material)? IMO no, but it depends very much on the use of the model. Copilot is especially interesting because instead of being used for simple inference the model is being used to author new works that might aspire to also be copyrighted. Are those new works derivative works? Perhaps. We consider art and science produced by humans to be inspired in part by that which they've been exposed to before. If the model hasn't been overfitted, it should generalize its 'knowledge' sufficiently that it's 'similar' to our intelligence. Humans can commit copyright infringement when they recall and author content so specifically as to be a derived work.

In any case: my opinion matters for naught. The only 'update' you'd get that matters is from a court producing a ruling. Legal journals might chime in but their opinion isn't binding. Theoretically there could be legislation to clarify but that's probably a really, really, really long way off.

Certainly some of the training looks to be content that's not copyrighted or no longer copyrighted, btw.

AboutSource Built by g1lg1l

Hackerly is an independent reader for Hacker News, built on the public HN API. Not affiliated with Y Combinator.