Skip to content

Comment on If you can't reproduce the model then it's not open-sourceparent

Comments

If a LLM knows a song as part of its training data, then it is copyright infringement.

No it isn't. You can feed whatever you want into your LLM, including copyrighted data. The issues arise when you start reproducing or distributing copyrighted content.

>You can feed whatever you want into your LLM, including copyrighted data.

That's currently the subject of considerable legal debate.

https://edition.cnn.com/2023/07/10/tech/sarah-silverman-open...

That is mostly an issue of the latter, whether the service that Meta/OpenAI offers outputs content that is a violation of copyright. Technically, derivative works are a copyright violation, but if you're not distributing them, you normally have a good fair use argument, and/or nobody knows.

AboutSource Built by g1lg1l

Hackerly is an independent reader for Hacker News, built on the public HN API. Not affiliated with Y Combinator.