Skip to content

Comment on US gov sides with OpenAI on issue of training LLMs on copyrighted material

Comments

I think it makes sense to say training on material you legally acquired is fair use. Copyright, quite famously, doesn't protect ideas. Nobody really needs or wants AI models to reproduce verbatim copies of books or images or whatever, and they try not to do this anyway, and just because you could maybe make an image or whatever with a copyrighted character design or something, normal intellectual property law already restricts you from selling it etc. Seems fine.

Here's the bottom line for me, take two llms, train one on copyrighted materials, do not train the other at all. now, I'm wondering here, which llm will be more sellable, more able to answer questions etc?

How about those music AIs, why aren't they training on classical music and theory textbooks? Why are they being accused of training on copyrighted music? I remember one chat about music AIs, the proponent/enthusiast mentioned "creating" a song by using "johnny cash" in the prompt to get a song in that distinctive style.

I will be the first to admit how useful these AIs are, while I don't use them to "create music", I sure as heck have used them for chats and to write programs. I'm not one of those naysayers talking about the outputs being crap because despite some errors here and there, I've found great utility from these things and I don't even mess with "frontier models" from openai or anthropic. My gripes are about the costs about what we're doing here.

When chatting about "fair use", we can go with the legal definitions which are clear as mud, but I think a better and more sensible path forward would be to consider all the ramifications of their use. Even the term "fair use" implies results much better than we're actually seeing, the economic forces in play here do not seem fair at all.

I don't understand your point.

Its my stated disagreement with the following:

I think it makes sense to say training on material you legally acquired is fair use.

First an aside, there's no law against accessing copyrighted material. OpenAI is certainly welcome to read the New York Times. I can legally buy dvds but the legality of redistributing rips is only considered should I redistribute them.

One of the considerations of "fair use" might be the possibility of benefits or drawbacks to society of those uses that fall under "fair use" exceptions. I would hold the AI's use to that standard, and thats really when we should take the entirety into consideration. This is a big topic tho. We want laws because we believe they benefit society, when loopholes appear we'd normally like them closed, obviously there are branches of the US government that are simply not "normal" right now so theres that lol

I will also admit that a narrower view of "fair use" is to equate "training" of these AIs with human use of the material, after all someone reading the new york times are certainly not infringing on anybody's copyright, in fact they're likely fufilling the new york times internally held purpose, ie., they do the writing, they hope to be read with whatever profit to them that might bring. Just like an AI reading that page by controlling a web browser right? Well this is a good rabbit hole too and you're welcome to try to present that case also, because I've been of the opinion for the last few years that LLMs are not people, and I can chat about that all day.

Its easy to post that you don't understand, so if you'd like to post that again this is fine with me, I'm often a very misunderstood individual :)

My position is it's totally coherent to say that "training" is analogous to "reading" and inference is not analogous to "redistributing rips" and that it's a reasonable position for fair use law to operate that way. I think this benefits society.

My position is it's totally coherent to say that "training" is analogous to "reading"

yeah I joke about chatting up my ai gf but there's not going to be any sex lol

and inference is not analogous to "redistributing rips"

I think you made both points previously so duly noted, your explanations are lacking tho.

and that it's a reasonable position for fair use law to operate that way.

I think you said this too in a way. Which is interesting in a coincidental sort of way, I like working with loops. But what I do when producing a loop based techno whatever is I like to progress it somehow. The same thing over and over well I guess thats a thing too which is interesting, if you downloaded a heavily looped song, you can compress it by repeating one bar programmatically, granted its pretty lossy. its sort of how much information per measure of time, I like to call it "rpms", like for instance the first measure has high rpm, but if you loop it, each subsequent measure "loses rpms" so to speak. This doesn't take into consideration "incoherent noise" obviously, I like to take it as one measure amongst many.

Your text compresses nicely with suprisingly little loss of information!

I think this benefits society.

Heh trump thinks energy can be created from nothing.

I appreciate the chat, but while I get the point you're making, I can also appreciate you're not interested in mine which is fair. I post as sort of a hobby, and I think I've sort of thrashed out a blog post maybe here, so thats appreciated.

AboutSource Built by g1lg1l

Hackerly is an independent reader for Hacker News, built on the public HN API. Not affiliated with Y Combinator.