Skip to content

Comment on I fear for the unauthenticated web

Comments

I would think all you need to do is add a copyright statement of some kind.

Sad things are getting to this point. Maybe I should add this to my site :)

(c) Copyright (my email), if used for any form of LLM processing, you must contact me and pay 1000USD per word from my site for each use.

The argument the AI companies are making is that training for LLMs is fair use which means a copyright statement means fuck all from their point of view. (Even if it does, assuming you're in the US, unless you register the copyright with the US copyright office, you can only sue for actual damages, which means the cost of filing a lawsuit against them--not even litigating, just the court fee for saying "I have a lawsuit"--would be more expensive than anything you could recover. Even if you did register and sued for statutory damages, the cost of litigation would probably exceed the recovery you could expect.)

Of course, the big AI companies are already trying to get the government to codify AI training as fair use and sidestep the litigation which doesn't seem to be going entirely their way on this matter (cf. https://arstechnica.com/google/2025/03/google-agrees-with-op...).

In addition, we need to start paying attention to the growing legislation about AI and copyright law. There was an article on HN I think this week (or last) specifically where a judge ruled AI cannot own copyright on its generated materials.

IANAL, but I do wonder how this ruling will be used as a point of reference whenever we finally ask the question "Does material produced by GenAI violate copyright laws?" Specifically if it cannot claim ownership, a right that we've awarded to trees and monkeys, how does it operate within ownership laws?

And don't even get me ranting about HUMAN digital rights or Personified AIs.

Fair use requires transformation. LLM is as transformative as it gets. If I'm on the jury, you're going to have to make new copyright law for me to convict.

I am personally happy to have everyone, people and LLM alike, learn from my wisdom.

Fair use requires transformation.

No, it doesn't. There are four factors for fair use, and whether the use is transformative is part of one of them. And you don't need to win on all four factors.

LLM is as transformative as it gets.

The current ruling precedent for "transformative" is the Warhol decision, which effectively says that to look at whether or not something is transformative, you kind of have to start by analyzing its impact on the market (and if you're going "doesn't that import the fourth factor into the first?" the answer is "yes, I don't like it, but it's what SCOTUS said"). By that definition, LLMs are nowhere near "transformative."

Even pre-Warhol, their role as "transformative" is sketchy, because you have to remember that this is using its legal definition, not its colloquial definition.

If I'm on the jury

Fortunately, for this kind of question, the jury isn't going to be involved in determining fair use, so it doesn't matter what you think.

That's untrue. See my comment elsewhere around here. It doesn't rely on the commercial aspect, though if it's not commercial the bar for fair use is set lower.

The argument in Warhol relies on the fact that the derivative work, ie, Warhol's painting, is substantially similar in function to the original photograph. If Warhol had used the picture as stuffing for a soft sculpture, it would not infringe.

LLM is closer to the latter than the former.

"so it doesn't matter what you think"

A perfectly fine, if incorrect reply, but then you have to be a dick. Why?

Copyright is for topics like redistribution of the source material. You can’t add arbitrary terms to a copyright claim that go beyond what copyright law supports.

I think you’re confusing copyright with a EULA. You would need users to agree to the EULA terms before viewing the material. You can’t hide contractual obligations in the footer of your website and call it copyright.

What about if my index says "This are the EULA, by clicking "Next" or "Enter", you are accepting them", and a LLM scrapper "clicks" Next to fetch the rest of the content?

That's how the big software companies have been doing it to us for years, so it does seem like turnabout would be fair play.

"Ah, but for us, it's unenforceable, it's for you little people."

"Copyright? Well if you are a big label, we probably need to talk. Little people? Oh fuck you, just give us your money and creative output."

It's reasonably likely, but not yet settled, that LLM training falls under fair use and doesn't require a license. This is what the https://githubcopilotlitigation.com/ class action (from 2022) is about, and its still making its way through the court. This prediction market has it at 12% likely to succeed, suggesting that courts will not agree with you: https://manifold.markets/JeffKaufman/will-the-github-copilot...

It's reasonably likely, but not yet settled, that LLM training falls under fair use and doesn't require a license.

I would say it's not reasonably likely that LLM training is fair use. Because I've read the most recent SCOTUS decision on fair use (Warhol), and enough other decisions on fair use, to understand that the primary (and nearly only, in practice) factor is the effect on the market for the original. And AI companies seem to be going out of their way to emphasize that LLM training is only going to destroy the market for the originals, which weighs against fair use. Not to mention the existence of deals licensing content for LLM training which... basically concedes the point.

Of the various options, a ruling that LLM training is fair use I find the least likely. More likely is either that LLM training is not fair use, that LLM training is not infringing in the first place, or that the plaintiffs can't prove that the LLM infringed their work.

I do not read it that way at all. The Goldsmith decision mainly turns on the idea that an artist protections include that for derivative works. Warhol produced a work that does substantially the same things as Goldsmith's, ie, is a picture that can be viewed.

When talking about parody, they note that the usage as the foundation for parody is always substantially different from the original and thereby allowed, even if it would otherwise infringe. LLMs are always substantially different from the original, too.

If I want to write software that draws that picture exactly, the code would not be a copyright violation. It is text and cannot be printed in a magazine as a picture. If I used it to print a picture that was a derivative work and sold that, it might be.

A large language model has no intersection with the picture or, for that matter, anything that it absorbs. It is possible that someone might figure out how to prompt it to do exactly the same picture as Goldsmith did but fairly unlikely.

Unless you could show that this was easy, common and part of the intent of the LLM creator, I can see no possibility that it is infringing.

This prediction market has it at 12% likely to succeed

Randos on the internet with a betting addiction are distinctively different from a court of law. I wish people would stop talking about prediction market as if they mattered.

Participants in prediction market do not need to be experts for their collective input to be informative.

There's a long history of economic research on the "wisdom of crowds" that backs up their value.

this isn't about copyright but about computer access. the CFAA is extremely broad; if you ban LLM companies from access on grounds of purpose you have every legal right to do so

in theory that legislation has teeth, too. they are not allowed to access your system if you say they are not; authentication is irrelevant.

every GET request to a system that doesn't permit access for training data is a felony

why are we pretending that these gambling sites have any weight on anything

What do you mean by weights?

I'd certainly trust their predictions more than those given by most "experts".

Such a notice is legally meaningless, though. Doubly so if the courts rule that scraping for AI purposes counts as fair use.

This is pretty naive.

The only reason copyright is so strong in the US is that there are big players (Disney, Elsevier) who benefit from it. But gig tech is much bigger, and LLMs have created a situation where big tech has a vested interest in eroding copyright law. Both sides are gearing up for a war in the court systems, and it's definitely not a given who will win. But, if you try to enter the fray as an individual or small company, you definitely aren't going to win.

The reality is that a lot of these small websites have very permissive licenses. I really hope we don't get to the point where we must all make our licenses stricter.

The reality is that none of these LLM scrapers give a damn about copyright, because the entire AI industry is built on flagrant copyright violation, and the premise that they can be stopped by a magic string is laughable.

You could sue, if you can afford it, meanwhile all of your data is already training their models.

A class action, funded by their rivals could hurt quite a bit, especially for sites damaged monetarily by these LLM scrapers.

Sure, because Meta certainly followed copyright law to the letter when they torrented thousands of copyrighted books from hundreds of published and known authors to train Lama. Forgive me if I doubt a text disclaimer on the page will slow them down.

Unfortunately copyright is no limit to these companies.

Meta is stating in court that knowingly downloading pirated content is perfectly fine (ref https://news.ycombinator.com/item?id=43125840) so they for one would have absolutely no issue completely ignoring your copyright notice and stated licensing costs. Good luck affording a legal team to try force them to pay attention.

Copyright is something for them to beat us with, not the other way around, apparently.

AboutSource Built by g1lg1l

Hackerly is an independent reader for Hacker News, built on the public HN API. Not affiliated with Y Combinator.