People keep saying this without attribution. Meta and others got caught, and brought into court, over using torrented files. But those models trained on that data have long since been retired and replaced with new models based on new from-scratch training runs. OpenAI, Anthropic, etc all pay studios, newspapers, Reddit and others for access to data for training. They scrape the open web, but if that's illegal a court hasn't said so. The open web is open, after all. And they don't seem to be stealing books, they seem to be buying physical copies and scanning. Seems legit, that's what a human would do to learn from a book. They also pay big bucks for commercially curated data and training sets.
Just feels like there's enormous CCP effort to put their labs on equal moral footing with everyone else when it's not demonstrably the case. They want the West to hate themselves so we're happy to squander our technological lead.
There really isn't any moral argument against distillation, which is itself pretty goddamn benign. It's pretty simple. Someone pays for Claude access. Claude outputs tokens that are not copyrighted. Then you train on those tokens, which doesn't create a derivative work in the first place.
Is it theft? Well, no. There's no authentication bypass here, no Claude model leak. At best it is violating the terms of use, kind of like how it is violating the terms of use to scrape many websites that AI scrapers scraped.
Is it immoral? Why would it be, exactly? Distillation is not a forbidden technique with moral implications. In fact, there is quite compelling evidence that Anthropic themselves were distilling from OpenAI in early Claude models. It helped them bootstrap if nothing else. There is no special moral code that makes distillation forbidden any more than training off of people's works without permission, or even express non-consent, is forbidden.
Really the more concerning aspect of this is the deception of using Kimi and expecting Kimi output and getting Claude instead, but I would like some independent confirmation that this is even something Moonshot really did before raking them over the coals, rather than just assuming it's true because Anthropic said so. How exactly did they figure out, considering ZDR? It deserves more information.
I do agree that there is a tendency for people to justify CCP human rights violations by trying to equate them to much lesser but similarly shaped transgressions from Western governments, but that's an unrelated issue entirely. The story regarding distillation is consistent: Sorry, but I can't afford enough tiny violins to express my lack of giving a shit. I harbor no ill will, I truly hope the golden parachutes that Sam and Dario fly out on are adorned with the finest materials.
they seem to be buying physical copies and scanning. Seems legit, that's what a human would do to learn from a book
There’s still the open question on learn vs copy/mimic/repeat.
As a human, I can read a book I bought. I’m definitely not allowed to scan it and post its pages online and upload them to an archive of scanned PDFs without the authors’ and publishers’ permission.
IIRC the Meta legal case wasn’t even about LLMs, they just torrented and shared pirated files, whether with strangers or among employees. Those may or may not have been later used for training, but it was already illegal to just share among employees.
I’m definitely not allowed to scan it and post its pages online and upload them to an archive of scanned PDFs without the authors’ and publishers’ permission.
That's explicitly not what they are doing. They are scanning it and then training on the scan. They are allowed to do this in much the same way you are: format shifting for personal use is also allowed (much as the DMCA likes to get in the way with DRM'd media).
I didn't realize Anthropic is a person and doing all of this for his/her/their personal use that is never shared with anyone else, and even more never for financial gain. What a fun hobby. /s
It might still be allowed for other reasons, but "personal use" isn't what they claim in court.
I am not allowed to read a book many times until I memorize it, and later record an audiobook of one of its chapters for money.
While many things might be legal (or hasn't been ruled clearly illegal yet) /under different jurisdictions, we are still in the process of figuring out what we accept as ethical. As the output of models can't be easily copyrighted, destillation is equally disputed. Particularly if the primary model interaction was not destillation (as in this case) IMHO it will be legally quite difficult to restrict secondary use for training. In the end we have to find a legislation and probably even international treaties that account for the fact that classical copyright is beginning to become an obsolete concept.
If American AI labs can learn from the open web why can't Chinese AI labs learn from American ones?? The argument is that learning and distillation are transformative and legal, is it not?? What's good for the goose is good for the gander.
Except they use thousands of stolen credentials which is definitely not legal. They also don’t notify users they send data to US labs which can include sensitive information.
EU labs like mistral cannot legally do any of these things. So people cheerleading China for it is strange.
Just because the material was legally required doesn't mean you can do anything you want with it. I can't (legally) buy a physical book, scan it, and put the scan on my web site. It seems to me that an LLM is a derived work of the training materials that went into it, and thus needs permission from the copyright holders.
But OK, the law seems to disagree with me there. But OK, let's say it's fine for AI companies to train their models on copyrighted content as long as they didn't torrent it or whatever. What then makes it illegal, or morally wrong, to do the same thing with their competitors' model outputs? Why is it OK for Anthropic to scrape this comment and feed it into their system, but not OK for Moonshot to scrape the output of Anthropic's system and feed it into theirs?
you can buy a book, scan it, and upload the counts of every letter, distribution of apostophies, use it as the input to some convoluted process to produce weights or a search index though. They got slapped for illegally obtaining the files, not for producing derivative works of them.
distilling another llm is a clear tos violation but no one really knows how much teeth those have. financially probably none all they can do is whack a mole on the accounts doing it which won’t work.
so they’re trying to lobby copyright changes i guess; unlikely to succeed as doing so would also make all search engines illegal
Not sure what part of my comment this is meant to address. I explicitly acknowledged that the law seems to consider this to be legal. My point is that if it's legal to feed random web sites into the training system, why would it not also be legal to feed competitors' model outputs into it?
Because it’s TOS violation. Anthropic have no agreed tos with the websites they’re scraping. The accounts being used to distill Anthropic’s models are all bound by their terms
That's a contract violation, not an illegal act. And I'd bet that plenty of sites that Anthropic et al have trained on have ToS that forbid using them for model training. This site does. Do we think the AI companies aren't training on HN comments?
Comments
People keep saying this without attribution. Meta and others got caught, and brought into court, over using torrented files. But those models trained on that data have long since been retired and replaced with new models based on new from-scratch training runs. OpenAI, Anthropic, etc all pay studios, newspapers, Reddit and others for access to data for training. They scrape the open web, but if that's illegal a court hasn't said so. The open web is open, after all. And they don't seem to be stealing books, they seem to be buying physical copies and scanning. Seems legit, that's what a human would do to learn from a book. They also pay big bucks for commercially curated data and training sets.
Just feels like there's enormous CCP effort to put their labs on equal moral footing with everyone else when it's not demonstrably the case. They want the West to hate themselves so we're happy to squander our technological lead.
There really isn't any moral argument against distillation, which is itself pretty goddamn benign. It's pretty simple. Someone pays for Claude access. Claude outputs tokens that are not copyrighted. Then you train on those tokens, which doesn't create a derivative work in the first place.
Is it theft? Well, no. There's no authentication bypass here, no Claude model leak. At best it is violating the terms of use, kind of like how it is violating the terms of use to scrape many websites that AI scrapers scraped.
Is it immoral? Why would it be, exactly? Distillation is not a forbidden technique with moral implications. In fact, there is quite compelling evidence that Anthropic themselves were distilling from OpenAI in early Claude models. It helped them bootstrap if nothing else. There is no special moral code that makes distillation forbidden any more than training off of people's works without permission, or even express non-consent, is forbidden.
Really the more concerning aspect of this is the deception of using Kimi and expecting Kimi output and getting Claude instead, but I would like some independent confirmation that this is even something Moonshot really did before raking them over the coals, rather than just assuming it's true because Anthropic said so. How exactly did they figure out, considering ZDR? It deserves more information.
I do agree that there is a tendency for people to justify CCP human rights violations by trying to equate them to much lesser but similarly shaped transgressions from Western governments, but that's an unrelated issue entirely. The story regarding distillation is consistent: Sorry, but I can't afford enough tiny violins to express my lack of giving a shit. I harbor no ill will, I truly hope the golden parachutes that Sam and Dario fly out on are adorned with the finest materials.
There’s still the open question on learn vs copy/mimic/repeat.
As a human, I can read a book I bought. I’m definitely not allowed to scan it and post its pages online and upload them to an archive of scanned PDFs without the authors’ and publishers’ permission.
IIRC the Meta legal case wasn’t even about LLMs, they just torrented and shared pirated files, whether with strangers or among employees. Those may or may not have been later used for training, but it was already illegal to just share among employees.
The topic was training. You cropped text to change it. Maybe consider integrity.
Which part from the cropped text do you think changes anything?
When you talk about this part:
That's explicitly not what they are doing. They are scanning it and then training on the scan. They are allowed to do this in much the same way you are: format shifting for personal use is also allowed (much as the DMCA likes to get in the way with DRM'd media).
I didn't realize Anthropic is a person and doing all of this for his/her/their personal use that is never shared with anyone else, and even more never for financial gain. What a fun hobby. /s
It might still be allowed for other reasons, but "personal use" isn't what they claim in court.
I am not allowed to read a book many times until I memorize it, and later record an audiobook of one of its chapters for money.
Cool, but that's also not what they're doing. They are claiming the same rights you have, the main inequality is the resources to defend it in court.
While many things might be legal (or hasn't been ruled clearly illegal yet) /under different jurisdictions, we are still in the process of figuring out what we accept as ethical. As the output of models can't be easily copyrighted, destillation is equally disputed. Particularly if the primary model interaction was not destillation (as in this case) IMHO it will be legally quite difficult to restrict secondary use for training. In the end we have to find a legislation and probably even international treaties that account for the fact that classical copyright is beginning to become an obsolete concept.
If American AI labs can learn from the open web why can't Chinese AI labs learn from American ones?? The argument is that learning and distillation are transformative and legal, is it not?? What's good for the goose is good for the gander.
Except they use thousands of stolen credentials which is definitely not legal. They also don’t notify users they send data to US labs which can include sensitive information.
EU labs like mistral cannot legally do any of these things. So people cheerleading China for it is strange.
Just because the material was legally required doesn't mean you can do anything you want with it. I can't (legally) buy a physical book, scan it, and put the scan on my web site. It seems to me that an LLM is a derived work of the training materials that went into it, and thus needs permission from the copyright holders.
But OK, the law seems to disagree with me there. But OK, let's say it's fine for AI companies to train their models on copyrighted content as long as they didn't torrent it or whatever. What then makes it illegal, or morally wrong, to do the same thing with their competitors' model outputs? Why is it OK for Anthropic to scrape this comment and feed it into their system, but not OK for Moonshot to scrape the output of Anthropic's system and feed it into theirs?
you can buy a book, scan it, and upload the counts of every letter, distribution of apostophies, use it as the input to some convoluted process to produce weights or a search index though. They got slapped for illegally obtaining the files, not for producing derivative works of them.
distilling another llm is a clear tos violation but no one really knows how much teeth those have. financially probably none all they can do is whack a mole on the accounts doing it which won’t work.
so they’re trying to lobby copyright changes i guess; unlikely to succeed as doing so would also make all search engines illegal
Not sure what part of my comment this is meant to address. I explicitly acknowledged that the law seems to consider this to be legal. My point is that if it's legal to feed random web sites into the training system, why would it not also be legal to feed competitors' model outputs into it?
Because it’s TOS violation. Anthropic have no agreed tos with the websites they’re scraping. The accounts being used to distill Anthropic’s models are all bound by their terms
That's a contract violation, not an illegal act. And I'd bet that plenty of sites that Anthropic et al have trained on have ToS that forbid using them for model training. This site does. Do we think the AI companies aren't training on HN comments?