Skip to content

Comment on Moonshot serves Claude instead of Kimi and collects exchanges for model trainingparent

Comments

Just because the material was legally required doesn't mean you can do anything you want with it. I can't (legally) buy a physical book, scan it, and put the scan on my web site. It seems to me that an LLM is a derived work of the training materials that went into it, and thus needs permission from the copyright holders.

But OK, the law seems to disagree with me there. But OK, let's say it's fine for AI companies to train their models on copyrighted content as long as they didn't torrent it or whatever. What then makes it illegal, or morally wrong, to do the same thing with their competitors' model outputs? Why is it OK for Anthropic to scrape this comment and feed it into their system, but not OK for Moonshot to scrape the output of Anthropic's system and feed it into theirs?

you can buy a book, scan it, and upload the counts of every letter, distribution of apostophies, use it as the input to some convoluted process to produce weights or a search index though. They got slapped for illegally obtaining the files, not for producing derivative works of them.

distilling another llm is a clear tos violation but no one really knows how much teeth those have. financially probably none all they can do is whack a mole on the accounts doing it which won’t work.

so they’re trying to lobby copyright changes i guess; unlikely to succeed as doing so would also make all search engines illegal

Not sure what part of my comment this is meant to address. I explicitly acknowledged that the law seems to consider this to be legal. My point is that if it's legal to feed random web sites into the training system, why would it not also be legal to feed competitors' model outputs into it?

Because it’s TOS violation. Anthropic have no agreed tos with the websites they’re scraping. The accounts being used to distill Anthropic’s models are all bound by their terms

That's a contract violation, not an illegal act. And I'd bet that plenty of sites that Anthropic et al have trained on have ToS that forbid using them for model training. This site does. Do we think the AI companies aren't training on HN comments?

AboutSource Built by g1lg1l

Hackerly is an independent reader for Hacker News, built on the public HN API. Not affiliated with Y Combinator.