you can buy a book, scan it, and upload the counts of every letter, distribution of apostophies, use it as the input to some convoluted process to produce weights or a search index though. They got slapped for illegally obtaining the files, not for producing derivative works of them.
distilling another llm is a clear tos violation but no one really knows how much teeth those have. financially probably none all they can do is whack a mole on the accounts doing it which won’t work.
so they’re trying to lobby copyright changes i guess; unlikely to succeed as doing so would also make all search engines illegal
Not sure what part of my comment this is meant to address. I explicitly acknowledged that the law seems to consider this to be legal. My point is that if it's legal to feed random web sites into the training system, why would it not also be legal to feed competitors' model outputs into it?
Because it’s TOS violation. Anthropic have no agreed tos with the websites they’re scraping. The accounts being used to distill Anthropic’s models are all bound by their terms
That's a contract violation, not an illegal act. And I'd bet that plenty of sites that Anthropic et al have trained on have ToS that forbid using them for model training. This site does. Do we think the AI companies aren't training on HN comments?
Comments
you can buy a book, scan it, and upload the counts of every letter, distribution of apostophies, use it as the input to some convoluted process to produce weights or a search index though. They got slapped for illegally obtaining the files, not for producing derivative works of them.
distilling another llm is a clear tos violation but no one really knows how much teeth those have. financially probably none all they can do is whack a mole on the accounts doing it which won’t work.
so they’re trying to lobby copyright changes i guess; unlikely to succeed as doing so would also make all search engines illegal
Not sure what part of my comment this is meant to address. I explicitly acknowledged that the law seems to consider this to be legal. My point is that if it's legal to feed random web sites into the training system, why would it not also be legal to feed competitors' model outputs into it?
Because it’s TOS violation. Anthropic have no agreed tos with the websites they’re scraping. The accounts being used to distill Anthropic’s models are all bound by their terms
That's a contract violation, not an illegal act. And I'd bet that plenty of sites that Anthropic et al have trained on have ToS that forbid using them for model training. This site does. Do we think the AI companies aren't training on HN comments?