I understand where you are coming from and agree that this is a massive abuse of authorship rights, but I do not think it should be required to go through this complex process.
We know the network only works thanks to training data. We know training data was scraped without consent. We know the output is derivative of training data. We know OAI/MS get paid for those derivative works and are not in turn paying the producers of the original content. We know it has real impact on those producers and will cost future incentives to publicly share original research, along with other societal implications.
I wonder when large publishers in the US and EU wake up to the fact that their books and films were also scraped the very same way. Unless MS is paying them behind the scenes, they have a corporation profiting off their publications and they are not known to be especially charitable.
Comments
I understand where you are coming from and agree that this is a massive abuse of authorship rights, but I do not think it should be required to go through this complex process.
We know the network only works thanks to training data. We know training data was scraped without consent. We know the output is derivative of training data. We know OAI/MS get paid for those derivative works and are not in turn paying the producers of the original content. We know it has real impact on those producers and will cost future incentives to publicly share original research, along with other societal implications.
I wonder when large publishers in the US and EU wake up to the fact that their books and films were also scraped the very same way. Unless MS is paying them behind the scenes, they have a corporation profiting off their publications and they are not known to be especially charitable.