The most interesting conspiracy theory in this lawsuit is that Scribd might be doing the whole PDF-hosting business as a loss leader in order to enable an automatic infringement-detection business.
In other words, Scribd's document-hosting is sort of a honeypot, designed to attract both legit and infringing uploads, as well as complaints from publishers regarding infringing uploads. Collecting all of this, they have both a document hosting site and a database of stuff publishers have claimed as copyrighted.
You couldn't very well call up every publisher and say "hey, send me a PDF of everything you have a copyright on so I can start a business doing automatic copyright-checking." But if you have a popular document hosting site, you can rely on the publishers to incrementally develop that database for you. And your business goals line up very well with your DMCA legal requirements.
This would be a very clever hack. I doubt Scribd chose to pursue it intentionally, but either way, an "Is this blob of text copyrighted by someone else?" API would be a very valuable service.
Comments
The most interesting conspiracy theory in this lawsuit is that Scribd might be doing the whole PDF-hosting business as a loss leader in order to enable an automatic infringement-detection business.
In other words, Scribd's document-hosting is sort of a honeypot, designed to attract both legit and infringing uploads, as well as complaints from publishers regarding infringing uploads. Collecting all of this, they have both a document hosting site and a database of stuff publishers have claimed as copyrighted.
You couldn't very well call up every publisher and say "hey, send me a PDF of everything you have a copyright on so I can start a business doing automatic copyright-checking." But if you have a popular document hosting site, you can rely on the publishers to incrementally develop that database for you. And your business goals line up very well with your DMCA legal requirements.
This would be a very clever hack. I doubt Scribd chose to pursue it intentionally, but either way, an "Is this blob of text copyrighted by someone else?" API would be a very valuable service.