Not if you told them first - if popular CMSs/blog platforms notified Google when you clicked published they'd know just before the page even existed on the internet. Other content producers would rush to integrate it as well.
I believe they already ping Google to inform them of new content, it would be a natural extension of that.
So how about - content creator pushes content to Google, and list of people they think will scrape it. If sufficiently large sufficiently close copy of content shows up on scraper page they're penalized (adwords, search, banned from Blogger whatever).
Edit: I guess that only works if they're fixed IPs.
You don't need the list of potential scrapers, it could apply to anyone who published the same content after.
The issue though is if you write a blog and don't know to do this, the scrapers do know and nick your content and inform Google of "their" new article. At that point you're now seen as a scraper for an article you actually wrote because the real scraper plays the game better.
And that's the real problem - the scrapers play the game better than the rest of us because they can commit larger amounts of time to it (not having to waste time writing actual content and all).
What I was thinking is if you have a list of scrapers supplied and a publish notification, Google could then index the alleged scraper sites directly (instead of waiting for a general index to find them), and then at some interval later. For true scraper, this should result in no match followed by match. Then Google can establish order of publishing, assuming Googles push-notification-to-index time is shorter than the scraper-polling-interval.
And it seems people often know of their regular scrapers.
Comments
Not necessarily. That would only help if everyone used it. Otherwise it'd be another tool for the scrapers to claim legitimacy.
Not if you told them first - if popular CMSs/blog platforms notified Google when you clicked published they'd know just before the page even existed on the internet. Other content producers would rush to integrate it as well.
I believe they already ping Google to inform them of new content, it would be a natural extension of that.
So how about - content creator pushes content to Google, and list of people they think will scrape it. If sufficiently large sufficiently close copy of content shows up on scraper page they're penalized (adwords, search, banned from Blogger whatever).
Edit: I guess that only works if they're fixed IPs.
You don't need the list of potential scrapers, it could apply to anyone who published the same content after.
The issue though is if you write a blog and don't know to do this, the scrapers do know and nick your content and inform Google of "their" new article. At that point you're now seen as a scraper for an article you actually wrote because the real scraper plays the game better.
And that's the real problem - the scrapers play the game better than the rest of us because they can commit larger amounts of time to it (not having to waste time writing actual content and all).
What I was thinking is if you have a list of scrapers supplied and a publish notification, Google could then index the alleged scraper sites directly (instead of waiting for a general index to find them), and then at some interval later. For true scraper, this should result in no match followed by match. Then Google can establish order of publishing, assuming Googles push-notification-to-index time is shorter than the scraper-polling-interval. And it seems people often know of their regular scrapers.
They might be able to claim it but at least it'll Also give the content creator a tool to deny that claim.