Skip to content

Comment on Google penalizes original content site because of scrapers

Comments

The headline is factually inaccurate. It looks like a mistake that he was denied access to Adwords due to original content, but that has no connection to his ranking in search. It looks like the adwords representative may not have access to fine-grained enough tools to assess the site accurately, which is an organizational failing, but there's no bad intent there. In some sense it reflects how disconnected search and ads are from each other that they're using crude tools to assess original content.

Google cares a great deal about putting the original source of a piece of content first. If we're doing that incorrectly, it's because we screwed up, not because that's how things are designed. It's a hard problem and an area we are still working on intensely. It would be great if someone involved could post the queries on which we are screwing up so we can debug what's causing it.

You don't seem to have read the article.

The search:

["a superb app for iPad and iPhone that lets you quickly and easily transfer photos and videos between iOS devices and computers – has been updated this week, to Version 2.3."]

returned results from content scrapers above the original content.

For me, the original content doesn't even show up in search results, even though it's in Google's index:

http://ipadinsight.com/ipad-apps/photo-transfer-app-updateda...

Further, Google wouldn't let the site owner buy AdWords to drive traffic to their site.

Google owns both search and AdWords. This makes the headline here:

  Google penalizes original content site because of scrapers
accurate, as far as I'm concerned.

The results for that query are horrible, but his complaint about losing traffic certainly isn't due to his ranking on 30-word quoted queries. I'm hoping to get an example of a normal query where he's losing out to scrapers so we can debug what's going on.

The headline implies that he's penalized in search due to scrapers, which isn't happening.

> The headline implies that he's penalized in search due to scrapers, which isn't happening.

In this case the phrase "Google penalizes" = "Google denies access to adwords". The word penalize doesn't always refer to site penalties in the Google search index.

The reason it's on top of HN is almost certainly because that's the conclusion everyone jumped to. That seems to be how the word "penalize" is consistently used with respect to Google.

I agree, I am a bit surprised this is on the frontpage. What is ironic though, is than now when I search for the 30 word phrase, two scraper sites appear with the article scraped from seobook.com. ...nicely backdated by 37 seconds.

I am not sure what you mean by the headline is inaccurate. It would be more clear if it said, "Google penalizes an original content site because of scrapers", but the meaning is the same. You seem to agree that an original content site can get penalized if Google thinks the scrapers are the original source of the content.

I personally don't see what is news about this. It has been known for a long time that newer or less frequently updated sites can get beaten by scrapers, though it usually resolves itself later unless the scraper is a decently reputable source like the Huffington Post.

He was denied access to adwords due to a mistaken impression that he's a scraper. He did not lose search traffic due to that.

The headline conflates the two things, so is inaccurate.

I've never seen an actual instance of a site with original content being penalized because it is getting scraped (though it is theoretically possible.) Our systems for this are robust and quite conservative. When a scraper outranks the original site it's because we weren't aggressive enough in demoting the scraper, or don't have enough data about it, not because the original was penalized.

...and how is being denied access to AdWords over mistaken identity not penalization? If you get thrown into jail because of mistaken identity, that doesn't make the situation any fairer or just for you. You're still being punished.

Here is one that ranks the original content bottom of the first page. The first two results are exact replicas of each other and a copy of original article. The query is for the title of his most recent post.

https://encrypted.google.com/search?q=Quick+Look+%E2%80%93+W...

I admit I struggled to come up with a headline for it and my choice of wording was not very good. However, he did lose search traffic as a scraper site took over his top ranking for certain searches.

It appears he has more details about specific queries on this page: http://www.google.com/support/forum/p/AdWords/thread?tid=0bb...

A good headline would have been "Google rejects an original content site from Adwords due to scrapers." He may have a quite legitimate complaint, but it doesn't seem to involve search.

I sent the site owner an email to see if he can comment on more general search queries he may have lost ranking on. If it is just that long quoted one I agree that your suggested headline would be better.

AboutSource Built by g1lg1l

Hackerly is an independent reader for Hacker News, built on the public HN API. Not affiliated with Y Combinator.