Skip to content

Comment on Matt Cutts is looking for scraper sites

Comments

I wonder how Google chooses which Wikipedia articles they scrape and which ones they don't.

In testing, they definitely don't seem to scrape every article:

http://i.imgur.com/ujDqZhB.png

This is a good question...I've long since surmised that Google has a set of heuristics for every site that has an API that allows for easy domain-specific ranking. With Wikipedia, you have number of page edits, frequency of page edits, and (to an extent) quality of recent page edits. StackOverflow provides an even easier metric for what's considered high quality, and Google appears to apply its own layer on top of that (and in my non-scientific perception, looking something up by Google is almost always more fruitful on the first search than by going directly to SO)

AboutSource Built by g1lg1l

Hackerly is an independent reader for Hacker News, built on the public HN API. Not affiliated with Y Combinator.