Skip to content

Comment on Wouldn't it be fun to build your own Google?

Comments

A really nice idea.

I volunteered a bit early this year for Common Crawl (not much, just some Java and Clojure examples for fetching and using the new archive format).

Common Crawl already has many volunteers (and a professional management and technical staff) so it would seem like a good idea to merge some of the author's goals with the existing Common Crawl organization. Perhaps more frequent Common Crawl web fetches and also making the data available on Azure, Google Cloud, etc. would satisfy the author's desire to have more immediacy and have the data available from multiple sources.

some edits:

Most of the Common Crawl data is on Microsoft Azure, but not all of it.

The Common Crawl is a great resource that deserves attention from more companies and developers.

AboutSource Built by g1lg1l

Hackerly is an independent reader for Hacker News, built on the public HN API. Not affiliated with Y Combinator.