Skip to content

Comment on Ask HN: What are the viable alternatives to DuckDuckGo?parent

Comments

Wouldn't you do better say indexing Wikipedia, then what it links directly, then what those link directly?

As you note, you got into an adult content 'bubble'. Going to random sites doesn't mean that site is good in anyway. Wikipedia, or some other proxy for quality at least gives you a good baseline.

If you're aiming to index all of the web, you should have the same endpoint, but in the meantime my way, you are indexing 'good' sites

My crawling approach did hit Wikipedia several times, and crawled a number of pages, but not the entire site at once.

With either approach, we would both eventually end up with all the websites - that being a surprisingly small number (~1 billion of which we only ever access ~1000).

I got in an adult content bubble from starting with "A" ("Adult") and would have hit another one presumably at "X". But "B" had quite a few too, lol.

Most of the words were bubbles - where it spent a decent amount of time on each one and I wondered if it would ever move on or was stuck. There are just a lot of backlinks and tangents for every word you can think of.

Another fun thing about this experiment is you will find a lot more international websites that no other search engine will find - some weird personal sites and some surprisingly good forums.

But yeah, you could enter anywhere - Wikipedia, or some other random Array of words. It doesn't have to be a dictionary.

AboutSource Built by g1lg1l

Hackerly is an independent reader for Hacker News, built on the public HN API. Not affiliated with Y Combinator.