Skip to content

Comment on Robots.txt Disallow: 20 Years of Mistakes To Avoidparent

Comments

How is IA supposed to distinguish a new website from a sincere wish to delete old stuff? A change in domain registration data means nothing; I have a domain that I registered for an association in my name, and which I then sold to them (for a symbolic price), but it was only an administrative issue - the site was the same.

IA is on iffy territory w.r.t. copyright as it is; if they stop respecting robots.txt, they could get into a world of hurt.

Your last sentence is key. As I understand it, there's no real legal precedent for IA which basically copies everything out there on an opt-out basis. I personally am glad they do but one of the ways they get off with it is by treading as lightly as possible, including respecting robots.txt even retroactively.

They're also non-commercial, broad in scope, arguably serve a valuable scholarly function and have other characteristics that have kept them mostly out of legal hot water. But it's unclear to what degree they're legally different from a site that decided to create an archive of all comics, commercial and otherwise, and slap advertising up.

AboutSource Built by g1lg1l

Hackerly is an independent reader for Hacker News, built on the public HN API. Not affiliated with Y Combinator.