Skip to content

Comment on Who needs to scrape millions of pages, or monitor them?

Comments

Cool !! Is it robots.txt compliant? If not, it might be a good idea to make this available as an option/parameter.

For 'quick and dirty' tasks, wget -r can come in handy too.

Currently it doesn't look at the robots.txt file. I had taken note of this, and will be added in future releases. Thanks for the suggestion.

AboutSource Built by g1lg1l

Hackerly is an independent reader for Hacker News, built on the public HN API. Not affiliated with Y Combinator.