Comment on Who needs to scrape millions of pages, or monitor them?Comments−webvet13yCool !! Is it robots.txt compliant? If not, it might be a good idea to make this available as an option/parameter.For 'quick and dirty' tasks, wget -r can come in handy too.−calufaOP13yCurrently it doesn't look at the robots.txt file. I had taken note of this, and will be added in future releases. Thanks for the suggestion.
Comments
Cool !! Is it robots.txt compliant? If not, it might be a good idea to make this available as an option/parameter.
For 'quick and dirty' tasks, wget -r can come in handy too.
Currently it doesn't look at the robots.txt file. I had taken note of this, and will be added in future releases. Thanks for the suggestion.