Is there a delay between the time the robots.txt changes and the time when the content becomes inaccessible via Wayback Machine? How often does the archive.org_bot crawl robots.txt?
Can a script check robots.txt periodically for changes and if changes are detected, then download the content from Wayback Machine before it becomes inaccessible?
Additionally, can a script check the domain registration for an anticipated expiration date, or perhaps monitor domainname "drop lists"?
Comments
Is there a delay between the time the robots.txt changes and the time when the content becomes inaccessible via Wayback Machine? How often does the archive.org_bot crawl robots.txt?
Can a script check robots.txt periodically for changes and if changes are detected, then download the content from Wayback Machine before it becomes inaccessible?
Additionally, can a script check the domain registration for an anticipated expiration date, or perhaps monitor domainname "drop lists"?