Skip to content

Comment on Internet Archive is now a federal depository libraryparent

Comments

If anyone reading knows an easy way to download and mirror IA pages please make it easier to find. A bot told me they offer downloads of the underlying WARC files but I could not find it

A bot told me they offer downloads of the underlying WARC files but I could not find it

The "bot" is wrong. Most of the crawl data used by the Internet Archive, particularly the Alexa crawls, isn't publicly accessible. (This is because some of it includes archived pages which have since been suppressed by the site owner - removing those pages from the archived crawl data isn't practical.)

https://archive.org/details/alexacrawls

Common Crawl data is public, but less comprehensive than IA - https://commoncrawl.org/

There are utilities to help, waybackpack comes to mind, but I haven't looked in a while. https://github.com/jsvine/waybackpack

I used wayback-machine-downloader, I think you need one of the forks to make it work though.

https://github.com/hartator/wayback-machine-downloader

They locked away most .warc files due to the AI harvesting crunch.

AboutSource Built by g1lg1l

Hackerly is an independent reader for Hacker News, built on the public HN API. Not affiliated with Y Combinator.