There have been collaborative computing projects like SETI@home [1] and Folding@Home [2] where unused computing power could be used for productive purposes. Could there be something similar for storage? Software that provides unused storage for Internet archiving? In the best case scenario, we could have redundant backups of the Internet Archive distributed around the world.
archive.org does use torrents and I have one such torrent laying around in my client, which occasionally connects to peers although the the trackers are currently offline. I suppose a new client would find me and other peers through the DHT. I'd share a magnet link for someone to try, but it's a copyright-ignoring ROM dump archive so it may not be the best idea to post it here.
It's interesting that torrents may not be the first thing that comes to mind. They have the "PR issue" of being the now seemingly mundane way in which we've been downloading DVD rips for the last 20 odd years. Newer technology like IPFS does a better job making the cool core of this technology actually sound cool.
I didn’t know that archive.org already has torrents. I guess what we would need, on top of that, is a system for assigning those torrents to new peers.
That already exists. Peers find each other through a distributed hash table which can be bootstrapped from a variety of sources.
I would say the problem is discoverability and actual deployment.
For suddenly popular files it can be a way to donate bandwidth most of all, because then there might suddenly be a lot of peers. For the vast majority of files however, there won't be any other peers and they have to be web seeded by Archive.org either way.
Then there's the discoverability problem. Ultimately you need something like a magnet link to connect to a swarm.
IPFS on the tin seems pretty awesome however when I attempted to dig into it for an hour I still had no idea how to actually do anything with it. Their usability needs to go a long way before I give it another go. it's definitely not a two step process where you download a client and click on a link to start load sharing an archive in my past experience.
I really wish the EU would have their own organisation for creating an internet archive that at the very minimum mirrored IA. This is our history and there's only a single place now that has any significant archive of it. It seems like the EU should have a significant interest in preserving it for generations to come.
Demonising it is more fashionable right now. Everyone I know that contributes to the Internet archive is right of Center enough to be considered a horrible person.
The INTERNETARCHIVE.BAK experiment has come to a close a number of years ago.
Much was learned in the process, and many thanks are given to the dozens of people who donated time, space and coding efforts to make the system work as long as it did. A number of useful facts and observations came from the project.
The Internet Archive continues to explore methods and code to decentralize the collection, to have a mirror running in various ways - these include IPFS, FileCoin, and others. The INTERNETARCHIVE.BAK project also added general mirroring and tracking code to a number of projects that are still in use.
IA called this their Postmortem, but it sounds... intentionally opaque. Also, I'm not sure if this website is affiliated with archive.org, since they say at the bottom of their homepage:
Archive Team is in no way affiliated with the fine folks at ARCHIVE.ORG Archive Team can always be reached by e-mail at archiveteam@archiveteam.org or by IRC at the channel #archiveteam (on hackint).
Yeah, it’s a completely [1] separate team (they do run a bunch of archiving projects that end up in the IA / Wayback Machine though). Just wanted to share – it’s sad there isn’t much more info though apart from some code; maybe worth looking into IRC logs?
[1]: On paper, at least; the founder, Jason Scott, seems pretty involved with the IA as well, and I’m not really sure how much the teams intersect.
Comments
This incident brings up a good point: Who archives the archives?
There have been collaborative computing projects like SETI@home [1] and Folding@Home [2] where unused computing power could be used for productive purposes. Could there be something similar for storage? Software that provides unused storage for Internet archiving? In the best case scenario, we could have redundant backups of the Internet Archive distributed around the world.
[1]: https://setiathome.berkeley.edu/
[2]: https://foldingathome.org/
Perhaps torrents?
archive.org does use torrents and I have one such torrent laying around in my client, which occasionally connects to peers although the the trackers are currently offline. I suppose a new client would find me and other peers through the DHT. I'd share a magnet link for someone to try, but it's a copyright-ignoring ROM dump archive so it may not be the best idea to post it here.
It's interesting that torrents may not be the first thing that comes to mind. They have the "PR issue" of being the now seemingly mundane way in which we've been downloading DVD rips for the last 20 odd years. Newer technology like IPFS does a better job making the cool core of this technology actually sound cool.
I didn’t know that archive.org already has torrents. I guess what we would need, on top of that, is a system for assigning those torrents to new peers.
That already exists. Peers find each other through a distributed hash table which can be bootstrapped from a variety of sources.
I would say the problem is discoverability and actual deployment.
For suddenly popular files it can be a way to donate bandwidth most of all, because then there might suddenly be a lot of peers. For the vast majority of files however, there won't be any other peers and they have to be web seeded by Archive.org either way.
Then there's the discoverability problem. Ultimately you need something like a magnet link to connect to a swarm.
https://github.com/anacrolix/btlink
The vision behind IPFS is that (to an extent) https://ipfs.tech/
IPFS on the tin seems pretty awesome however when I attempted to dig into it for an hour I still had no idea how to actually do anything with it. Their usability needs to go a long way before I give it another go. it's definitely not a two step process where you download a client and click on a link to start load sharing an archive in my past experience.
https://github.com/anacrolix/btlink
There is currently ArchiveTeam going on
I really wish the EU would have their own organisation for creating an internet archive that at the very minimum mirrored IA. This is our history and there's only a single place now that has any significant archive of it. It seems like the EU should have a significant interest in preserving it for generations to come.
Demonising it is more fashionable right now. Everyone I know that contributes to the Internet archive is right of Center enough to be considered a horrible person.
r/DataHoarder
(or r/archiveteam ?)
Personally, I have archived a few of the magazine collections.
https://archiveteam.org/index.php/IA.BAK
IA called this their Postmortem, but it sounds... intentionally opaque. Also, I'm not sure if this website is affiliated with archive.org, since they say at the bottom of their homepage:
Yeah, it’s a completely [1] separate team (they do run a bunch of archiving projects that end up in the IA / Wayback Machine though). Just wanted to share – it’s sad there isn’t much more info though apart from some code; maybe worth looking into IRC logs?
[1]: On paper, at least; the founder, Jason Scott, seems pretty involved with the IA as well, and I’m not really sure how much the teams intersect.
The co-founder, Jason Scott, retired from Archive Team years ago and stays around as a cheeleader and advisor. He is employed by Internet Archive.
Must be a busy guy, fancy seeing him here. (Thanks for all the great work!)