One of the questions I've had to really grapple with is that I just cannot personally host all the terabytes of data I'm collecting.
Hosting is a big challenge for open science. Whenever I see someone calling for open data, I feel half enthusiasm for the movement to cooperate more and half dread at the prospect of trying to comply. I used OSF to publish one of my datasets, but they instituted storage limits that would prevent doing the same in the future. And exposure to thousands of dollars in surprise fees from personal archival in cloud hosts isn't acceptable. Which leaves the status quo of email the author for data, hope they respond, as disappointingly the best option.
Open code is trivial to provide in comparison, but also less useful.
Sequence read archive has been such a boon for developing reproducible biological pipelines without having to worry about data. A paper references a dataset by ID and I can use it as input for my pipelines and keep the raw data locally only for as long as its needed to generate analysis within the running pipeline. I can even set threshold levels of how much local or cloud compute resources should be used at a time if I didn't want to exhaust my systems with one job.
Not for these size datasets. Torrents are just too small and unreliable, except for the very most popular items. Slightly out of mainstream movies, which are small and likely more popular than data for an obscure science experiment, are nearly impossible to find seeders for.
So I'd guess no seeders want to host tens to thousands of terabyte torrents, and then thousands to millions of those for all the different datasets grabbed by all the science projects all over the world. The odds of being able to download one of these on demand is just about zero.
Comments
Hosting is a big challenge for open science. Whenever I see someone calling for open data, I feel half enthusiasm for the movement to cooperate more and half dread at the prospect of trying to comply. I used OSF to publish one of my datasets, but they instituted storage limits that would prevent doing the same in the future. And exposure to thousands of dollars in surprise fees from personal archival in cloud hosts isn't acceptable. Which leaves the status quo of email the author for data, hope they respond, as disappointingly the best option.
Open code is trivial to provide in comparison, but also less useful.
Sequence read archive has been such a boon for developing reproducible biological pipelines without having to worry about data. A paper references a dataset by ID and I can use it as input for my pipelines and keep the raw data locally only for as long as its needed to generate analysis within the running pipeline. I can even set threshold levels of how much local or cloud compute resources should be used at a time if I didn't want to exhaust my systems with one job.
Do people ever make a torrent of their open data and link to it?
Not for these size datasets. Torrents are just too small and unreliable, except for the very most popular items. Slightly out of mainstream movies, which are small and likely more popular than data for an obscure science experiment, are nearly impossible to find seeders for.
So I'd guess no seeders want to host tens to thousands of terabyte torrents, and then thousands to millions of those for all the different datasets grabbed by all the science projects all over the world. The odds of being able to download one of these on demand is just about zero.