Skip to content

Comment on How a PhD astrophysicist thinks about dataparent

Comments

Sequence read archive has been such a boon for developing reproducible biological pipelines without having to worry about data. A paper references a dataset by ID and I can use it as input for my pipelines and keep the raw data locally only for as long as its needed to generate analysis within the running pipeline. I can even set threshold levels of how much local or cloud compute resources should be used at a time if I didn't want to exhaust my systems with one job.

AboutSource Built by g1lg1l

Hackerly is an independent reader for Hacker News, built on the public HN API. Not affiliated with Y Combinator.