Skip to content

Comment on Sergey Brin’s Search for a Parkinson’s Cure parent

Comments

Yeah, we haven't even had to deal with data directly from the instrument yet. So far, we've been getting data from collaborators for analysis via terabyte usb drives (FedEx throughput can't be beat). For the actual analysis, we've found the same thing... disk IO is a limiting factor. Well, that and the 16GB human genome indexes in RAM. And we aren't on an Isilon system yet (probably won't be either).

However, we just got our own instrument, so this will definitely be an issue, but our University knows a thing or two about dealing with big data (http://kb.iu.edu/data/avvh.html).

I've only dealt with a few GWAS style datasets, and the next-gen stuff dwarfs the GWAS data in terms of size. But when looking for linkages between variations, we're still talking more time than the universe is old level of calculations for more than 3 combinations. Which is really scary, because like you said, all the genetics people are going to be using sequencing for most things from here on out, so its like you have complexity on top of complexity...

AboutSource Built by g1lg1l

Hackerly is an independent reader for Hacker News, built on the public HN API. Not affiliated with Y Combinator.