> they needed to download two thirds of a terabyte of compressed data. “We’ve got 30 to 40 servers pulling down data from NASA,”
40 servers to download 600 gigabytes of data? Something does not sound right here. If they wanted to avoid overloading the nasa pipes, they could have asked nasa to fedex that amount on couple of hard-drives or something. At this day and age bulk transferring a terabyte or two of data should not be a challenge.
Heh, yes, we seriously considered the Honda Civic full of external drives option.[0]
Some of our images were from NASA’s hot new GIBS[1] service, which was extremely fast. However, it’d only backfilled about 2/5 of the data we needed, so for the rest of it we were using a legacy endpoint that was launched around 2004ish and has some, let’s say, idiosyncratic caching and throttling systems.
Instead of setting up a special channel, we talked to them and figured out how to shotgun the downloads in a way that wouldn’t kill their cache or require them to give us special treatment. We thought of this on a Friday and wanted to have it ready when we came in on Monday, and it worked. Thus the somewhat blunt methods.
0. For any young persons in the audience, the reference is to “Never underestimate the bandwidth of a station wagon full of tapes hurtling down the highway.” from Tanenbaum’s “Computer Networks”.
My guess would be they wanted to uncompress and process it in parallel? I doubt it was a question of bandwidth—probably more a question of convenience.
I think it's fair to say that they thought it through. They also talked to NASA [1] so NASA would have told them if there was a better way (e.g. hard drive) than downloading over the net.
[1] “We called them up and said, ‘hey we’re going to hit you hard, what’s the best way we can do it for you?’”
Comments
> they needed to download two thirds of a terabyte of compressed data. “We’ve got 30 to 40 servers pulling down data from NASA,”
40 servers to download 600 gigabytes of data? Something does not sound right here. If they wanted to avoid overloading the nasa pipes, they could have asked nasa to fedex that amount on couple of hard-drives or something. At this day and age bulk transferring a terabyte or two of data should not be a challenge.
Heh, yes, we seriously considered the Honda Civic full of external drives option.[0]
Some of our images were from NASA’s hot new GIBS[1] service, which was extremely fast. However, it’d only backfilled about 2/5 of the data we needed, so for the rest of it we were using a legacy endpoint that was launched around 2004ish and has some, let’s say, idiosyncratic caching and throttling systems.
Instead of setting up a special channel, we talked to them and figured out how to shotgun the downloads in a way that wouldn’t kill their cache or require them to give us special treatment. We thought of this on a Friday and wanted to have it ready when we came in on Monday, and it worked. Thus the somewhat blunt methods.
0. For any young persons in the audience, the reference is to “Never underestimate the bandwidth of a station wagon full of tapes hurtling down the highway.” from Tanenbaum’s “Computer Networks”.
1. https://earthdata.nasa.gov/about-eosdis/system-description/g...
My guess would be they wanted to uncompress and process it in parallel? I doubt it was a question of bandwidth—probably more a question of convenience.
I think it's fair to say that they thought it through. They also talked to NASA [1] so NASA would have told them if there was a better way (e.g. hard drive) than downloading over the net.
[1] “We called them up and said, ‘hey we’re going to hit you hard, what’s the best way we can do it for you?’”