To corroborate this, it has not been our experience that Splunk doesn't scale. We have a ~40GB/day environment and according to support can scale that to ~100GB/day on a single quad-core server with SAN back-end, license permitting. The lag between reception and indexing is negligible; we actually have some developers looking at their own debug logs in it.
Unfortunately, I don't work in a TB/day environment (though that would be fun), so I can't comment on that directly, but our experience with Splunk has been positive at this scale and I see no reason so far why a distributed installation with multiple 100GB/day nodes would change that. The only customer complaints we have are basically that long-term searches don't return instantly, and that's primarily a factor of I/O speed and can be resolved by setting up summary indexing (which requires forethought about what data you're interested in... therefore, it very rarely happens in my organization).
Comments
To corroborate this, it has not been our experience that Splunk doesn't scale. We have a ~40GB/day environment and according to support can scale that to ~100GB/day on a single quad-core server with SAN back-end, license permitting. The lag between reception and indexing is negligible; we actually have some developers looking at their own debug logs in it.
Unfortunately, I don't work in a TB/day environment (though that would be fun), so I can't comment on that directly, but our experience with Splunk has been positive at this scale and I see no reason so far why a distributed installation with multiple 100GB/day nodes would change that. The only customer complaints we have are basically that long-term searches don't return instantly, and that's primarily a factor of I/O speed and can be resolved by setting up summary indexing (which requires forethought about what data you're interested in... therefore, it very rarely happens in my organization).