From a data analyst's perspective, let's go through what he says.
First he states something along the lines of "More data does not always help." This is right from a theoretical perspective. But: it never hurts. This is also right from a theoretical perspective, it's a result from probability theory: additional observations will always lead to less or equal variance in your estimations. There is no data like more data. There is no down side with more data.
I am not sure in what way (2) and (3) relate to big data. I'd even say that (3) is pro big data.
Then there is this term "intelligent data". Actually, I can't emphasize how badly chosen this term is. Intelligence is related to the quality of actions someone takes. Data does not take actions, It just "is". Data cannot be intelligent, just as a stone cannot be intelligent.
He also thinks that data measurements should be repeatable. Guess what, in all interesting cases data measurements are not repeatable due to randomness in the source itself. One of the main challenges of data analysis is to still get robust results.
He also thinks that data should be concise, e.g. that the data set at hand should be as minimal as possible to lead to the same actions. This sounds like a chicken and egg problem. How would you be able to even assess this without trying it out?
You're neglecting that data costs money and time to collect and process. More data is more cost.
I am in agreement with the spirit of this post (I'm not interested in arguing whether intelligent data is a good term or not.) Heck I even blogged along similar lines just a few days ago: http://noelwelsh.com/streaming-algorithms/2012/08/29/lean-da... Here are a few problems with collecting everything:
- Big Data infrastructure like Hadoop is expensive and slow. It's very much not in the turn-on-a-dime spirit of startups.
- If you collect everything, the value per data item is low. This impacts the analyses you profitably do. Compare the value that Klaviyo can deliver per data point vs, say, Mixpanel. (And then ask yourself why Mixpanel is moving into "People" analytics. My suggestion: because it's much more valuable.)
Disclaimer 1: My startup, Myna [http://mynaweb.com/], had a shout-out in the blog post.
Disclaimer 2: I'm a Klaviyo user, as of a few days ago.
Today the storage of more data is negligible. There are alternatives to Hadoop. However, to capture new variables that exist in your market place definitely takes time and money.
I think you make a great point here about the confusion that often gets put out there between data and analysis - a confusion which I'd say is implicit in the term big data as well (and hence I ran with - caveat, I'm the author).
As far as the problem with the term "intelligent data" - I think what you say is exactly true if you do data analysis one time; however, the issue is that for those of us running startups, we find ourselves doing analysis over and over - so intelligently selecting data (in a way that takes us less time together and leads to the same decisions) is a huge win. Read intelligent data as being data + intelligence - not a new type of data.
Likewise, the problem with asserting that more data is always better ignores how most companies are making decisions. At the end of the day, our analysis is completely meaningless without a new action. So a better analysis that doesn't get implemented is worth far less (nothing) than an analysis that gets implemented successfully and drives results.
Good point. There is a cost/benefit balancing for attaining more data which may turn out to be useless. However, technology is increasingly pushing us in a direction that allows us to capture more data at a cheaper cost.
Agreed. Intelligent Data seems like a very messy term. In Big Data we use many different processes to figure out what the "intelligent" relationships are, or what variables really express strong relationships. This is one of the biggest challenges in Big Data. A quote from Alex Pentland:
"the data scientists themselves don't have much of intuition either…and that is a problem. I saw an estimate recently that said 70 to 80 percent of the results that are found in the machine learning literature, which is a key Big Data scientific field, are probably wrong because the researchers didn't understand that they were overfitting the data."
Comments
From a data analyst's perspective, let's go through what he says.
First he states something along the lines of "More data does not always help." This is right from a theoretical perspective. But: it never hurts. This is also right from a theoretical perspective, it's a result from probability theory: additional observations will always lead to less or equal variance in your estimations. There is no data like more data. There is no down side with more data.
I am not sure in what way (2) and (3) relate to big data. I'd even say that (3) is pro big data.
Then there is this term "intelligent data". Actually, I can't emphasize how badly chosen this term is. Intelligence is related to the quality of actions someone takes. Data does not take actions, It just "is". Data cannot be intelligent, just as a stone cannot be intelligent. He also thinks that data measurements should be repeatable. Guess what, in all interesting cases data measurements are not repeatable due to randomness in the source itself. One of the main challenges of data analysis is to still get robust results. He also thinks that data should be concise, e.g. that the data set at hand should be as minimal as possible to lead to the same actions. This sounds like a chicken and egg problem. How would you be able to even assess this without trying it out?
You're neglecting that data costs money and time to collect and process. More data is more cost.
I am in agreement with the spirit of this post (I'm not interested in arguing whether intelligent data is a good term or not.) Heck I even blogged along similar lines just a few days ago: http://noelwelsh.com/streaming-algorithms/2012/08/29/lean-da... Here are a few problems with collecting everything:
- Big Data infrastructure like Hadoop is expensive and slow. It's very much not in the turn-on-a-dime spirit of startups.
- If you collect everything, the value per data item is low. This impacts the analyses you profitably do. Compare the value that Klaviyo can deliver per data point vs, say, Mixpanel. (And then ask yourself why Mixpanel is moving into "People" analytics. My suggestion: because it's much more valuable.)
Disclaimer 1: My startup, Myna [http://mynaweb.com/], had a shout-out in the blog post.
Disclaimer 2: I'm a Klaviyo user, as of a few days ago.
Today the storage of more data is negligible. There are alternatives to Hadoop. However, to capture new variables that exist in your market place definitely takes time and money.
I think you make a great point here about the confusion that often gets put out there between data and analysis - a confusion which I'd say is implicit in the term big data as well (and hence I ran with - caveat, I'm the author).
As far as the problem with the term "intelligent data" - I think what you say is exactly true if you do data analysis one time; however, the issue is that for those of us running startups, we find ourselves doing analysis over and over - so intelligently selecting data (in a way that takes us less time together and leads to the same decisions) is a huge win. Read intelligent data as being data + intelligence - not a new type of data.
Likewise, the problem with asserting that more data is always better ignores how most companies are making decisions. At the end of the day, our analysis is completely meaningless without a new action. So a better analysis that doesn't get implemented is worth far less (nothing) than an analysis that gets implemented successfully and drives results.
Good point. There is a cost/benefit balancing for attaining more data which may turn out to be useless. However, technology is increasingly pushing us in a direction that allows us to capture more data at a cheaper cost.
Agreed. Intelligent Data seems like a very messy term. In Big Data we use many different processes to figure out what the "intelligent" relationships are, or what variables really express strong relationships. This is one of the biggest challenges in Big Data. A quote from Alex Pentland:
"the data scientists themselves don't have much of intuition either…and that is a problem. I saw an estimate recently that said 70 to 80 percent of the results that are found in the machine learning literature, which is a key Big Data scientific field, are probably wrong because the researchers didn't understand that they were overfitting the data."
http://www.edge.org/conversation/reinventing-society-in-the-...