Ah, but it often belies and smooths over the important distinguishing factor of data: statistical significance.
Simply having "plural" is not enough. You need a well-designed experiment, with well-controlled variables, good methods and measurements, good consistent recording, and enough points of that nature to show significance. And all that without natural biases, such as the confirmation bias so common with anecdotes. You can have as many self-reported stories from a self-selecting population as you like and it still won't be good data, due to the inherent bias in the method.
It's difficult to recognize these biases, which is why we say "the plural of anecdote is not data." Not because some data isn't a series of anecdotal points, but because good data is often so much more, and it's important to respect that. Sure, use it to guide your instincts, but don't mistake it for rigorous science.
You're making a lot of assumptions, though. You're assuming self-selection, and you're assuming what kinds of situations the data is being used. The plural of anecdote isn't data, but a group of anecdotes isn't inherently not data.
Further, a collection of many anecdotes is not necessarily great data, but the quality of data can be taken into account when making assertions about it, and you can attach confidence ratings to assertions made based on data with known faults.
The problem here is that "the plural of anecdote isn't data" is just a short, snappy thing to say, that doesn't capture any of the nuances of what data is and how it's used. A lot of times it's used to express a legitimate point, but I think we could express that point more accurately by actually talking about what data is instead of oversimplifying the concern.
Statistical significance? At what p-value? Controlled studies are great, but claims of "statistical significance" are not as strong as we make them to be. I also don't like this idea of an arbitrary threshold.
For the nice case of comparing two hypotheses, I'd rather talk about decibels of evidence: which hypothesis does it support, and how strongly does it support it. That way you don't even have to chose a null hypothesis. (You still have to work with some "priors", but in practice you have to assume things anyway).
Sure, use it to guide your instincts, but don't mistake it for rigorous science.
"Rigorous science" is the easy part. It's the point where you already have some insight or theory in mind, and you just have to test it. But first, you need to get that insight, or formulate your theory. At that point, your instincts is pretty much all you have.
Comments
Ah, but it often belies and smooths over the important distinguishing factor of data: statistical significance.
Simply having "plural" is not enough. You need a well-designed experiment, with well-controlled variables, good methods and measurements, good consistent recording, and enough points of that nature to show significance. And all that without natural biases, such as the confirmation bias so common with anecdotes. You can have as many self-reported stories from a self-selecting population as you like and it still won't be good data, due to the inherent bias in the method.
It's difficult to recognize these biases, which is why we say "the plural of anecdote is not data." Not because some data isn't a series of anecdotal points, but because good data is often so much more, and it's important to respect that. Sure, use it to guide your instincts, but don't mistake it for rigorous science.
You're making a lot of assumptions, though. You're assuming self-selection, and you're assuming what kinds of situations the data is being used. The plural of anecdote isn't data, but a group of anecdotes isn't inherently not data.
Further, a collection of many anecdotes is not necessarily great data, but the quality of data can be taken into account when making assertions about it, and you can attach confidence ratings to assertions made based on data with known faults.
The problem here is that "the plural of anecdote isn't data" is just a short, snappy thing to say, that doesn't capture any of the nuances of what data is and how it's used. A lot of times it's used to express a legitimate point, but I think we could express that point more accurately by actually talking about what data is instead of oversimplifying the concern.
Statistical significance? At what p-value? Controlled studies are great, but claims of "statistical significance" are not as strong as we make them to be. I also don't like this idea of an arbitrary threshold.
For the nice case of comparing two hypotheses, I'd rather talk about decibels of evidence: which hypothesis does it support, and how strongly does it support it. That way you don't even have to chose a null hypothesis. (You still have to work with some "priors", but in practice you have to assume things anyway).
"Rigorous science" is the easy part. It's the point where you already have some insight or theory in mind, and you just have to test it. But first, you need to get that insight, or formulate your theory. At that point, your instincts is pretty much all you have.