The author has an earlier post where he cites a particular study as an example of the overreliance on probalistic models: http://www.argmin.net/2021/09/13/effect-size/ . The author seems to think this kind of inappropriate use of models is fairly prevalent.
What are your thoughts on that mask example? To me, this seems like a reasonable critique, but (as I just posted in this thread) I don't have a deep understanding of statistics, so I am a little uncertain if my interpretation of the blog post is correct.
I think that blogpost is very reasonable, and actually when I first heard about the mask study my thinking was very similar, "whoah, how do you even start to control for all of the kinds of interference that'll play havoc with the randomization?". But I don't agree with the claim that "P-values and confidence intervals associated with a regression are valid only if the model is true. What if the model is not true?" The key to statistical thinking is to think in terms of quantities rather than dichotomies, so a Poisson distribution may not be ideal, okay, but how far away from ideal is it; the randomization is not perfect, okay, but how much bias will that introduce? At the end of the day, you're still going to need a model to answer those questions.
That makes sense. I think the author would argue that the model is never needed, just go back to the actual population, and make inferences based on random sampling. From another prior blog post[1], he states:
"So what is the remedy here? The thing is, we already know the answer: if we randomized the assignment ... it is critical to decouple the randomness used to probe a system from the randomness inherent in its system itself. Statistical models are not necessary for statistical inference, but randomness itself is amazingly… let’s say… useful for understanding natural phenomena."
This makes sense to me if there's always a population that you can random sample. And yes, you'd have to sample a lot for the sample and population summaries to converge, but this seems fine in certain contexts (i.e. when you have access to computer simulations). Your counterpoint seems correct when you aren't able to randomly sample the population in this manner.
Would you agree with the author based on the specific framing I'm making (randomly sampling population beats building a probabalistic model if you are able to do the former). Again, I'm a stats novice so apologies if I'm making an obvious point.
Comments
The author has an earlier post where he cites a particular study as an example of the overreliance on probalistic models: http://www.argmin.net/2021/09/13/effect-size/ . The author seems to think this kind of inappropriate use of models is fairly prevalent.
What are your thoughts on that mask example? To me, this seems like a reasonable critique, but (as I just posted in this thread) I don't have a deep understanding of statistics, so I am a little uncertain if my interpretation of the blog post is correct.
I think that blogpost is very reasonable, and actually when I first heard about the mask study my thinking was very similar, "whoah, how do you even start to control for all of the kinds of interference that'll play havoc with the randomization?". But I don't agree with the claim that "P-values and confidence intervals associated with a regression are valid only if the model is true. What if the model is not true?" The key to statistical thinking is to think in terms of quantities rather than dichotomies, so a Poisson distribution may not be ideal, okay, but how far away from ideal is it; the randomization is not perfect, okay, but how much bias will that introduce? At the end of the day, you're still going to need a model to answer those questions.
That makes sense. I think the author would argue that the model is never needed, just go back to the actual population, and make inferences based on random sampling. From another prior blog post[1], he states:
"So what is the remedy here? The thing is, we already know the answer: if we randomized the assignment ... it is critical to decouple the randomness used to probe a system from the randomness inherent in its system itself. Statistical models are not necessary for statistical inference, but randomness itself is amazingly… let’s say… useful for understanding natural phenomena."
This makes sense to me if there's always a population that you can random sample. And yes, you'd have to sample a lot for the sample and population summaries to converge, but this seems fine in certain contexts (i.e. when you have access to computer simulations). Your counterpoint seems correct when you aren't able to randomly sample the population in this manner.
Would you agree with the author based on the specific framing I'm making (randomly sampling population beats building a probabalistic model if you are able to do the former). Again, I'm a stats novice so apologies if I'm making an obvious point.
1. http://www.argmin.net/2021/09/21/models-are-wrong/