Skip to content

Comment on Statistics as algorithmic summarization

Comments

So I've read this article a couple of times and I still don't quite understand it. Can anyone summarize the key argument for a layperson?

Here's my best interpretation thus far:

This particular post is a continuation of his two earlier posts[1], [2] which seems to critique the way statistical significance is simplistically calculated from a gaussian distribution when effect size is small. This leads to an argument that the statistics community is too dependant on generative models for their interpretation, given that these models are typically to simplistic outside of hard sciences:

"But in biology, medicine, social science, and economics, our models are much less accurate and less grounded in natural laws. Most of the time, models are selected because they are convenient, not because they are plausible, well motivated from phenomenological principles, or even empirically validated. Freedman built a cottage industry around pointing out how poorly motivated many of the common statistical models are."

This part I am a little unclear on, but it seems like this leads him to suggest focusing on the random sampling as a way to get counts that you can then plug into various statistical formulas, without the need to assume a probabalistic model:

"So what is the remedy here? The thing is, we already know the answer: if we randomized the assignment, we can estimate log odds by counting the number of positive outcomes under treatment and control, and then just plugging these values into the odds ratio. If you do this, you find an estimate whose median is precisely equal to the true log odds. No covariate adjustment is required."

So I think, what the author means by "algorithmic summation" is that we focus on random experiment design, and discard model assumptions. Is that right?

If so I think this makes sense. I believe this is something Allen Downey has talked about before, specifically saying statistical experiments can now take advantage of cheap computational simulation to hit the large numbers needed for the sample to approximate the population, without a need for the typical model approximations developed in a pre-computational era. Downey's post here: http://allendowney.blogspot.com/2016/06/there-is-still-only-...

1. http://www.argmin.net/2021/09/13/effect-size/

2. http://www.argmin.net/2021/09/21/models-are-wrong/

AboutSource Built by g1lg1l

Hackerly is an independent reader for Hacker News, built on the public HN API. Not affiliated with Y Combinator.