Skip to content

Comment on Statistics as algorithmic summarizationparent

Comments

Why would we select a subset at random? Why shouldn't we try to select a representative subset?

Isn't this similar to the point this article is trying to make? Looking at statistics as a collection of algorithmic recipes shifts attention to how those procedures are designed, and when do they yield us useful summaries. When you want a "representative mean" rather than the population average, you'd just tweak your algorithm.

The mean from a representative sample is, if you got the representation right, "equal" to the mean from a random sample, and both are "equal" to the population average.

The problem with representative samples is that you can't know whether you've actually constructed one, and even if you have, you'll find it hard to compute the errors of your estimations.

My beef is not with thinking of traditional frequentist-objectivist statistics as a set of algorithms solving specific problems (because that's what it is), my beef is with not starting out explaining exactly when the algorithms apply and why.

When you take a black boxy algorithmic approach to something, you have to be particularly clear about things like that.

AboutSource Built by g1lg1l

Hackerly is an independent reader for Hacker News, built on the public HN API. Not affiliated with Y Combinator.