I'm not qualified to comment on Cam's math, but Evan Miller's tools [1] are my go-to source for calculating sample sizes and significance. How does this differ? (If it does.)
The real (and all to common to run into ) problem isn't having 100x the samples you really need, it is having 1/10 the samples you really need... And not knowing that.
Comments
I'm not qualified to comment on Cam's math, but Evan Miller's tools [1] are my go-to source for calculating sample sizes and significance. How does this differ? (If it does.)
http://www.evanmiller.org/ab-testing/sample-size.html
That looks right to me. This is a very important calculation to do in order to avoid getting 100x the number of samples that you really need.
Essentially the same calculation applies to classification error.
The real (and all to common to run into ) problem isn't having 100x the samples you really need, it is having 1/10 the samples you really need... And not knowing that.