EDIT: Just confirm, I think I was incorrect in the message below.
I'm not completely sure I'm correct (please correct me if I'm not!) here but as I understand it the article does not support the claim in the headline. The headline claims that signups increases by 28% in the changed version and that this was all attributable to the change.
It's the second bit that isn't supported, they say that the result was statistical significant but what I understand them as saying was that it was statistically significant that the new variation was better than the control. But it could be better by an amount more or less than 28%, all we know from that is it's almost certainly (95%) at least a little better. We would need to know the number of trails to be able to get a certainty for the amount of improvement.
Could someone with a slightly better understanding of statistics chip in maybe? I could use some more information in my own A/B tests, sometimes I know a change is going to be a pain to maintain so I want to know not just if it is better, but by how much.
Since I am the one who wrote the case study, perhaps I can explain. Loosely saying, what the article actually means is that the observed difference of 28% or more may actually be true in 95% of cases. However, there is a 5% chance that that difference is a lesser than 28%.
If you have any specific question, please feel free to ask.
How are you using to arrive at that conclusion though? When I do a A/B testing I (or rather, the software I use) use the chi-squared test to give a confidence value expressed as a percentage. When I get over 95% I know I have statistical significance and I end the test. At that point I also have an "improved by" number but I would have to collect much more data for that to be statistically significant as far as I understand.
I'm not an expert in statistics, but I'm careful about what I do, and I know the limitations of my knowledge.
However ...
I believe it is flawed methodolgy to run a test until you get a significant result. I'm pretty sure I read something lunk to from HN that discussed this at some length. It's possible - if you simply run your test until you get significance and then stop early - that you will get significance because of random fluctuations in the middle of your trial and stop early when you shouldn't.
As I recall, you should decide on the length of your trial at the beginning, then run your stats at the end.
I'll try to find the article in question, but my Google-Fu is pretty poor today for some reason.
HN won't let me reply to your answer below (presumably to stop me from flaming you :p). But thanks, I'll read that blog post later but it looks like it answers my questions plus a bit more.
Seems my understanding was even less than I thought :)
you are right - without the sample sizes (for before and after) there's no way to know whether this measured increase of 28% is statistically significant or not.
Comments
EDIT: Just confirm, I think I was incorrect in the message below.
I'm not completely sure I'm correct (please correct me if I'm not!) here but as I understand it the article does not support the claim in the headline. The headline claims that signups increases by 28% in the changed version and that this was all attributable to the change.
It's the second bit that isn't supported, they say that the result was statistical significant but what I understand them as saying was that it was statistically significant that the new variation was better than the control. But it could be better by an amount more or less than 28%, all we know from that is it's almost certainly (95%) at least a little better. We would need to know the number of trails to be able to get a certainty for the amount of improvement.
Could someone with a slightly better understanding of statistics chip in maybe? I could use some more information in my own A/B tests, sometimes I know a change is going to be a pain to maintain so I want to know not just if it is better, but by how much.
Since I am the one who wrote the case study, perhaps I can explain. Loosely saying, what the article actually means is that the observed difference of 28% or more may actually be true in 95% of cases. However, there is a 5% chance that that difference is a lesser than 28%.
If you have any specific question, please feel free to ask.
I must be have been wrong then, my apologies.
How are you using to arrive at that conclusion though? When I do a A/B testing I (or rather, the software I use) use the chi-squared test to give a confidence value expressed as a percentage. When I get over 95% I know I have statistical significance and I end the test. At that point I also have an "improved by" number but I would have to collect much more data for that to be statistically significant as far as I understand.
I'm not an expert in statistics, but I'm careful about what I do, and I know the limitations of my knowledge.
However ...
I believe it is flawed methodolgy to run a test until you get a significant result. I'm pretty sure I read something lunk to from HN that discussed this at some length. It's possible - if you simply run your test until you get significance and then stop early - that you will get significance because of random fluctuations in the middle of your trial and stop early when you shouldn't.
As I recall, you should decide on the length of your trial at the beginning, then run your stats at the end.
I'll try to find the article in question, but my Google-Fu is pretty poor today for some reason.
EDIT: It's here:
http://news.ycombinator.com/item?id=1277004
See this http://visualwebsiteoptimizer.com/split-testing-blog/tag/mat...
HN won't let me reply to your answer below (presumably to stop me from flaming you :p). But thanks, I'll read that blog post later but it looks like it answers my questions plus a bit more.
Seems my understanding was even less than I thought :)
you are right - without the sample sizes (for before and after) there's no way to know whether this measured increase of 28% is statistically significant or not.