This is a pretty thin article - as far as I can tell entire article boils down to "apply a statistical test" to ab test results?? The four step process laid out doesn't magically protect you from confirmation bias, for example there are ample opportunities for it to manifest while "identifying the ideal data".
The toy example provided is actually a perfect example of why tools (like chat gpt) can't protect you from fucking this up. Why is the view count so different between versions a and b? Was the test set up as a 60-40 split, or had you intended it to be closer to 50-50? Does the massive increase in conversion rate pass the smell test? If anything seems suspicious, you probably should do some further investigations...
Comments
This is a pretty thin article - as far as I can tell entire article boils down to "apply a statistical test" to ab test results?? The four step process laid out doesn't magically protect you from confirmation bias, for example there are ample opportunities for it to manifest while "identifying the ideal data".
The toy example provided is actually a perfect example of why tools (like chat gpt) can't protect you from fucking this up. Why is the view count so different between versions a and b? Was the test set up as a 60-40 split, or had you intended it to be closer to 50-50? Does the massive increase in conversion rate pass the smell test? If anything seems suspicious, you probably should do some further investigations...