The sample is also much smaller than the article suggests. The paper itself says they found about 1800 'linked pairs' of users with a stated ZIP code in that sample of 100K. And in the graphs for the larger (4500) sample of email data, the fit is quite bad, with a pronounced 'hump' in the tail. There are just far too few sample points here to make these kinds of sweeping conclusions.
Comments
The sample is also much smaller than the article suggests. The paper itself says they found about 1800 'linked pairs' of users with a stated ZIP code in that sample of 100K. And in the graphs for the larger (4500) sample of email data, the fit is quite bad, with a pronounced 'hump' in the tail. There are just far too few sample points here to make these kinds of sweeping conclusions.