The methodology here is quite simplistic, and suspect.
The way they “objectively” decided on “quality” of posts, by which to judge the ‘before’ picture of a poster was via some machine learning analysis of the keywords in the post in question, implicitly assuming that their keyword analysis was a better judgment of quality than the explicit voting on the site. This seems like a very poor assumption to me.
It’s quite plausible (to me) that the posts’ keywords in the negatively-voted posts looked okay to their algorithm while the content still turned out to be trolling or bullshit (hence the downvotes). It would be quite reasonable to assume that posts full of trolling or bullshit were by posters who were inclined to future trolling/bullshit posts.
I’d be very interested to see what results they’d get if they ran the clock in reverse. I.e. pulled some posts of similar “quality” based on their metric but different voted scores, and then looked at the several posts before that, from those posters.
Comments
The methodology here is quite simplistic, and suspect.
The way they “objectively” decided on “quality” of posts, by which to judge the ‘before’ picture of a poster was via some machine learning analysis of the keywords in the post in question, implicitly assuming that their keyword analysis was a better judgment of quality than the explicit voting on the site. This seems like a very poor assumption to me.
It’s quite plausible (to me) that the posts’ keywords in the negatively-voted posts looked okay to their algorithm while the content still turned out to be trolling or bullshit (hence the downvotes). It would be quite reasonable to assume that posts full of trolling or bullshit were by posters who were inclined to future trolling/bullshit posts.
I’d be very interested to see what results they’d get if they ran the clock in reverse. I.e. pulled some posts of similar “quality” based on their metric but different voted scores, and then looked at the several posts before that, from those posters.