What would happen if you actually did try to textmine it.
surely this is a game of chicken because no company wants to lose tens of thousands of dollars a year for no reason.
Elsevier is willing to play this game of chicken; they are convinced that their customers (i.e., research universities) cannot do without their product. Hence the OP reporting that twice, they instantly and without warning cut off all access to their product to the entire university, because of detected scraping.
Telling people what software they can use to read text doesn't scale.
Unfortunately, I think it does. If individual users, even a large number of them, use a plug-in, they might get away with it-- but this would result in an incomplete data set. To systematically textmine the corpus (which is the task at hand) requires some kind of systematic access to the data, and this is where Elsevier steps in and shuts it down.
What actually happen with that guy? I remember all the mainstream media stories what kind of a bad ass hacker he was, but somewhat never caught up with the "happy end".
Not really a bad ending. The feds have no case against him, MIT and JSTOR aren't going to sue him, and JSTOR decided that releasing their archive was a good idea. Hopefully the FBI will drop their case.
35 years in prison is what you would get in federal court for walking across the street as the light was changing from yellow to red. It is there to give the government a better position to offer a plea agreement. It is highly unlikely that he would be convicted of every count, and even if he was, it is highly unlikely that he would receive the maximum sentence for each count.
But, 35 years is scary, and if I were in his shoes and the government said "pay a $100,000 fine", I'd probably agree without much argument. And that's exactly the point of saying he faces "up to" 35 years.
Comments
What would happen if you actually did try to textmine it. surely this is a game of chicken because no company wants to lose tens of thousands of dollars a year for no reason.
Elsevier is willing to play this game of chicken; they are convinced that their customers (i.e., research universities) cannot do without their product. Hence the OP reporting that twice, they instantly and without warning cut off all access to their product to the entire university, because of detected scraping.
Telling people what software they can use to read text doesn't scale.
Unfortunately, I think it does. If individual users, even a large number of them, use a plug-in, they might get away with it-- but this would result in an incomplete data set. To systematically textmine the corpus (which is the task at hand) requires some kind of systematic access to the data, and this is where Elsevier steps in and shuts it down.
Possibility number two is to break in to a router closet at MIT to do the scraping :)
What actually happen with that guy? I remember all the mainstream media stories what kind of a bad ass hacker he was, but somewhat never caught up with the "happy end".
http://en.wikipedia.org/wiki/Aaron_Swartz#JSTOR
Not really a bad ending. The feds have no case against him, MIT and JSTOR aren't going to sue him, and JSTOR decided that releasing their archive was a good idea. Hopefully the FBI will drop their case.
35 years in prison would be a pretty bad ending, and he's still being prosecuted for that.
35 years in prison is what you would get in federal court for walking across the street as the light was changing from yellow to red. It is there to give the government a better position to offer a plea agreement. It is highly unlikely that he would be convicted of every count, and even if he was, it is highly unlikely that he would receive the maximum sentence for each count.
But, 35 years is scary, and if I were in his shoes and the government said "pay a $100,000 fine", I'd probably agree without much argument. And that's exactly the point of saying he faces "up to" 35 years.