For some reason I am not really moved by a lot of the hand wringing I am seeing lately.
It's a not a binary thing to me: LLMs are not god, but even without AGI, they have proven wildly useful to me. Calling them "shitty chat bots" doesn't sway me.
Further I have always assumed that everything that I post to the web is publicly accessible to everyone/everything. We lost any battle we thought we could wage some 2+ decades ago when web crawlers started hoovering up data from our sites.
This article isn’t about that. It’s about the externalized costs that LLM companies are pushing onto webmasters because of their aggressive scraping. It’s one thing to believe that LLMs are a good thing, it’s another thing to believe that individuals and cooperative groups that run small internet services ought to be the ones to pay for that good.
It's not about secret vs public, it's about resource overload on the websites. Existing crawlers so far mostly respected robots.txt, the LLM crawlers don't.
You, as a user, might not care, but as servers keep going down, more and more website owners start blocking LLMs. Good riddance, hopefully all good stuff gets locked down.
Or to use an analogy, your comment is similar to: "sure those delivery vans violate speed limits and occasionally hit the pedestrians. I don't care, those fast deliveries have been proven wildly useful to me"
I have published a lot of content about a particular topic online and I want it to be publicly accesible. A lot of people use it to create YouTube videos, that's fine (and a lot of them even cite me). I have a problem with LLM profiting from them.
Which now I realize is not different from people amaking YouTube videos. I feel there is a difference but I don't know how to explain it. Maybe there isn't. Ouch, writing this comment was not a good idea...
This difference in emotional reaction is because of the effort involved in the process. Functionally, we see YouTube video creation as a fundamentally difficult exercise (to do well) and results in a singular product (one video). Any additional content would need an ongoing investment of time and money from the creator. The LLMs though would not require an ongoing investment beyond the first training run, that is probably why you have a problem with it, they're an extremely high leverage way of taking advantage of content.
Individuals who have to do work in order to use your content to do work to create their own content is qualitatively different than automation trivially doing whatever.
To me this feels almost like the news complaining that they want a "link tax." Weren't their headlines and summaries used? It seems inconsistent to somehow say that AI and scraping is not okay; but that news companies should also not be entitled to their link tax. It's okay to index, but not that kind of index.
It seems pretty cut-and-dry to me: the website owner should opt-out if they don't like the deal (being indexed in this case).
In the "link tax" case, there were plenty of trivial ways to opt out of headline usage - robots.txt, http headers, http tags. The problem was newspapers did not want to opt out (as they were benefiting from Google themselves), so they wanted a 3rd option. Which was pretty stupid of course - if you don't like the deal, don't take it; suing the offering party for a better deal is not a good long-term strategy.
In the AI case, there is no opt-out. All those websites already indicated they want to opt-out via robots.txt, but the AI companies ignore robots.txt, change user-agent, fake IPs, and so on - do the things that are normally done by shady malwar-ish services rather than multi-billion-dollar companies.
It really bothers me when people don't see the difference between those two cases.
Comments
For some reason I am not really moved by a lot of the hand wringing I am seeing lately.
It's a not a binary thing to me: LLMs are not god, but even without AGI, they have proven wildly useful to me. Calling them "shitty chat bots" doesn't sway me.
Further I have always assumed that everything that I post to the web is publicly accessible to everyone/everything. We lost any battle we thought we could wage some 2+ decades ago when web crawlers started hoovering up data from our sites.
This article isn’t about that. It’s about the externalized costs that LLM companies are pushing onto webmasters because of their aggressive scraping. It’s one thing to believe that LLMs are a good thing, it’s another thing to believe that individuals and cooperative groups that run small internet services ought to be the ones to pay for that good.
It's not about secret vs public, it's about resource overload on the websites. Existing crawlers so far mostly respected robots.txt, the LLM crawlers don't.
You, as a user, might not care, but as servers keep going down, more and more website owners start blocking LLMs. Good riddance, hopefully all good stuff gets locked down.
Or to use an analogy, your comment is similar to: "sure those delivery vans violate speed limits and occasionally hit the pedestrians. I don't care, those fast deliveries have been proven wildly useful to me"
I have published a lot of content about a particular topic online and I want it to be publicly accesible. A lot of people use it to create YouTube videos, that's fine (and a lot of them even cite me). I have a problem with LLM profiting from them.
Which now I realize is not different from people amaking YouTube videos. I feel there is a difference but I don't know how to explain it. Maybe there isn't. Ouch, writing this comment was not a good idea...
This difference in emotional reaction is because of the effort involved in the process. Functionally, we see YouTube video creation as a fundamentally difficult exercise (to do well) and results in a singular product (one video). Any additional content would need an ongoing investment of time and money from the creator. The LLMs though would not require an ongoing investment beyond the first training run, that is probably why you have a problem with it, they're an extremely high leverage way of taking advantage of content.
You are confronted with automation.
Individuals who have to do work in order to use your content to do work to create their own content is qualitatively different than automation trivially doing whatever.
What you are feeling is described in The Work of Art in the Age of Mechanical Reproduction [1]
1. https://en.wikipedia.org/wiki/The_Work_of_Art_in_the_Age_of_...
I agree even if I have mixed feelings on it.
To me this feels almost like the news complaining that they want a "link tax." Weren't their headlines and summaries used? It seems inconsistent to somehow say that AI and scraping is not okay; but that news companies should also not be entitled to their link tax. It's okay to index, but not that kind of index.
It seems pretty cut-and-dry to me: the website owner should opt-out if they don't like the deal (being indexed in this case).
In the "link tax" case, there were plenty of trivial ways to opt out of headline usage - robots.txt, http headers, http tags. The problem was newspapers did not want to opt out (as they were benefiting from Google themselves), so they wanted a 3rd option. Which was pretty stupid of course - if you don't like the deal, don't take it; suing the offering party for a better deal is not a good long-term strategy.
In the AI case, there is no opt-out. All those websites already indicated they want to opt-out via robots.txt, but the AI companies ignore robots.txt, change user-agent, fake IPs, and so on - do the things that are normally done by shady malwar-ish services rather than multi-billion-dollar companies.
It really bothers me when people don't see the difference between those two cases.