Comment on Musk’s X Corp sues Israel’s Bright Data for scraping dataparentComments−mschuster913yOr in Germany, see §44b UrhG.[1] https://www.gesetze-im-internet.de/urhg/__44b.html−nmlracx3yIt seems to me that Twitter's robots.txt only allows Googlebot:https://twitter.com/robots.txtTherefore, this disallows other bots in "maschinenlesbarer Form" and the scraping is illegal.−gnfargbl3yGoogle is, to a first approximation, the web's only search engine. As such I think there's an argument to be made that declaring your bot's User-Agent as Googlebot is a technical necessity, similar to how every browser declares itself as Mozilla.−mschuster913yThis is fascinating. Seems like at least Bing openly ignores robots.txt?
Comments
Or in Germany, see §44b UrhG.
[1] https://www.gesetze-im-internet.de/urhg/__44b.html
It seems to me that Twitter's robots.txt only allows Googlebot:
https://twitter.com/robots.txt
Therefore, this disallows other bots in "maschinenlesbarer Form" and the scraping is illegal.
Google is, to a first approximation, the web's only search engine. As such I think there's an argument to be made that declaring your bot's User-Agent as Googlebot is a technical necessity, similar to how every browser declares itself as Mozilla.
This is fascinating. Seems like at least Bing openly ignores robots.txt?