Skip to content

Comment on Musk’s X Corp sues Israel’s Bright Data for scraping dataparent

Comments

It seems to me that Twitter's robots.txt only allows Googlebot:

https://twitter.com/robots.txt

Therefore, this disallows other bots in "maschinenlesbarer Form" and the scraping is illegal.

Google is, to a first approximation, the web's only search engine. As such I think there's an argument to be made that declaring your bot's User-Agent as Googlebot is a technical necessity, similar to how every browser declares itself as Mozilla.

This is fascinating. Seems like at least Bing openly ignores robots.txt?

AboutSource Built by g1lg1l

Hackerly is an independent reader for Hacker News, built on the public HN API. Not affiliated with Y Combinator.