Skip to content

Meta has tried to scrape this site 1 million times in 2 weeks

robertmay.photography
14 pointsrobotmay2 comments
On HN

Comments

I still don’t get why AI scrapers are so much worse than search engine scrapers. Why are they executing this so poorly?

Meta presumably has some at least marginally competent people on hand that don’t need random bloggers to tell them to sort out their scraping tech…

Yeah I've been really surprised by how inept they are. Lots of others have fallen for the trap but extract themselves fairly quickly and probably blacklist my site. Archive.org got stuck for a while, which was unfortunate, so I had to exclude them manually. There's something on an Oracle network that keeps coming back but it's very slow compared to the ridiculous rate of requests from Meta's scraper.

It's up to 1.5 million requests now, sigh. My guess is that Meta has too much money and is in the "move fast and break things" stage where they're just throwing money at a problem. Not setting a max crawl depth is hilarious though.

AboutSource Built by g1lg1l

Hackerly is an independent reader for Hacker News, built on the public HN API. Not affiliated with Y Combinator.