Skip to content

Comment on Amazon has a way to scrape GitHub and feed its AI modelparent

Comments

I think that is very different from cloning repos from Github in accordance with the licences.

I agree crawling sites to the extent it causes problems is a problem.

Googlebot and Bingbot do follow robots.txt and respect HTTP 429 responses and usually have reasonable default crawl rates.

Is it possible that these are scrapers using fake UA strings?

Is it possible that these are scrapers using fake UA strings?

I now believe this is the case

AboutSource Built by g1lg1l

Hackerly is an independent reader for Hacker News, built on the public HN API. Not affiliated with Y Combinator.