For somebody who knows a bit how things are set up, or is willing to spend 10 minutes researching, it's a no-brainer that you can just "git clone" entire linux kernel development history, or download entire wikipedia [0].
Alas, large number of scrapers are not willing to spend those 10 minutes, it would appear. So, here we are.
Comments
This is exactly the problem, unfortunately.
For somebody who knows a bit how things are set up, or is willing to spend 10 minutes researching, it's a no-brainer that you can just "git clone" entire linux kernel development history, or download entire wikipedia [0].
Alas, large number of scrapers are not willing to spend those 10 minutes, it would appear. So, here we are.
[0] https://dumps.wikimedia.org/
You can tell Claude to clone from github for Linux stuff all you want... it's still going to try web, and fail, before doing what you asked it to do.