Meta has tried to scrape this site 1 million times in 2 weeks

(robertmay.photography)

13 points | by robotmay a day ago ago

2 comments

  • Havoc 14 hours ago ago

    I still don’t get why AI scrapers are so much worse than search engine scrapers. Why are they executing this so poorly?

    Meta presumably has some at least marginally competent people on hand that don’t need random bloggers to tell them to sort out their scraping tech…

    • robotmay 14 hours ago ago

      Yeah I've been really surprised by how inept they are. Lots of others have fallen for the trap but extract themselves fairly quickly and probably blacklist my site. Archive.org got stuck for a while, which was unfortunate, so I had to exclude them manually. There's something on an Oracle network that keeps coming back but it's very slow compared to the ridiculous rate of requests from Meta's scraper.

      It's up to 1.5 million requests now, sigh. My guess is that Meta has too much money and is in the "move fast and break things" stage where they're just throwing money at a problem. Not setting a max crawl depth is hilarious though.