A Project to Poison LLM Crawlers

Disillusionist@piefed.world · 2 days ago

A Project to Poison LLM Crawlers

vane@lemmy.world · 7 hours ago

I have around 10-20GB github / gitlab mirror. I am constantly under attack from crawlers from top US technology corporations and LLM startups. Whenever I ban one IP range they switch to other - I don’t know if those fuckers have tickets in their systems to do it manually or they just deploy this shit all over the planet. From what I observe during attacks that I mitigate the best way to poison them is to just create gitea instance with poisoned code repository and couple hundred revisions. It’s because what they are most interested in is html representation of diff between two git revisions.

A Project to Poison LLM Crawlers

A Project to Poison LLM Crawlers

RNSAFFN