AI haters build tarpits to trap and trick AI scrapers that ignore robots.txt(arstechnica.com)

posted 2 days ago

pelespirit@sh.itjust.works

Building on an anti-spam cybersecurity tactic known as tarpitting, he created Nepenthes, malicious software named after a carnivorous plant that will “eat just about anything that finds its way inside.”

Aaron clearly warns users that Nepenthes is aggressive malware. It’s not to be deployed by site owners uncomfortable with trapping AI crawlers and sending them down an “infinite maze” of static files with no exit links, where they “get stuck” and “thrash around” for months, he tells users. Once trapped, the crawlers can be fed gibberish data, aka Markov babble, which is designed to poison AI models. That’s likely an appealing bonus feature for any site owners who, like Aaron, are fed up with paying for AI scraping and just want to watch AI burn.

Sort:

Hot Top Controversial New Old

You are viewing a single thread.

View all comments View context

[ - ]

vrighter@discuss.tchncs.de

21 points

2 days ago

an infinite loop detector detects when you’re going round in circles. They can’t detect when you’re going down an infinitely deep acyclic graph, because that, by definition doesn’t have any loops for it to detect. The best they can do is just have a threshold after which they give up.

permalink

report

parent

[ - ]

LovableSidekick@lemmy.world

-5 points

2 days ago

You can detect pathpoints that come up repeatedly and avoid pursuing them further, which technically aren’t called “infinite loop” detection but I don’t know the correct name. The point is that the software isn’t a Star Trek robot that starts smoking and bricks itself when it hears something illogical.

permalink

report

parent

[ - ]

Crassus@feddit.nl

8 points

2 days ago

It can detect cycles. From a quick look at the demo of this tool it (slowly) generates some garbage text after which it places 10 random links. Each of these links loops to a newly generated page. Thus although generating the same link twice will surely happen. The change that all 10 of the links have already been generated before is small

permalink

report

parent

[ - ]

LovableSidekick@lemmy.world

-4 points

2 days ago

I would simply add links to a list when visited and never revisit any. And that’s just simple web crawler logic, not even AI. Web crawlers that avoid problems like that are beginner/intermediate computer science homework.

permalink

report

parent

[ - ]

dev_null@lemmy.ml

9 points

2 days ago

They are no loops and repeated links to avoid. Every link leads to a brand new, freshly generated page with another set of brand new, never before seen links. You can go deeper and deeper forever without any loops.

permalink

report

parent

[ - ]

LovableSidekick@lemmy.world

1 point

1 day ago

You can limit the visits to a domain. The honeypot doesn’t register infinite new domains.

permalink

report

parent

Show more comments

[ - ]

vrighter@discuss.tchncs.de

2 points

2 days ago

sure, if you have enough memory to store a list of all guids.

permalink

report

parent

[ - ]

LovableSidekick@lemmy.world

1 point

1 day ago

It doesn’t have to memorize all possible guids, it just has to limit visits to base urls.

permalink

report

parent

Show more comments

Technology

!technology@lemmy.world

Create post

This is a most excellent place for technology news and articles.

Our Rules

Follow the lemmy.world rules.
Only tech related content.
Be excellent to each other!
Mod approved content bots can post up to 10 articles per day.
Threads asking for personal tech support may be deleted.
Politics threads may be removed.
No memes allowed as posts, OK to post as comments.
Only approved bots from the list below, to ask if your bot can be added please contact us.
Check for duplicates before posting, duplicates may be removed
Accounts 7 days and younger will have their posts automatically removed.

Approved Bots

Community stats

16K
Monthly active users
13K
Posts
592K
Comments

Our Rules

Approved Bots

Community stats

Community moderators