User avatar
eri neofox_flag_nb fox_floofing @eri@mk.moth.zone
1w
blocking random IPs and shoving everything through mazes of captchas and proof of work and giving poisoned data to anyone who isn't using a "correct" web browser to try and stop "AI scrapers" it just feels a bit like we've created a medicine that's worse than the disease here
5
1
1
0
User avatar
eri neofox_flag_nb fox_floofing @eri@mk.moth.zone
1w
you guys do realize that all these hallmarks of "AI" browsers/scrapers are also the same things that accessibility tools do, right? because both interact with a webpage as a big pile of text, not as a bunch of 2d images you can click on.

but it's fine. we should be feeding deliberately fake data to any browser that doesn't act exactly like the latest version of google chrome with all tracking features enabled. we should block entire ASNs permanently because a single connection did something that looked kind of weird in one of your access logs. we should force every device to burn compute hours solving proof of waste puzzles because we need to fight those evil AI companies wasting compute hours. we are the good guys and we are doing the right thing.
2
1
0
0
User avatar
eri neofox_flag_nb fox_floofing @eri@mk.moth.zone
1w
clearly, the only solution is to ensure that everyone connecting to your server is actually a person. i propose that every ID card should contain a smart chip which contains a cryptographic key. by verifying the connection has a valid user, you can ensure that there are no more automated scrapers! simply rate limit each individual person's access. This system will not be used by any other people for any other things, and will be a net good for society and the freedom of information
2
1
0
0
User avatar
eri neofox_flag_nb fox_floofing @eri@mk.moth.zone
1w
those who judiciously block random traffic are just as evil to me as big tech companies running poorly behaved crawlers. i hate both of you because you both want to stand over your little fiefdoms of the information you hoard privately as if the capacity of your hard drives is a proxy for the size of your metaphorical penis, and your kink is making people beg to get just a few more nybbles of that sweet sweet data from you.
4
0
0
0
User avatar
Consensus Tullyality @tully@cathode.church
1w
@eri I get this, but AI scrapers are literally DDoSing every website they know about, so the options in 2026 seem to be:
have a website that doesn't work for anyone because it's overloaded 100% of the time
have a website that rejects Chrome user agents specifically (because all the scrapers are pretending to be Chrome)
don't have a website
5
2
2
0
User avatar
eri neofox_flag_nb fox_floofing @eri@mk.moth.zone
1w
@tully that's weird because they aren't doing that to my website
1
0
0
0
User avatar
illy [Shrimple-mode] protomoji_orange_flag_lesbian @illyBytes@shrimp.imsofucking.gay
1w
@eri @tully your lucky then :p
1
0
1
0

User avatar
illy [Shrimple-mode] protomoji_orange_flag_lesbian @illyBytes@shrimp.imsofucking.gay
1w
@eri @tully ive seen friends actually get their stuff assaulted by bots scrapping
1
0
1
0
User avatar
Consensus Tullyality @tully@cathode.church
1w
@illyBytes @eri yeah… I'm dreading the day the scrapers find our NextCloud instance, or our GTS, or any of the web-facing things we have
1
0
1
0
User avatar
eri neofox_flag_nb fox_floofing @eri@mk.moth.zone
1w
@tully @illyBytes if you have an HTTPS cert they already know about your stuff since those are all publicly logged. maybe there's a reason only some sysadmins are having problems. maybe that reason is poorly programed applications being ran by people who don't understand how they work
2
0
0
0
User avatar
illy [Shrimple-mode] protomoji_orange_flag_lesbian @illyBytes@shrimp.imsofucking.gay
1w
@eri @tully yes, bots scrap publicly available stuff, ofc thats literally what they do lol
0
0
0
0
User avatar
Consensus Tullyality @tully@cathode.church
1w
@eri @illyBytes if that's the case then we're fucking doomed, because literally none of our stuff is set up well lmaoooo

like, our NC is years out of date because the RPi it's running on is running a 32-bit kernel and updating it at this point is likely to need a full system rebuild.

tbh I think the reason we haven't been destroyed by scrapers is much simpler: we don't have any incoming hyperlinks. which, if we were ever to set up an actual site that had things on it that others wanted to read, would quickly change.
0
1
1
0