frontpage.
newsnewestaskshowjobs

Open Source @Github

fp.

What 2026 looks like (2021)

https://www.alignmentforum.org/posts/6Xgy6CAf2jqHhynHL/what-2026-looks-like
1•mbeavitt•1m ago•1 comments

Coding agents supercharge Murphy's Law

https://hatchet.run/blog/murphys-law
1•abelanger•2m ago•0 comments

Prompt Caching in Agents

https://earendil.com/posts/prompt-caching/
1•lobo_tuerto•3m ago•0 comments

Builders lack an audience, Creators lack building skills

1•yurimhln•3m ago•0 comments

Neutrinos from Deep Inside Earth Provide a New Picture of the Mantle

https://www.quantamagazine.org/neutrinos-from-deep-inside-earth-provide-a-new-picture-of-the-mant...
1•lschueller•3m ago•0 comments

Evolution, Not Reset: Prepare Platform Engineering 2.0 for Autonomous Agents

https://platformengineering.com/features/evolution-not-reset-prepare-platform-engineering-2-0-for...
2•CrankyBear•4m ago•0 comments

McFarthest – The Greatest Distance from McDonald's

https://www.themeateater.com/conservation/public-lands-and-waters/bar-room-banter-mcfarthest-the-...
1•Aeroi•6m ago•0 comments

Harica Revoked 63,525 SSL Certificates in Five Days

https://tlsradar.com/blog/harica-mass-revocation-2026
1•tlsradar•6m ago•0 comments

Signal planning one-time payment for optional sign up without a phone number

https://aboutsignal.com/news/signal-login-registration-without-a-phone-number/
3•rvz•7m ago•0 comments

Show HN: Vaultak – Security for AI agents (built before the breaches started)

https://vaultak.com
1•samoladji•7m ago•0 comments

Automated and easy EU AI Act compliance

https://scanara.io/en/
1•scanara•7m ago•0 comments

Best temp mail Firefox extension

https://addons.mozilla.org/en-GB/firefox/addon/tempfox-multi-mail-otp/
1•skyyfisk•7m ago•0 comments

Neal: Codex and Claude in a loop shipped a 549-commit migration

https://navels.dev/blog/neal/
1•navels•8m ago•0 comments

Origins of Life on Earth by Tracing Early Chemical Reactions

https://www.smithsonianmag.com/science-nature/a-new-study-points-to-two-origins-of-life-on-earth-...
1•nhatcher•9m ago•0 comments

Decompiling the Android Developer Verifier App

https://blck-b.github.io/post/android-verifier/
1•blck-b•10m ago•0 comments

Turning Claude into Postgres so I can raise a Series A

https://byteofdev.com/posts/turning-claude-postgres/
1•JeanSebTr•11m ago•0 comments

Show HN: We Built MicroLeague Sports Vol. 3

https://medium.com/@esolar_18069/why-we-built-microleague-sports-vol-3-f3d7cedfed9d
2•esolar07•11m ago•0 comments

Nobody Died – and That's the Scary Part – Oshkosh 2026 [video]

https://www.youtube.com/watch?v=4rJEtpQ0r6A
2•marklit•13m ago•1 comments

Show HN: A workflow for building community skill catalogs

https://github.com/thinkingoracle/luke-skills
1•noobdabber•15m ago•0 comments

Qwen 3.8-Max Preview

https://manish.sh/writings/models/inside-qwen-3-8-max-preview-reverse-engineering-an-ai-assistant...
2•ms7892•16m ago•0 comments

Building an AI Chat for My Blog: Gemini, Cloud Functions, $5-15/Month

https://emergencemachine.com/building-an-ai-chat-for-my-blog/
1•speckx•17m ago•0 comments

Authentication and User Management Software: FusionAuth [audio]

https://sourceforge.net/articles/authentication-user-management-software-fusionauth-sourceforge-p...
1•mooreds•20m ago•0 comments

Improving mouse traps with Home Assistant

https://highlyprobable.io/posts/improving-mouse-traps-with-automation
1•darkpicnic•20m ago•0 comments

US Wildfire Incident Information System

https://inciweb.wildfire.gov:443/
3•mooreds•21m ago•0 comments

Simon Willison on Technical Blogging

https://writethatblog.substack.com/p/simon-willison-on-technical-blogging
1•mooreds•21m ago•0 comments

A Year of Saving Our Signs

https://www.datarescueproject.org/a-year-of-sos/
1•hn_acker•22m ago•0 comments

An all-sky map of half a million supermassive black holes

https://www.sdss.org/black-hole-mapper-release-20/
1•MarcoDewey•23m ago•0 comments

Even The New York Times "isn't immune" to declining search traffic

https://www.niemanlab.org/2026/08/even-the-new-york-times-isnt-immune-to-declining-search-traffic...
1•ChrisArchitect•23m ago•0 comments

I wrote my first C project after a long time. Unix wc clone. NO AI

https://github.com/emanueleoggiano/word_count
2•emanueleoggiano•23m ago•1 comments

Towards a Local-First Desktop (Guadec 2026) [video]

https://www.youtube.com/watch?v=qnF9NzBidPA
2•amlib•24m ago•0 comments
Open in hackernews

99% of My Website Traffic Is Bots

https://patronview.com/news/99-percent-of-my-website-traffic-is-bots/
95•petercooper•56m ago

Comments

lazerg•53m ago
Even this article was written by AI. So why are you angry at bots visiting your website?
Yiin•44m ago
what gave you this idea?
pixl97•40m ago
Because the ubermensch on HN can detect AI text from 10 miles away with 100% accuracy... like that time they called documents from 2015 AI written.
conartist6•30m ago
I'm just beyond thrilled that it's actually starting to be a public embarrassment to cough up a glob of AI text.
chrisandchris•38m ago
Take a look at the dashes. It's all there.

(i skimmed the whole post, there are none)

rhdunn•33m ago
I knew it... Wuthering Heights was written by AI and Lucy Maud Montgomery, Edgar Allan Poe, et. al. were AI bots churning out content!

Or maybe -- just maybe -- using dashes isn't a sign of content being AI written, just a style it picked up from the training data.

askl•37m ago
I didn't read the article, but the AI illustrations were off-putting enough to close the tab.
runjake•31m ago
Not the OP, but because it appears structured just like AI output?

(This isn't a condemnation. AI can often do a better job of representing thoughts than humans.)

Examples:

- The article organization

- The general language flow

- The bullet points with a bolded gist, a colon, and then elaboration (and bonus with details stats).

- The images are almost certainly AI generated. They look AI generated.

conartist6•28m ago
That's the thing: even if you don't use AI, if most of what you read is AI, soon this is what you'll sound like, just because it's so much of the training data your own brain's model has.

Using it slowly sucks the uniqueness out of you.

throwaway219450•21m ago
Using AI to write blogs doesn’t bother me in principle, but I hate the prose that gets left in. Even if it’s not AI writing, being human doesn’t give you a pass for writing like a self-help guru. This sort of grammar is straight out of Claude:

> A human on a VPN sees one CAPTCHA and passes, but a headless browser fleet sees a wall.

I do wish we’d stop complaining about em dashes though, that’s lazy criticism. The sentence structure is far worse, and a big tell is subheadings that are all variants of “The <adjective> <noun phrase>”.

behole•36m ago
HN's blanket AI allegations are almost as annoying and tired as the thing they are rallying against. I think you are AI.
bookofjoe•30m ago
>HN's blanket AI allegations are almost as annoying and tired as the thing they are rallying against.

Best thing I've read on HN so far this year.

mysterydip•30m ago
That’s exactly what an AI prompted to reply to AI accusations would say!
pixl97•27m ago
I'm not a bot, you're a bot!

So this is how AI wins, humans kill each other off because we might be bots and the bots inherit the earth.

qbane•40m ago
> And yes, my site gets its data by scraping those public documents. So I'm a scraper writing a blog post complaining about scrapers. I'm aware of how that sounds.
ethersteeds•25m ago
Who scrapes the scrapemen?
ihuman•15m ago
There's a difference between someone running a scraping tool occasionally and bots constantly and rapidly re-scraping the same site over and over again
alansaber•11m ago
Live by the scraper, die by the scraper
aeturnum•10m ago
Similar to the dose making the poison - the thing that jumped out at me in this blog was the ratio of scraping to visits. Unless OP is scraping thousands of times a day I don't really think they're in the same class as the bots they are blocking.
nickgray•6m ago
OP here: I'm not scraping thousands of times per day! Usually just a few times per year.
dzonga•39m ago
blocking by geo yeah might work - but what happens when someone is traveling abroad ? they've to use a VPN to access your site ?

my take with all the bots - the web is gonna be a bunch of private walled gardens. with most sites set to no index. you will only discover them via referral from someone real.

yoursred•36m ago
What happens when someone is from a shit country?
mcraiha•31m ago
AFAIK you already have to use VPN if you are e.g. westerner visiting China or Belarus.
zuzululu•38m ago
pretty crazy that this article is seemingly written by an AI used a detector and it is 95% confident its generated
storus•37m ago
Is there any cheap way to run personal websites without getting cooked by bots these days, degrading the performance? A $5/month VPS won't cut it anymore.
yoursred•35m ago
Sef-host with tailscale or something similar
rglover•34m ago
All of my boxes are cheap VPS. Highly recommend people throw Cloudflare in front of their stuff. I switched all of my load balancers over to there and all of that bot crap went away. That combined with a proper UFW setup keeps the weather clear for me.
inigyou•19m ago
Highly do not recommend centralising the internet.
Athas•30m ago
I run my personal website (and a bunch of other websites and services) off a somewhat more expensive but still reasonable VPS (I think 20€ at TransIP - it's so little that I forgot). Load is basically nil most of the time anyway. I think the bot problem is not so bad for personal websites.
archerx•28m ago
A lot of my sites are on $5 vps and run very well.
ddxv•37m ago
I'm in a similar boat. Probably 99.999% is bots. I have nearly 1m unique "visitors" according to cloudflare and my real users are in the dozens a day. That being said, I love the open internet and am holding on to keeping as much open as I can.
andai•37m ago
> There's a real conflict here. I want Google and Bing and DuckDuckGo to crawl my site and send me new readers. But I don't want everyone else strip-mining it.

Kinda sounds like we're missing a peer to peer network here.

Instead of downloading the same data over and over again we can just download it once and then share it.

Wouldn't that be better for everyone involved?

It would also function as a distributed WayBack Machine, in case anything ever happens to the Internet Archive. (Which I think is desperately needed, bot apocalypse aside.)

amazingamazing•33m ago
The problem there is trust
NicuCalcea•33m ago
Common Crawl?

https://commoncrawl.org/

inigyou•21m ago
Anyway, Google doesn't send traffic to your site any more. Only important sites and obvious scams seem to get indexed.
righthand•33m ago
> Then in November 2025, four thousand "visitors" showed up over a few days. Each visited exactly one page with a bounce rate of 99%. More telling was that they had no referrer. That's usually the easiest way I spot a bot.

> And they were only crawling my fund pages (like this one, this one, and this one), which only 10% of real visitors ever touch.

> But those 4,000 bots were just the warm-up.

I just hate this style of writing like you're on Twitter. Why does the above need to be 3 different paragraphs? A paragraph break indicates a separate thought but the author is still talking about the same data and still making their point. The sentence "But those 4,000 bots were just the warm-up." is effective when still the last line of a paragraph and it signals respect for your readers. I stopped reading after this because it's just a terrible reading experience.

Here's a correct version that doesn't read like the author left for a week to think about what the next sentence would be or having some sort of anxiety-induced mental pause:

> Then in November 2025, four thousand "visitors" showed up over a few days. Each visited exactly one page with a bounce rate of 99%. More telling was that they had no referrer, which is usually the easiest way I spot a bot. They were only crawling my fund pages (like this one, this one, and this one), which only 10% of real visitors ever touch. But those 4,000 bots were just the warm-up.

nickgray•16m ago
Hey! I'm the OP - thanks for feedback on my writing style. I went ahead and fixed this in the article. It should be updated by the time you read this:

https://patronview.com/news/99-percent-of-my-website-traffic...

And you're totally right: I mostly post on X (nee Twitter) and I probably have ADHD or just a low attention span, so I prefer to read things broken up into paragraphs. But for a smarter audience like this, and that reads long-form blog posts, I should tighten it up.

Thank you for the suggestion. LMK any other edits and I'll be happy to tighten it up.

johnorourke•32m ago
Anubis[1] is a superb fix for sites not behind Cloudflare/Fastly/Bunny etc. We had millions of bot requests, on a site serving all countries so we couldn't block by country, with fake user-agents so we couldn't block using that. It uses 'proof of work' to detect real browser software.

[1] https://anubis.techaro.lol/

drum55•28m ago
Which is trivially bypassed by an actual implementation of the proof of work in non-javascript, rendering it absolutely useless. The website is approximately 3800x times slower than native code, and hundreds of thousands of times slower than the CUDA kernel claude wrote. The "proof of work" is just non existent at that point, they're solved in milliseconds for what would take the browser version 10 minutes or more, it's security by obscurity being dressed up as something more.

  pow_server  http://127.0.0.1:8080   backend avx512-x16
  ──────────────────────────────────────────────────────────
  uptime   00:03:12
  solver   ● BUSY  difficulty 9, 0.3s
  queue    [####################............] 5/8   peak 12
  ──────────────────────────────────────────────────────────
  accepted 1240        solved 1180
  503 shed 48      504 timeout 2      4xx/5xx 10
  ──────────────────────────────────────────────────────────
  last     difficulty 5  nonce 645376  in 9 ms  (101.6MH/s, avx512-x16)
  hashes   3.90GH total   avg 65.3MH/s
  Ctrl-C to stop
Claude even made a nice little API server for it after implementing midstate compression, AVX multi way hashing, and a CUDA kernel. This doesn't stop the literal LLM it's trying to block from solving the challenges, it's really annoying that everybody is using it and claiming that it's something that's usable in the real world as a result of it using proof of work. It's obscure, and obscure is fine so long as nobody is pretending that it is secure.
gum_wobble•24m ago
how so, can you link to any sources?
Bender•26m ago
Seems about right. I rotated my logs this morning. Most humans go to access.log and most bots go to botpoop.log. This is the line count:

     2 access.log [1]
    40 botpoop.log [2]
2 is really 1 since a human will grab the CSS file. Most bots do not bother with the style-sheet so that's a 40:1 bots to humans. I could cut that down by blocking data-centers but then I inadvertently block a lot of VPN's which I really don't need to do for a static compressed blog served from ram. The bots just get a TCP Reset but it's still fun to log and study them. The most interesting one I've seen recently is ReadYou which may be a reader but it appears to be much more, possibly acting as a cell phone distributed bot collecting data for a centralized site.

[1] - https://nochan.net/logs/access.log

[2] - https://nochan.net/logs/botpoop.log

spockz•4m ago
[delayed]
harshreality•26m ago
Not mentioned: Anubis.

There's nothing preventing bots from running JS and solving challenges, or implementing native code solvers. In practice, however, they rarely do. If bots are degrading your sites's performance, or substantially increasing costs, or you just don't like seeing 500:1 bot:human ratios, it's an option.

It's also much better for the average human visitor than using, say, cloudflare's interactive challenge mode (like under-attack mode). Anubis introduces a mandatory short (tunable) delay for a cookie that expires after a week by default. Cloudflare challenges require interaction, and are often configured to be much more frequently than anubis's default.

It may not work a year from now. So what? The open internet may not be usable a year from now.

Venn1•25m ago
I'm blocking the Amazon search crawler, anything coming from Googleusercontent, and limiting AI crawlers to search rather than allowing AI assistants. The residential proxy waves are something to behold, but Cloudflare does an okay job catching those in the AI labyrinth. Still, it's all a bit silly, and I can't imagine what large sites deal with when I'm tangoing with this much nonsense on a small tech blog.
inigyou•18m ago
They just don't. Just serve the page unless it's a really expensive page to serve.
ashu1461•23m ago
I wonder if the author tried out the recently released feature by cloudfare to block ai bots

https://developers.cloudflare.com/bots/additional-configurat...

smolder•8m ago
Some people don't want to use cloudflare on principle. Like that putting the whole internet behind cloudflare or AWS is a bad thing, in principle.
mgbmtl•23m ago
I run scripts on my servers on an hourly basis to check which are the top 25 IPs visiting the server (aggregated by /24). If anyone in those top 25 IPs are from China, Vietnam, etc, or from Alibaba/Amazon/etc, the /24 gets blocked by iptables.

It's far from perfect, but it was a quick way to get rid of bots, while not completely blocking people from countries such as Vietnam.

However, on a Gitlab instance I manage (500 users), we have to restrict viewing of git logs and pretty much everything except issues. The bots were too aggressive. Chinese crawlers have access to a huge range of IPs and they often do only 10-20 requests per day, while generating in total over 50k requests per day. Our server load went from 99% down to 0.1% after that (and it's a fairly big server).

tarr11•20m ago
> My normal bill for running this whole site is around $90 a month. During one bad spike month, it jumped about 500%.

This is D1 - which has very surprising costs. you may just want to drop D1 and move to a static site. There’s no reason your site should cost this much.

inigyou•16m ago
There's a lot of people who host their site at extremely expensive places and then do everything they can to minimise unneeded traffic - instead of just moving to a cheaper host. Vercel is another popular extremely expensive host.
nickgray•13m ago
Thank you! I need to tighten up my KV compression, which is actually carrying a lot of D1's load otherwise. We also had some bad queries some months, as the pages and database grew, that were counting the wrong things (or extremely inefficiently) and those have since been fixed.
sp1982•18m ago
The annoying part is a large percentage of misbehaving bots (not obeying robots.txt for example) are via end user proxies across the world. However most of these aren't doing full-browser loop, so if you are behind cloudflare, you can do non-interactive challenge and that can help quite a bit.
nickgray•11m ago
Yes! OP here. I did the non-interactive challenge, and yet all those Chinese bots in my article got through (which surprised me).
AdrianB1•16m ago
I checked the comments to see if anyone pointed to this: I can imagine so many memes with this line :)
FerretFred•15m ago
Sigh .. same here. I don't write blog posts often (enough) but the ones I do write are from personal experiences and I take a lot of care with them. I look at my logs snd see bots everywhere, but now I just let them get on with it. AI scrapers are different though; they get to read my content which, just for them contains a smsttering of finest digital toxin. A pox on your datasets!
nromiun•4m ago
This is a static website running on Cloudflare infra. What on earth costs $90 per month? First optimize yoir infra before throwing up types rules in front of your visitors. I have several websites on Cloudflare too and I don't even check how many millions requests I get. Because it does not cost me anything.

> And yes, my site gets its data by scraping those public documents. So I'm a scraper writing a blog post complaining about scrapers. I'm aware of how that sounds.

Being self aware does not make it okey. Either you are okey with scraping (like me) or against it. Don't use it yourself and block your site at the same time.

These same people will be crying about how Cloudflare ruins the internet because they get these captchas.

ashton314•25m ago
I'm on a $4/mo droplet on Digital Ocean and my static site has been just fine. I'm using Caddy and it seems to handle the load like a champ. My site is very lightweight though, so ymmv.
ashu1461•23m ago
Cloudfare has a very generous plan usually for hobby projects. The article did mention that cloudfare did not work for them, but they have recently introduced few features to block AI crawling as well.
jerf•20m ago
"A $5/month VPS won't cut it anymore."

Are you speaking from experience, or inferring from articles like this?

I serve a static site on the second lowest Linode $5/month VPS and it is grotesquely overprovisioned for that use case. It is not the case that every site is getting slammed every second by hundreds of requests per second.

Now, if you have some sort of dynamically-computed website that is generated by a slow scripting language that is poorly optimized and hits the database too many times for a single page, yeah, it doesn't take many RPS to take you out. But that isn't the only option; it's the slowest of the slow options. Realistic, there are plenty of sites that match that description, but I concatenated that many clauses on purpose. Drop any one of them and your personal site will be fine.

marginalia_nu•14m ago
It's almost always the DBMS that's the bottleneck when websites drop from traffic. As long as you don't do anything fancier than primary key look-ups you're probably fine though.

I survived handling the search search traffic generated by this thread[1] on PC hardware off a residential broadband connection without any sort of degradation. Only time I've gone offline from traffic was when Elon Musk tweeted a link to one of my blog posts, and that was just a short temporary blip.

[1] https://news.ycombinator.com/item?id=28550764

jerf•5m ago
Yeah, IIRC my django site was 3 queries, all correctly indexed, for a main page hit, and 2 for the actual posts. I don't recall the exact perf numbers, but I'm pretty sure it was easily in the 50/rps range for a small dual-CPU host... which doesn't sound like much in "requests per second" but is enough to cover a front-page-HN'ing just fine. And that "rps" was just the database-backed pages, all the static content was served over nginx, so that 50rps is a "real", 50 humans per second rps, not something getting consumed by only two or three humans.
somehnguy•20m ago
A $5/month VPS should cut it completely fine unless you're doing something very complicated.
inigyou•20m ago
Yeah, you just make sure your site is fast enough to handle more than 1 RPS.

But if you want them to actually stop, you can also just serve a little JavaScript page that sets a cookie and refreshes, to anyone who hasn't set the cookie. The DDOS attacker doesn't run JavaScript.

marginalia_nu•18m ago
Static files on literally any hardware from the last 15 years on modern server software simply won't get cooked by bots. The network switch will bottleneck you before the server will. Your ephemeral port range will run out before the server will.
kube-system•14m ago
If you have a static site, GitHub pages is free
strenholme•6m ago
As someone who is moving my static sites over to GitHub pages: While they are free and work really nice, the problem is that GitHub frequently doesn’t deploy updates to the pages.

I have frequently have had to update a GitHub page, push the change, and then GitHub’s actions puke instead of deploying the change. The workaround is that I have a .txt file with a list of GitHub actions which failed, and when GitHub actions fails, I update that .txt file and push the updated site, which GitHub actions will hopefully successfully deploy.

GitHub pages are OK for pages which aren’t updated very frequently, but they are not OK for pages which update frequently.

speak_plainly•13m ago
Cloudflare offers a free plan that's fantastic. The free plan gives you effectively unlimited DNS/CDN traffic for a normal site, while the main practical cap is 100,000 Worker invocations per day, with 10 ms CPU per Worker request and a 100 MB request body limit.

The next tier up from free is $25/month or $240 per year.

https://www.cloudflare.com/plans/ https://www.cloudflare.com/plans/free/

strenholme•11m ago
I have a low cost VPS (actually two in two different pre-AI datacenters) and it runs fine. I use nginx to serve the web pages, and the content is about 99% static content.

The vBulletin and PHPbb style forums have issues with slowdown (I haven’t had a forum since 2015; even back then those forums were overrun with spambots), but static content on a nginx site can be served lightning fast.

drum55•11m ago
The prompt used for Opus 4.8 was:

    write a implementation of the anubis proof of work in native c code, optimized for speed above all else. use every trick available to make the proof of work as efficient and fast as possible, including modern processor tricks on the x86 platform. your code should avoid using external libraries where possible, include tests, and be readable and concise. a reference for what needs to be met is in this repository. https://github.com/TecharoHQ/anubis
Then

    let’s develop this more. turn this solver into a local HTTP server that can be given work in the request, and it returns solved work. make an end to end tester that sends test work to the solver and waits for a valid response. add support for solving with a GPU using cuda.

Then it was done more or less, it happily made a local server that supports solving the challenges given to it in bulk with priority based queue and can tolerate potentially tens of thousands of requests a second with no issue. The CPU time spent solving the challenges is less than the SSL setup for the connections. The GPU version does in excess of 20GH/s though I didn't really test it, I'm not using this for anything but proving a point that the LLM itself can write the bypass tools and run them happily.
inigyou•22m ago
But the people you're defending against don't do that.

They also don't load CSS but for some reason the security theater PoW won the mindshare.

Galanwe•22m ago
The point of PoW access is not that its hard to bypass, it's that you cannot bypass it at scale.
petu•16m ago
If algorithm used is static and GPU-friendly, then what stops bypass at scale?
drum55•12m ago
It's more or less designed for it, it's SHA256 with a break in the middle for midstate compression to be effective, and the difficulty system is based on a misunderstanding of how bitcoin PoW works ("number of zeros" is never, ever a consideration in bitcoin, it's a match to a floating point target).
harshreality•21m ago
It is not absolutely useless, and saying so means you've never had a website getting hammered by bots and experimented with anubis as a countermeasure.

While dedicated scrapers/attackers could work around it, and they could do so much more efficiently than the client-side js, almost none of them do. Unless you like paying additional hosting resource fees to serve bots, it's a worthwhile option, and is less annoying to typical human visitors than cloudflare's interactive captcha/challenge which is what most people use.

Don't let the perfect be the enemy of the good enough. For now.

Targeted attacks may not be repelled at all. That's not the point.

kro•6m ago
It does not even require the PoW thing Anubis does. I've setup a simple logic that just:

Checks for existence of a specific static cookie, if it does not exist, output a small page that sets the cookie via JS and reloads. Sadly this kills Noscript, but it would be possible to add a form in <noscript> that when submitted sets the cookie serverside.

Is this trivial to bypass? Yes. It still keeps out 95% of unwanted bots. Reality is most do not target you specifically they just want to mass-scrape with low effort. Running headless browsers is way more expensive for their op

I've extended this with a FCRDNS checked exclusion for Googlebot.

Another quite effective measure I figured out was checking the existence of Sec-Fetch-Dest header if the User-Agent claims to be a modern browser.

RattlesnakeJake•27m ago
I wish they'd ease up on the whole "don't change the logo without paying us" thing. The furry anime character is a turnoff for anyone with a brand or personal image that doesn't mesh with those subcultures.
gfaster•25m ago
I think that's partly why they do it? If you care about that, you should pay?
smolder•12m ago
Or just have an LLM write the same thing for next to nothing...
inigyou•24m ago
They're not stopping you. They're asking you politely not to.
teddyh•23m ago
If you have a brand or personal image that you are investing in, you can surely afford to invest in paying for a branded version of Anubis.
righthand•22m ago
You wish they'd give your preferable imagery for free and not make you feel bad for using their free software for personal gain. Your subculture is irrelevant and not special.
1bpp•21m ago
If you 'don't mesh' with that then you are not worth protecting anyway :)
tsunamifury•16m ago
I hope I’m missing the sarcasm here…
marklar423•22m ago
I'm assuming a bot running a headless browser instance can still get past it?

It's still valuable to raise the cost of scraping of course. I don't think anything can really stop a determined scraper from impersonating a human. I wonder though if a system similar to Anubis but mining some crypto would make bots _welcome_ - since they're paying for their traffic.

drum55•9m ago
People tried this in 2013 or so, there's no point to it. Doing proof of work in javascript in a browser is so crushingly, pointlessly slow that there's no value at all. Some browsers also intentionally detect attempts to do proof of work and attempt to block it entirely.