frontpage.
newsnewestaskshowjobs

Open Source @Github

fp.

DeepSeek Harness

https://www.deepseek.com/en/harness/
127•Kuyawa•3h ago•39 comments

Pi 1.0

https://earendil.com/posts/pi-1-0/
1050•sergiotapia•10h ago•328 comments

How Singapore's government-run dating service works

https://www.singapore-samizdat.com/p/how-singapores-government-run-dating-service-firstdate-works
82•danielfoster•4h ago•37 comments

Shimano Bicycle Museum Review

https://inrng.com/2026/10/shimano-bicycle-museum/
15•pietroppeter•1h ago•0 comments

Clef: Open-weight decision models, and new RL fine-tuning platform

https://blog.cloudflare.com/clef-decision-models/
480•jasondavies•14h ago•171 comments

SvelteKit 3

https://svelte.dev/blog/sveltekit-3-is-here
206•sampsn•10h ago•68 comments

Several vulnerabilities have been discovered in the Linux kernel

https://lwn.net/Articles/1097401/
224•luispa•7h ago•147 comments

Automatic Transmission – a data-privacy study of connected vehicles

https://automatictransmission.khoury.northeastern.edu/index.html
161•rafaelc•10h ago•153 comments

Pi Durable

https://earendil.com/posts/pi-durable/
314•paulsmith•11h ago•38 comments

Ask HN: Who is hiring? (October 2026)

177•whoishiring•15h ago•178 comments

Turbo Haskell

https://comonad.com/reader/2026/turbo-haskell/
58•pjmlp•16h ago•4 comments

StreetComplete on iOS is now in public beta

https://github.com/streetcomplete/StreetComplete/issues/5421
549•Snowly•19h ago•144 comments

Git 3.0's upcoming SHA-256 default will be a costly mistake

https://blog.gitbutler.com/git-3-sha-256
305•chmaynard•13h ago•299 comments

Using Opus 5.5 to discover a new eyewitness record of the dodo

https://resobscura.substack.com/p/using-opus-55-to-discover-a-new-eyewitness
114•benbreen•9h ago•28 comments

To grieve, or not to grieve?

https://xenaproject.wordpress.com/2026/10/01/to-grieve-or-not-to-grieve/
40•stabbles•20h ago•5 comments

CSS Bed: Classless CSS themes to use as starting points in web development

https://www.cssbed.com
94•sea-gold•9h ago•25 comments

RIP, vector database

https://turbopuffer.com/blog/rip-vector-database
301•razin•14h ago•82 comments

ArXiv's Updated Rate Limit Policy

https://blog.arxiv.org/2026/10/01/updated-rate-limit-policy/
101•50kIters•10h ago•39 comments

Building reliable (and fast) directory sync

https://www.firezone.dev/blog/building-reliable-directory-sync
3•jamilbk•1h ago•0 comments

Show HN: Janus – Go binary that runs GGUF models via Vulkan on AMD/Intel/Nvidia

https://github.com/Vibra-Ingenn/Janus
64•Maverick617•9h ago•9 comments

Vote on which of Hacker News' challenges for AI have been met

https://stoppels.ch/goalposts/
127•stabbles•12h ago•126 comments

Frog and Toad and the Increasingly Capable Machines

https://www.frogandtoad.ai/
109•supermdguy•8h ago•22 comments

Various Projects Find Hidden SDR Capabilities in ESP32 Microcontrollers

https://www.rtl-sdr.com/various-projects-independently-find-hidden-sdr-capabilities-in-esp32-micr...
200•nkw•15h ago•30 comments

Oxygen-deprived underwater zones may not be “dead zones” but clue to early life

https://agupubs.onlinelibrary.wiley.com/doi/10.1029/2026AV002570
92•gumby•11h ago•6 comments

Cloudflare K2: serverless event streams

https://blog.cloudflare.com/cloudflare-k2-streams/
226•elffjs•16h ago•89 comments

Butterflies use optical illusions to dodge predators

https://www.essex.ac.uk/news/2026/09/30/butterflies-use-optical-illusions-to-dodge-predators
35•gmays•7h ago•5 comments

Context Language Models

https://arxiv.org/abs/2609.37725
137•emersonmacro•15h ago•29 comments

Bez: Generating a browser engine from specs and tests

https://tangled.org/burrito.space/bez
108•nerdypepper•12h ago•44 comments

US tells France and Germany to release diesel stocks or face US export ban

https://www.reuters.com/business/energy/us-tells-france-germany-release-diesel-stocks-or-face-us-...
12•geox•1h ago•2 comments

How to speed up the Rust compiler in September 2026

https://nnethercote.github.io/2026/09/30/how-to-speed-up-the-rust-compiler-in-september-2026.html
243•trickypr•17h ago•130 comments
Open in hackernews

Meta's Muse is fantastic for web scraping

https://sigh.dev/posts/metas-muse-is-fantastic-for-web-scraping/
36•STRiDEX•1h ago

Comments

dvt•59m ago
I built an AI "web harness" running on a sandboxed Chromium (using a custom side-loaded plugin that talks over websockets to a "driver") to basically do anything a normal user could do in a browser. It totally bypasses any and all bot measures and only gets the ones you yourself would get as well (and passes those successfully, e.g. Cloudflare checkbox or those annoying OCR puzzles).

Not sure if I should release it, but I'm sure more people are catching onto the power of agentic browsing.

ares623•53m ago
By describing it here you've already released it no?
d0vs•50m ago
FYI Muse smashes through those captchas natively without prompting.
frabcus•49m ago
Right, but what happens when everyone uses that at scale? Without agreements and standards it isn't pretty.
dvt•37m ago
It's sad, but this train has left the station a quarter of a century ago imo. Many people have since become billionaires scraping the web without anyone agreeing (Google, Yahoo, and now possibly Anthropic and OpenAI).
csnover•13m ago
There is a huge difference between a clearly identifiable, robots.txt-respecting, largely symbiotic search engine spider versus this new AI-driven casual sociopathy of firing shotguns of deliberately masked bots at sites in ways which benefit no one except for, maybe, the bot-owner.

I agree with you, though, that the future has almost certainly already been decided, and see the eventual outcome of all this selfishness to be mandatory device attestation to access most services on the internet—which will of course never be allowed on any open platform. AI engineers who believed in freedom to compute (or just individual freedom more generally) have committed perhaps the greatest self-own in the history of our species to date.

bayindirh•48m ago
Thanks for letting us know that we need a new layer of detection systems.

Also it’s great(!) to see that we’re going from “but ethics” to “I got mine, who cares”.

Humans are interesting creatures.

pcthrowaway•33m ago
There is no detection method that will prevent AI from accessing systems without also blocking humans. The only thing we can do at this point is throttling.
kees99•17m ago
Agreed on inevitable collateral blockage of real people using real "headed" browser. I'm getting a ton of that already, personally.

Throttling is poor help though. Mass scrapers are using "residential proxy" loophole + rotating UA and other attributes. You can't throttle somebody without identifying them. Unless you're talking about a global rate-limit.

pcthrowaway•11m ago
throttling based on sessions kind of works; we're headed in the direction that sites like Reddit will probably prevent logged-out users from viewing threads (as they already do with mobile devices)

Once the LLMs create sockpuppets to get around that, the web services will need to resort to profiling users more aggressively so that they know which actual human an account corresponds to.

If someone has a malicious browser extension that uses their session to scrape Reddit then, they're probably going to see significant usage obstacles.

We are headed to a very user-hostile place.

afro88•24m ago
I'm becoming more and more convinced that a big source of outrage on the internet is caused by people assuming that all other people are a homogenous blob.

It's not that "we" are going from one thing to another. It's that these are two different people, with different ethical boundaries.

bel8•43m ago
Interesting, how is the chromium sandboxed? profile dir cli param? or deeper like chromium engine framework embeded in the application?
dvt•41m ago
Just simply a separate Chromium binary (not your usual browser).
rjtc•35m ago
anyone can spin these kind of side projects and do easy talk, but the moment you actually try to use this on signed-in Linkedin or Amazon its going to fail

the only solution is to drive your regular browser with all your sessions/cookies via an extension

dvt•32m ago
> the only solution is to drive your regular browser with all your sessions/cookies via an extension

This is exactly what I'm doing, but mocking/randomizing all the sessions/cookies/params (like resolution, OS, WebGL , etc.) in a separate Chromium binary. It's popular these days, but imo using your normal browser for agentic stuff is a very bad idea. These models do dumb stuff all the time.

carsoon•28m ago
no it doesn't fail. Agents are very good at comparing traffic characteristics from a real browser and a headless/automation browser and getting it to behave in the same manner. It's a cat and mouse game for sites stopping unauthorized access but right now llm agents are ahead.

For testing proxies should be used to avoid IP ban issues but given enough time modern agents can figure out how to bypass most of the modern scraping/automation prevention mechanisms.

I have built and used a lot of different automations and web scraping implementations for my business and it's never got permanently stuck yet, some take a bit longer, some shorter, but all within a reasonable time with little external help they have succeeded in their tasks.

kurisufag•31m ago
the 4get dev did the same thing a little while ago for his metasearch engine: https://git.lolcat.ca/lolcat/4play
dvt•23m ago
Yep, it's super similar to this! I think his is a bit overengineered, but tbh I haven't built Firefox plugins in forever so that might be the way you have to do stuff on FF. All you need is the websocket plugin to have "full access" to all visited webpages, and then your driver just hooks into the DOM like normal (with the extra ws layer—and even the websocket layer can be simplified away I think).
pprotas•28m ago
Camoufox bypasses most blocks with no problems https://camoufox.com/
onion2k•28m ago
(using a custom side-loaded plugin that talks over websockets to a "driver")

Is that necessary? You could start Chromium with an open debug port and use Chrome Devtools Protocol to send commands.

dvt•20m ago
It is, because `--remote-debugging-port` is detected via the root DOM object (and that can't be changed unless you want to recompile Chromium), so for example, if you try doing a Google search with debugging enabled, you'll get blocked (usually just by being served a blank page).
LunaSea•16m ago
Interesting, do you know what exactly changes on said root DOM node in case the remote debugging feature is enabled?
dvt•13m ago
`navigator.webdriver` afaik, but I think there's a couple more, I did some research on it a few months ago. If using Playwright/Puppeteer, they also have a few special things injected that you have to strip.
Cakez0r•3m ago
What model did you use to build this? I recently looked in to building something similar and got refusals.
raffraffraff•7m ago
Does remote debugging port itself get detected, or does the presence of a webdriver connection to the port get detected? I believe it's the latter.

I've had success launching the browser and using dumb dumb methods to get around the captcha before attaching playwright.

Dumb dumb methods = wmctrl, xdotool, bash (work fine for Cloudflare's "are you human" check)

kees99•12m ago
Some webshops (e.g. Aliexpress) already soft-block Chromium (at least Chromium-on-Linux). I.e. such users see larger-than-usual share of captchas.

Assuming "more people catching onto" this, expect Cloudflare and most everyone else to follow the suite.

dvt•9m ago
Are you saying they're blocking the `User-Agent`? Because that's trivial to bypass. I'm not exactly sure what you mean by "Chromium" because that's a whole class of browsers (Edge, Brave, Chrome, etc. are all forked "Chromium" browsers).
kees99•4m ago
Aliexpress specifically uses a javascript blob to do detection, and quite a sophisticated one. There was a recent story about that blob messing with bluetooth headphones.

> not sure what you mean by "Chromium"

This: https://www.chromium.org/getting-involved/download-chromium/

Google Chrome doesn't get flagged in the same way. Haven't tried anything else from your list.

Cakez0r•7m ago
I think it's inevitable that eventually one of the big AI companies makes something like this. It's such an obvious consumer win. "I am not a bot" checkboxes are a string and peg tethering the elephant.
arjunchint•53m ago
their static ip's were initially good and didn't get flagged, but now most sites are recognizing their ip ranges and blocking.

Muse's utility has significantly dropped with the blockages.

To become truly useful again they will need to use residential proxies, but I can't see them use those due to the risks and reputational damage.

sejje•29m ago
they can just use the user ip. i think grok already does this.
koolala•18m ago
How can it do this? Wouldn't you see a hundred fetch requests in your browser network tab?
Gareth321•25m ago
Most residential proxies are already far more blocked and rate limited than any Meta IP. The internet is becoming a very weird place, where individual and "trusted" personal IPs are becoming a kind of commodity. Some sites are already scoring IPs based on usage activity - like a credit score. It's only a matter of time until this data is collated and commoditised. AI analysis is turning this up to 11.
iamacyborg•9m ago
That doesn’t track with reality as far as I can tell.
simoncion•7m ago
> Some sites are already scoring IPs based on usage activity - like a credit score. It's only a matter of time until this data is collated and commoditised.

Spamhaus is nearly thirty years old and the notion of electronic distribution of IP and domain "reputation" lists is at least that old.

I'll bet my hat that the Internet "advertising" [0] industry has been calculating and determining the reputation of individual households (if not individual users) for at least a decade.

[0] The scare quotes are because its primary purpose these days is for dragnet private-sector surveillance.

kevmo314•28m ago
Muse ran into a captcha and asked me if I wanted it to solve it.

So of course I clicked yes and it dutifully convinced the site that it was not a bot.

xnickb•24m ago
For 2(3?) decades we've been training the robots to tell traffic lights from fire hydrants. It's finally paying off.
Gareth321•23m ago
I think this is the end of sites where it is expected that only humans may interact with them. It's been cat and mouse for a while, and some places like Reddit sell API access, but these agents are for all intents and purposes, humans interacting with the site. They're going to need to figure out new business models.
mejutoco•17m ago
Or the captchas will evolve.
therein•13m ago
This will be used as an excuse to normalize "scan QR code with your phone to verify you are human" along with Web Environment Integrity stuff.
wiether•16m ago
As someone having to fight Meta's bots everyday to keep websites accessible to actual customers, I'm not surprised to read that it can be seen as something positive on the other side of the fence.

But I'm wondering: at what cost?

STRiDEX•11m ago
I think for my own side projects i would require the user to login if they were making requests from those ip addresses or block.
gunalx•12m ago
Im guessing more of the internet will be login walled from now on.
Cakez0r•14m ago
You're framing this as if people are deliberately making decisions that they believe are unethical. The reality is that people have different ethical frameworks. For example, I believe that there is no ethical distinction between whether a web request originates from a browser or from an LLM on my behalf.