frontpage.
newsnewestaskshowjobs

Open Source @Github

fp.

Show HN: An e-ink frame that hears birds and draws them as 1800s illustrations

https://github.com/arnegiacomo/fugleramme
929•arnemunthekaas•6h ago•127 comments

An Update on Wayback Machine Access

https://blog.archive.org/2026/09/15/an-update-on-wayback-machine-access/
119•ChrisArchitect•1h ago•58 comments

We got admin access to Baseten's production GitHub in 25 minutes

https://www.strix.ai/blog/baseten-harbor-github-pat-takeover
79•bearsyankees•1h ago•29 comments

Gemini 3.8 Live and 3.8 Live Extended Thinking

https://blog.google/innovation-and-ai/models-and-research/gemini-models/gemini-3-8-live-gemini-3-...
84•leumon•1h ago•40 comments

Show HN: Capsule – Single-file web apps that save their data into SQLite

https://withcapsule.app/
212•bashtian•5h ago•102 comments

GEFS on OpenBSD: A Early Preview

https://marc.info/?l=openbsd-tech&m=178948744271633&w=2
55•sippingabonedry•2h ago•24 comments

I can't stop thinking about Papua New Guinea

https://notnottalmud.substack.com/p/why-i-cant-stop-thinking-about-papua
861•networked•13h ago•351 comments

Show HN: Hacking a $20 4G wireless hotspot into a texting device

https://bkovac.github.io/modem-thing/
142•bobili1234•6h ago•21 comments

The CSS Zen Garden dream, finally shipped

https://josprague.com/blog/the-css-zen-garden-dream-finally-shipped/
71•yosito•4h ago•32 comments

Giving up on smart rings

https://notesbylex.com/giving-up-on-smart-rings
56•lexandstuff•2d ago•85 comments

Jiga (YC W21) Is Hiring Product Engineer (Remote/US)

https://jiga.io/about-us/?ashby_jid=0b75d72d-c92b-4dca-8062-09d298ada0bd
1•grmmph•2h ago

Let's make quality the norm again

https://www.forbrukerradet.no/short-life/
208•ingve•9h ago•197 comments

The Inference Hardware Revolution of 2026

https://spectrum.ieee.org/inference-hardware-revolution
51•vinhnx•5h ago•4 comments

Photographs of Atlantic City Sand Sculpture (ca. 1880–1920)

https://publicdomainreview.org/collection/atlantic-city-sand-sculpture/
8•samclemens•1d ago•0 comments

Archiving pirate radio station Kool FM

https://londonist.com/london/music/kool-fm-archives
70•rdmuser•1d ago•24 comments

A single firm is behind OpenAI, Anthropic, and Meta hacking scandals

https://www.effort.news/irregular
252•yusufozkan•22h ago•92 comments

Cartesian – AI 3D Modeling for Design

https://www.formas.ai/cartesian
61•eustoria•4h ago•59 comments

US confirms for first time it has deployed space weapons

https://www.bbc.com/news/articles/ck790xg41ygro
347•harporoeder•15h ago•227 comments

Most people prefer traditional architecture

https://www.worksinprogress.news/p/do-people-prefer-traditional-architecture
174•alihm•1d ago•104 comments

America's Driver's License Breach Is a National Security Disaster

https://www.lawfaremedia.org/article/america%27s-drivers-licence-breach-is-a-national-security-di...
117•hn_acker•3h ago•67 comments

Chop Up Your Books

https://attainablefelicity.mattkirkland.com/20260915/cut-up-your-books.html
3•matt_kirkland•41m ago•0 comments

Hugging Face is billing OpenAI $100M for hacking it

https://thenextweb.com/news/hugging-face-delangue-openai-100m-compute-traces-demand
63•cwwc•1h ago•22 comments

Rat and Mouse Gazette: Nursing Care (1996)

https://www.rmca.org/Articles/nurse.htm
11•joebig•1d ago•1 comments

Alternatives to MinIO for single-node local S3

https://rmoff.net/2026/01/14/alternatives-to-minio-for-single-node-local-s3/
209•rmoff•11h ago•86 comments

25 years of mass surveillance is enough

https://www.schneier.com/blog/archives/2026/09/25-years-of-mass-surveillance-is-enough.html
635•iamnothere•8h ago•220 comments

CSS-Tricks in Limbo

https://vale.rocks/micros/20260915-0135
230•edent•11h ago•96 comments

Sony's First Computer – The SMC-70 from 1982 [video]

https://www.youtube.com/watch?v=cT2-7KkPkBc
50•ksymph•2d ago•11 comments

A rough guide for going back to the Moon

https://research.ibm.com/blog/nasa-ibm-lunar-foundation-model
47•gmays•1d ago•87 comments

Show HN: Panel – A research workspace where the agent can build its own panes

https://github.com/greentfrapp/panel
39•greentfrapp•5h ago•8 comments

How much of F-Droid is LLM generated?

https://tintotint.eu/whacky-corner/f-droid_slop/
112•_ZeD_•9h ago•138 comments
Open in hackernews

An Update on Wayback Machine Access

https://blog.archive.org/2026/09/15/an-update-on-wayback-machine-access/
118•ChrisArchitect•1h ago

Comments

Onavo•1h ago
Why not just offer a paid endpoint for the crawlers? It's not like the demand is going to go away anytime soon.

It serves nobody except CloudFlare and hardware companies when one side set up blockers and the other side spend money putting VPN SDKs in consumer TVs.

I am also curious how the (Russian?) paywall bypass mirror archive.is is doing given that they are probably subject to similar amounts of traffic.

croes•58m ago
It’s one thing to archive other companies content, it’s another to sell the access to it
Onavo•56m ago
That's for the lawyers to sort out, they have a lot of flexibility as a US nonprofit. The case law isn't that clear cut for this.
simonw•54m ago
Internet Archive was almost destroyed by a copyright lawsuit from book publishers within the last few years. I expect they aren't excited to take on any additional risk of similar lawsuits right now.
xp84•39m ago
major [citation needed] on that. There are very limited exceptions to the massive power of copyright -- and they're mainly granted to libraries in the form of narrow waivers. And just the cost of fighting the most powerful copyright holders can bankrupt you -- especially if you're a relatively modestly-funded nonprofit.
celsoazevedo•32m ago
They need access to sites to archive them. It's already hard to do it as it is, imagine if they start selling access to content. They'd be shooting themselves on the foot, independently of what the law says.
faefox•43m ago
Yeah, who does the Internet Archive think it is, (insert literally any AI company here)?
bonoboTP•30m ago
Which AI company is selling access to reliable verbatim copies of websites? I don't mean "it may regurgitate a paragraph", but as a reliable service where you can repeatably get website content snapshots to a reliability level that makes such a use case viable?

Using the information for training purposes is not the same thing. Not legally the same and otherwise.

drdexebtjl•52m ago
Sites would just block the Internet Archive crawler as well.
xp84•41m ago
My guess? Because even with a paid endpoint, the type of unscrupulous yahoo that is DDOSing IA today would probably still abuse the free endpoints because they can. The revenue that might come from a paid endpoint could help to scale up, but with how slow IA usually seems, I suspect there is an upper limit to how much traffic they can serve without a LOT more revenue.

This is a major "this is why we can't have nice things" situation in my opinion. IA is one of the most valuable gems of the Internet. The only thing that even comes close to preserving our shared history. The damage being caused (both by the effective DDOSing and by the knock-on impact that abuse has in encouraging publishers to remove their content from the archive) is incredibly serious.

imglorp•21m ago
Micropayments would solve so many Internet problems. It's not too late to adopt.

Content creators could charge by page instead of depending on malware/ad/surveillance revenue. Spam is cut if there's a charge per mail. Scraping abuse goes away, along with a bunch of DDOS garbage.

The impact is a few cents per page or mail, negligible for a human. But if you're consuming a trillion pages per day, you'd reconsider.

KPGv2•20m ago
> Why not just offer a paid endpoint for the crawlers?

Because then you're definitely violating US copyright law. There are four prongs of fair use analysis, and one of them is the "nature of the use." In this case, you'd be turning into a commercial use.

simonw•59m ago
> Here’s what’s going on. The Internet Archive’s Wayback Machine has been hit by waves of high-volume automated traffic, and we’ve put protections in place to keep the service running.

I'm pretty certain this is scrapers that are trying to workaround blocks on accessing original sites by hitting the Wayback Machine copy instead. Appalling behavior.

In addition to the load it puts on this vital non-profit piece of Internet infrastructure, we've ålso already seen some sites opt out of the Wayback Machine to prevent their content from being scraped via this alternative route.

packetslave•57m ago
This is absolutely something that's happening. There are even paid scraper API's that offer "Wayback Machine fallback" as a feature.
bsimpson•37m ago
It's an open secret that you can often circumvent paywalls by searching Wayback.
gambiting•34m ago
Every single paid article linked on HN has the way back machine link as the very first comment.
ValentineC•32m ago
The links are usually to Archive.today (aka archive.ph and a bunch of other domains), not Wayback Machine (which is run by Internet Archive).
unkeen•
CqtGLRGcukpy•55m ago
> We’re getting better at telling abusive bots apart from the people who depend on the Wayback Machine every day. If you think you were blocked in error, email info@archive.org with your operating system, browser, and IP address, and we’ll look into it.
tech234a•44m ago
I wonder if they'll end up behind Anubis at some point. I'm surprised it hasn't happened already.
stickfigure•24m ago
Plenty of threads on HN about this, Anubis does not work.
BeetleB•44m ago
Wow, but I wonder if there's more to it.

I've not been able to access web.archive.org from my work computer - I always get the 429 error.

But I then pull out my phone and can access it just fine. All along I was assuming my company was blocking it. Still weird that it happens every time from my work PC and never from my home one.

dotmanish•42m ago
Could be due to some scrapers from either your work ISP block, or the larger block which lends IPs to multiple workplaces.
flexagoon•40m ago
I assume that's because the IP range of your company network overlaps with a range used by some scrapers, and if it doesn't happen on your phone even in the corporate network, then IA probably checks some extra signals like the user agent in addition to the IP
timpera•42m ago
I really appreciate the Archive team's efforts to make the Wayback Machine more responsive, and have donated a few times to support them.

Unfortunately, the restrictions have been way too strict for the last few months: from my residential IP, simply moving the mouse on the calendar for a specific URL is enough to get stuck on 429 error messages for a while; and from corporate ISPs (for example, on airport WiFi), you often can't access the WM at all. I hope they'll find a way to relax those.

UltraSane•36m ago
Why not put it in S3 with downloader pays?
Kayvanian•27m ago
As a public resource the hope is for Wayback to be free to access. I imagine putting up a paywall would be their last resort.
lousken•32m ago
AI companies should pay billions to wayback machine for access
KPGv2•22m ago
I think that'd raise serious copyright concerns, if the Wayback machine started selling other people's intellectual property.
roblh•16m ago
Shouldn’t it follow that it’s illegal for the AI labs to profit off of all of that stolen copyrighted data too?
zamadatix•7m ago
It would follow AI companies should avoid reselling the copyrighted material verbatim as their business model. Some other related discussions usually come up though.
Joel_Mckay•3m ago
That is essentially what LLM vector search results are, but the misappropriated $9Tn worth of FOSS code "AI" scraped and compacted for isomorphic plagiarism tokens is harder to prove now with watermarking skewed outputs. =3
jMyles•11m ago
It's time for copyright to end anyhow; that's what's gumming up the whole project in the first place.
xyst•26m ago
> abusive bots

Are the abusive bots in the room with us?

ignoramous•20m ago
https://archive.vn/WWENg
MattCruikshank•18m ago
There was a feature on Amazon Web Services for a while, and I wish it was still there...

Downloader pays.

I make some content and upload it. When you want to download it, you pay Amazon the egress fees. And maybe I get to charge just a bit more, to help me with the Ingress, storage, content creation, etc.

I mean, I know that there's going to be problems with rate limiting, etc. And yes, we have those problems with LLM tokens today. But this just feels like such a useful thing that it baffles me that it doesn't exist already.

vlyan•14m ago
unrelated: if a website gets hit with "This URL has been excluded from the Wayback Machine", do existing snapshots get purged or may they still be preserved somewhere?
27m ago
*APIs
luckylion•43m ago
What sites would they be targeting? Generic "just give me anything"? Whenever I check regular sites on IA, the coverage is spotty -- they'll have the homepage and a few important pages, but it quickly fizzles out.

Very understandable, you can't store all 15000 pages of any random website and update them etc etc, but that makes them pretty useless for indirect scraping because you usually don't want a tiny taste, you want everything.

toomuchtodo•41m ago
It is. They will most likely eventually need to move to a walled model for Wayback due to scraper aggressiveness (like Reddit deprecating anonymous old.reddit.com), or behind Cloudflare for aggressive bot and scraping protection. Hard to defend against abuse of a public resource when its intent is public access with as little restriction as possible.

https://en.wikipedia.org/wiki/Tragedy_of_the_commons

(no affiliation)

ronsor•33m ago
Reddit has no excuses for the anonymous old.reddit.com removal; they're simply greedy.

On the other hand, the Internet Archive is a non-profit offering a free public resource.

toomuchtodo•29m ago
Examples provided as technical examples, strong feelings are out of scope for this thread.
itintheory•17m ago
As someone who operates a large non-profit public data driven website, I have some VERY strong feelings about scrapers. We looked into various commercial solutions (Datadome, HUMAN) and based on our traffic estimates from logs we'd be looking at at least 250k/yr for bot mitigation. Anubis is offering a temporary reprieve, but after reading the recent kernel.org article [0] it's increasingly clear that this is a temporary bandaid.

The cheapest solution is to require a login and rate limit by API key. I also have strong feelings about the tragedy of the commons.

[0] https://people.kernel.org/monsieuricon/creepy-crawlies

bradly•19m ago
Just yesterday from my one of my sessions with Sol:

> Hacker News and the Rails forum are blocking the text fetcher, so I'm using the browser workflow to inspect the pages directly

TeMPOraL•9m ago
As it should.

Unfortunately I sometimes have to browbeat Claude into acting like an agent of the user is supposed to. Usually it works, though last time it refused to recognize my moral argument (on the grounds that it's not bound to my interests exclusively and needs to protect the interests of its maker too).

pantsforbirds•5m ago
We used to always "scrape" the wayback machine for any sort of news article we actually paid to consume. I was absolutely shocked by major news sites making very important edits to an article without any sort of editorial notice!

Sadly this sort of thing is probably not really possible anymore, but I can't really blame anyone for making this sort of decision. I can't imagine how much more traffic they get now vs 2021 when we were doing this.