frontpage.
newsnewestaskshowjobs

Open Source @Github

fp.

AMD acquires Taalas to boost inference performance by etching models in silicon

https://www.theregister.com/systems/2026/08/06/amd-acquires-ai-chip-startup-taalas-to-boost-infer...
111•itvision•1h ago•58 comments

NSF Inouye Solar Telescope Enables Major Discovery of a Hidden Solar Process

https://nso.edu/press-release/nsf-inouye-solar-telescope-enables-major-discovery-of-a-hidden-sola...
76•neversaydie•1d ago•9 comments

Qwen3.8 Max now ranked as the best overall model by agentic index

https://artificialanalysis.ai/?intelligence=agentic-index
320•apitman•3h ago•195 comments

Mario Meets Pareto

https://www.mayerowitz.io/blog/mario-meets-pareto
792•theanonymousone•10h ago•140 comments

Herdr is joining Y Combinator. The runtime stays open

https://herdr.dev/blog/herdr-is-joining-y-combinator/
76•collinmanderson•2h ago•47 comments

Launch HN: ProvenMetal (YC S26) delivers circuit boards in days instead of weeks

https://provenmetal.com
159•willcarkner•5h ago•109 comments

Quake – 30th Anniversary Update

https://slayersclub.bethesda.net/en-US/news/quake-30th-anniversary-update
125•dsubburam•1h ago•60 comments

Almost no skill required to cook a steak

https://blog.sydorets.com/en/posts/almost-no-skill-required-to-cook-a-steak/
221•yusyd•6h ago•257 comments

Can you reverse engineer an ASIC?

https://blog.janestreet.com/can-you-reverse-engineer-an-asic/
35•bschne•2h ago•12 comments

Humans missed 1 in 3 threats approving AI agent commands across 40k game runs

https://scalex.dev/blog/ai-agent-permissions-stats/
222•Wirbelwind•9h ago•178 comments

Learn how chips are made with this Rollercoaster Tycoon-inspired animation

https://laurentiugabriel.github.io/ChipTycoon/
67•laurentiurad•6d ago•11 comments

Civilians under siege by Mexican cartel fight back with AK-47s, grenades

https://www.pbs.org/newshour/world/civilians-that-were-under-siege-by-a-mexican-cartel-fight-back...
41•starkparker•58m ago•10 comments

Crime Pays but Botany Doesn't

https://www.crimepaysbutbotanydoesnt.com/reading-list
618•DarkContinent•17h ago•191 comments

Improving GPT‑5.6 Sol in ChatGPT, expanding GPT‑5.6 Luna access for free users

https://openai.com/index/improving-gpt-5-6-sol-in-chatgpt/
91•tedsanders•4h ago•70 comments

Taste Is All That's Left

https://notashelf.dev/posts/taste-is-all-thats-left
98•tsak•4h ago•79 comments

The simple elegance of the integrated timing belt loopback fastener

https://danielmangum.com/posts/integrated-timing-belt-loopback-fastener/
84•hasheddan•4d ago•17 comments

How to Make a Nintendo 64 Game in 2026

https://phoboslab.org/log/2026/08/xibalba64-making-of
434•atan2•2d ago•212 comments

GitHub Actions and Pages are experiencing degraded availability

https://www.githubstatus.com/incidents/qcvjkzcs7j74
244•Footkerchief•5h ago•201 comments

Building Progressively Enhanced Forms Using htmx

https://www.rafa.ee/articles/progressive-enhanced-forms-htmx/
39•mpweiher•1w ago•6 comments

Pareto Front

https://en.wikipedia.org/wiki/Pareto_front
210•binyu•1w ago•89 comments

Show HN: The Channels SDK – Bring Any Agent to Any Channel (Slack, MS Teams)

https://github.com/CopilotKit/channels-sdk
75•davidmckayv•5h ago•19 comments

The Triangle Game, from Zero

https://muchmirul.github.io/conjectures/multicolor-ramsey/
8•jdkee•4d ago•1 comments

State-oriented consistency: Why we stopped looking for one right answer

https://keel-iot.eu/blog/state-oriented-consistency.html
4•fcravio•1h ago•1 comments

Tiny black holes may be exploding stars across the Milky Way

https://www.sciencedaily.com/releases/2026/07/260729051515.htm
60•jandrewrogers•1d ago•73 comments

Zapscape (CVE-2026-64561): Guest-to-Host Escape in KVM/x86

https://github.com/V4bel/Zapscape
55•john_strinlai•5h ago•9 comments

Astronomers capture highest-resolution image ever of the Sun's surface

https://physicsworld.com/a/astronomers-capture-highest-resolution-image-ever-of-the-suns-surface/
21•layer8•1h ago•4 comments

Four simple rules behind Japan's most liveable cities

https://www.bbc.com/travel/article/20260805-four-simple-rules-behind-japans-most-liveable-cities
61•tchalla•8h ago•76 comments

Towards Bottom-Up Enumeration in miniKanren via Pruning and Memoization

https://arxiv.org/abs/2607.25373
4•Jimmc414•3d ago•0 comments

Federal Communications Commission scraps limit on broadcast TV ownership

https://www.nbcnews.com/business/media/federal-communications-commission-scraps-limit-broadcast-t...
88•pseudolus•3h ago•73 comments

Unearthing my 1996 windowed OS in machine code for Am29000 homebrew computer

https://nanochess.org/the_am29000_computer.html
151•nanochess•5d ago•32 comments
Open in hackernews

AMD acquires Taalas to boost inference performance by etching models in silicon

https://www.theregister.com/systems/2026/08/06/amd-acquires-ai-chip-startup-taalas-to-boost-inference-performance-by-etching-models-into-silicon/5284344
99•itvision•1h ago
https://ir.amd.com/news-events/press-releases/detail/1296/am...

Comments

proxysna•1h ago
Really hoped to see their hw out in the wild one day
rvz•1h ago
Didn't even give them a chance to launch the hardware.
MarkWayneNewton•1h ago
While this design is self-limiting I think its a good approach. It doesn't take an entirely new architecture or infinite memory to produce significant performance improvement.
badatnames•1h ago
Well so much for that dream.

Guess we can look forward to picking these up ex-enterprise on ebay for under $5k a pop in a decade or two

whythismatters•59m ago
The demo: https://chatjimmy.ai/
nsxwolf•34m ago
It doesn’t believe it’s running on that chip, it’s arguing with me
shaewest•26m ago
It's running a very small, non-reasoning model at the moment. But more generally, almost all LLMs argue on the hardware/model they are/are on.
dumberquestions•22m ago
Which model? Or how many active parameters?
_whiteCaps_•19m ago
Llama 3.1 8B model
metadat•9m ago
What would tokens/sec performance look like for a reasoning model? An order of magnitude slower?
itvision•28m ago
OMFG this thing is fast.
A_D_E_P_T•54m ago
This is probably a win-win. The team gets paid, and we get greater assurance that their best ideas and architectures -- which are truly impressive -- are going to see the light of day in actual products.
badatnames•51m ago
They were too small for this to be a meaningfully sized purchase for AMD, there's real risk they get sucked into a team that ultimately delivers sqat, not to mention the chances of anything being delivered in an even remotely consumer-priced bracket are definitely out the window
ycui7•51m ago
so qwen3.x-27b on hardware? or better deepseek-v4-flash on hardware .
ilaksh•46m ago
I wrote them an email asking for PrismML Bonsai 27b Ternary which is like 6b or something crazy small and would be a lot easier for them to do initially.
syntaxing•46m ago
Honestly, this is starting to make more and more sense. SOTA models are starting to converge to certain architecture and capabilities. I wouldn’t be surprised we end up with a base model ASIC + “fine tune” card where it’s a physical LoRA style adapter.
smokel•40m ago
The technical aspects of SOTA models are not publicly documented. How do you know if something is converging?
cyanydeez•37m ago
if they were still exponentially increasing, they wouldn't be preparing for an IPO. IPO is where companies go to die and founders escape.
_aavaa_•37m ago
If we had deepseek v4 flash 0731 etched on a chip it would be more than capable enough and fast enough for so many people's needs, even hardcore engineer.
nurumaik•26m ago
Will be capable and fast enough for 2-3 weeks until new sota drops
amazingamazing•24m ago
If it is capable today why would a new model change this?
FridgeSeal
bhouston•43m ago
Toronto Canada startup btw.
cmrdporcupine•25m ago
Seems to be somehow some kind of offshoot from or connected to Tenstorrent, which is just down the road. Founder looks like he was/is maybe at Tenstorrent and previously associated with Keller?

Always fantasize about applying at Tenstorrent, but wrong side of Toronto. 2 hour commute.

kridsdale1•21m ago
Works well, I remember driving by the ATI building as a kid.
mikeayles•29m ago
AMD could have saved their money and used their own hardware! I've got a language model doing 60k tok/s on AMD hardware already, a Xilinx Kria K26 SOM, with the weights baked into URAM/BRAM with zero DRAM in the token loop. Same thesis as Taalas: single-stream decode is bandwidth bound, so stop fetching weights from far away.

Caveats stacked high, obviously. It's 3.16M parameters (tinystories, and I also have a kevin-speak lemmatised version), the tokens are characters, and the 60k record is 16 streams that each remember exactly one token of context, so it's blisteringly fast at saying nothing. The honest build with full context and KV caching still does ~19k tok/s on one stream though.

I keep messing with the blogpost with the live demo, but I'm planning on flipping it to live in the next day or two

tandr•17m ago
Well, technically it is their hardware now...
bob1029•26m ago
I feel like NAND process tech could become useful at solving some of these problems. A GPU where you can update the weights a few thousand times may be sufficient.
kridsdale1•24m ago
FPGA model storage?
addaon•13m ago
NAND hasn't been scaling great lately. It seems like PCM or MRAM would both be better fits.
mdp2021•5m ago
The basis of Taalas is "compute in memory" electronics - past Von Neumann's separation of processor and memory.

You need to be able to add|mul where the data (the weights) are stored.

fellowniusmonk•16m ago
Token quantity will have a quality all its own.
LarsDu88•10m ago
I'm surprised neither OpenAI nor Anthropic made this move first. The Chinese open weight models are pulling ahead and commoditizing their value proposition.

Baking models onto silicon would've been the next logical move to get a moat.

Google is already doing this and has an experimental project on top of already having TPUs and cramming their quantized flash onto individual TPUs for inference.

walrus01•9m ago
Imagine the size of chip needed to 'etch' something like Qwen 3.6 27B in size.
wxw•24m ago
I freakin' love this demo. It feels magical.
walrus01•6m ago
I know it's a relatively tiny model, but damn, is that thing fast.

It also mostly passes the "schlong" test

https://pastes.io/YcxSi8Fp

hendurhance•6m ago
I understand the appeal due to the speed
•
13m ago
Because new stuff instantly makes anything prior bad and incapable and garbage of course! Did you forget the hype-machine speaking notes??? /s
catchnear4321•12m ago
if capability is a commodity then the differentiator becomes taste.
syntaxing•32m ago
SOTA American models are not. SOTA Chinese models are. From a physics aspect, closed source models cannot be too far from open source ones in terms of size. There’s only so much you can squeeze out a B100 style cluster even with fancy Dflash style diffusion model for the speculative model.
cyanydeez•36m ago
I don't think there'll be a fine tune card; you'll have the base model vintage whatever year, and then your GPU will do whatever LoRA layers you want it to do; the LoRA will wrangle older dated models into the current of whatever your looking at.

But yeah, for things like programming, if it can do linux and python and some go and sql and javascript, larger domains can be threaded with LORA

VladVladikoff•32m ago
Wouldn't this mean someone with sufficient hardware could lift the SOTA model weights off the chip? Or are you saying that these chips would only be used internally by these companies and not sold to the public?
syntaxing•30m ago
I don’t get why this is an issue? You can run Claude/OpenAI SOTA models through Amazon bedrock. These weights have to live somewhere to run on Bedrock.
snek_case•24m ago
The weights are very unlikely to be on the chip itself. That wouldn't work for SOTA models that are terabyte scale, even quantized. This is probably an accelerator for specific kernels in the model, but the weights are likely loaded from memory. The chip may have SRAM to store some of the weights temporarily during inference.
amazingamazing•24m ago
One idea would be to use an open model.
dumberquestions•17m ago
I wouldn't expect companies not sharing their weights today to be any more likely to share them if they're on hardware, this doesn't sufficiently hide weights from a local user.
encyclopedism•29m ago
Imagine a multi-modal model with 1000's of tokens per second. Realtime inference for a host of applications. This is a BIG deal and will change the landscape in unfathomable ways.

The https://chatjimmy.ai demo was impressive.

Once models settle down this makes sense. Imagine a cartridge with a physical model on it. You purchase a cartridge and stick it in your computer/phone/server. Want to upgrade? By a new 'cartridge'.

This should bring inference cost down dramatically, I wonder how OpenAI/Anthropic feel about that.

kevin_thibedeau•14m ago
Then we can have machine psychologists pull cards when they run anok.