frontpage.
newsnewestaskshowjobs

Open Source @Github

fp.

Open in hackernews

AI is now capable of developing its own inference hardware

https://github.com/FeSens/openTPU
79•fsbonetto•1h ago

Comments

fsbonetto•1h ago
After using AI to develop risc-v CPU cores, the same technique was used for developing openTPU. An open source AI inference engine. It's able to run most of the modern models like Qwen 3.5, Gemma 4, and many others. The TPU started able to produce only a few tokens per second and trough a recursive self improvement loop got to 80+ tok/sec on the smallers models.
vatsachak•47m ago
I feel like there is a lot to be gained from an experienced user pointing an LLM in a tasteful direction.
skybrian•37m ago
This seems to be running on an FPGA board that costs ~$300? Anyone know more about the hardware?
fsbonetto•27m ago
Its a datacenter decommissioned board, really popular among hobbyists.

For a TPU focused on inference the name of the game is memory bandwidth. How much of the available bandwidth you can extract for as little logic/area/power as you can.

pcarolan•34m ago
Really dumb question from a software guy. Why aren't the labs burning their frontier models into chips already? Seems like the performance gains and cost per request would be worth it. That said, I understand neither the economics nor the physical challenges to doing this.
hehimself•32m ago
They do. It takes time to deploy those chips though. Check out OpenAI and Broadcom deal.
traverseda•31m ago
I'd presume because it take too long to go from design to tapeout to production. Their whole business is predicated on having better models.

Also can't keep them closed source if you do that.

skeskinen•31m ago
Lead times are so long that there is a lot of risk the chips would be obsolete by the time they come out.

Also, it's hard to get fab capacity for any project. Let alone something so experimental.

jcims•28m ago
Addressing these issues seems to a major driver behind the design of terrafab.
zitterbewegung•29m ago
Etched is a startup doing exactly this.

https://www.etched.com/progress/frontier-inference-clusters

rfgplk•29m ago
Yep, 99.9% of people are completely oblivious to what LLMs can do. Just wait until the next gen of CPUs/GPUs designed by LLMs start coming out (fyi chip development tools have advanced centuries in the last few months) and you'll start seeing exponential gains in hardware.
jetemple•19m ago
Which tools have made that leap? Faster design iteration makes sense, but what points to exponential hardware gains rather than shorter development cycles?
xg15•22m ago
"Recursive self-improvement will kill us all!"

Also: Here is our recursive self-improvement hard at work...

lelanthran•13m ago
> "Recursive self-improvement will kill us all!"

> Also: Here is our recursive self-improvement hard at work...

Soon we will see

token-providers: "The torment nexus is a cautionary tale"

Also token-providers: "Finally, we have created the torment nexus that we first told you about!"

nialse•10m ago
All will end up on same plateau eventually. RSI is just a phase on the way there.
dumberquestions•9m ago
Technology has always contributed to improving next iterations of itself, it's only a concern when it's fully autonomous.
athrowaway3z•20m ago
I haven't really dug into the results yet, but my guess is that a SOTA model has been able to produce an accelerator that runs a model since around December.

The obvious next step is to get enough memory throughput to run that SOTA model itself so that it develop its own hardware.

But perhaps the more interesting question is this: Can an AI be given a big FPGA and design a model architecture that takes advantage of the fabric being reconfigurable.

felixgallo•9m ago
I suspect an AI could design a purpose-built FPGA-like replacement that would be, for its purpose, significantly more effective than the current general-purpose FPGAs.
fsbonetto•8m ago
It could have a small improvement on power consumption, but the current design can already achieve 90% of the maximum theoretical speed of this hardware without giving up programability/flexibility
bitwize•16m ago
Colossus is building Colossus II.
srameshc•10m ago
This post brings me to question "What does it mean to be a software developer in future" ?
AnimalMuppet•7m ago
Can anyone comment on the performance of this hardware? How does it compare to state of the art, human-designed hardware? Is this actually an improvement? (To get to recursive self-improvement, you first have to improve at all.)
ohazi•28m ago
They [1] are [2].

[1] https://taalas.com/

[2] https://chatjimmy.ai/

birdatlaw•24m ago
From what I've read, not only are some labs doing it (other commenters already mentioned).

But it's complicated for other reasons, one being that the number of parameters for frontier models (especially with MoE models) are so high, and not always utilized (once again, thanks to MoE) that it would actually be incredibly cost prohibitive, if not impossible, to attempt to make giga-chips that would allow running it.

I definitely do believe that we will see more and more specialized chips over time, but putting the entire model on a chip is still a ways away.

I believe Taalas has a heavily handicapped llama 8-billion parameter model. And it still pulls >200W to run.

I can't imagine how anthropic or open ai would be able to burn a multi-trillion parameter model on a chip, we just aren't there yet.

zdragnar•23m ago
Model SOTA moves faster than chips can be designed or produced. You'd need to commit to a particular model for years to get payoff while still burning buckets of money producing new SOTA models to keep up with the competition.

It's why everyone and their dog runs these things on GPUs. When a new model supercedes the previous one, so long as you've got the memory for it your chips aren't obsolete.

I'm looking forward to someone picking a model to be "good enough" (say, qwen 4.0 or something) and selling them as peripheral hardware

fhdkweig•12m ago
I know FPGAs are more expensive than GPUs, but are they fast enough to justify the extra cost?
fsbonetto•11m ago
They are more like a way to proving the architecture of the accelerator before committing 100's of millions into a custom ASIC with TSMC
LoganDark•6m ago
1. No

2. They don't have enough capacity either

The current largest FPGA, the AMD Versal Premium VP1902 has 18.5 million logic cells. That's not even enough for the smallest whisper.cpp model (75M)

monocasa•5m ago
They're not magical go faster juice. I don't know of a microarch where they're faster than modern GPUs at ML training or inference.
zdragnar•4m ago
It isn't just a matter of speed, it's also a matter of model quality. If they take 6 months to burn Fable to chips, and it takes 2 years to break even between design, custom fab, energy savings, etc, are those chips even worth running when the new models that are running on GPUs at that point are producing 10x better quality results?

Sure, your 2.5 year old models are running faster, but you can't drop prices on them without pushing the break even point further out.

If the cost difference isn't incredibly significant, will people even want to pay for the 2.5 year old model, or will they get more value for their money paying more to get better results from the newer model?

There's a lot of open ended questions that I don't have the insiders knowledge for to suggest whether or not such a capital outlay would be a worthy investment.

My guess is that state of the art stuff will stay on GPUs and models burned into chips will be for "good enough" applications that people are still teasing out. Probably highly specialized models in automated sensor units and such.

fsbonetto•22m ago
The bottleneck, for inference at least, is memory bandwidth. And that you can't make any faster by making it specific to your model.

So companies try to maximize the memory bandwidth they can get, balancing tradeoffs of power/area/programability of their chip. Right now they feel like the economy on power/area is not worth the decrease in programability/flexibility.

fnordpiglet•11m ago
Presumably though the kernel has a pretty specific set of operations done against the weights in memory. Burning the weights into the memory with local memory cores capable of the kernel operations would be a lot more efficient than round tripping busses.

The primary constraint isn’t likely what’s possible to do, but that the kernel and weights are too variable right now and the patterns too poorly established to bake into hardware accelerators yet. Margin pressure is also not there yet.

I suspect as the marginal utility of the frontier improvement settles into diminishing returns (I suspect we are there already tbh) baking hardware models with ROM, working set, and kernel cores collocated will be the frontier space as the goal will become reducing capital spend to utility levels rather than research levels.

Once someone has a model that is sufficient for almost any practical use, making marginal inference cost effectively zero will be the competition frontier. I do shed a tear for all those lonely data centers as compute densities will almost certainly make most of them a terrible investment.

But such is the cycle

schleck8•22m ago
Because the iteration speed on models is so fast that by the time they have an ASIC ready for one model version, they are already significantly ahead in capability. Think of how big the jump between Opus 4.8 and 5.5 has been. They were released four months apart.
pmarreck•17m ago
Yeah, and what about FPGA? Which was the same interim state when Bitcoin went GPU -> FPGA -> custom chip fab?
fsbonetto•15m ago
GPUs are faster, but you can't make your own arch on GPUs. FPGAs offer you that possibility. Said that... There are a few beasty FPGAs used in crypto mining coming my way... I expect that OpenTPU will be able to run frontier models with those.
dmitrygr•14m ago
In addition to some of the other replies you got, here is one more:

Much of a model are weights, and high-density ROMs are very very very hard.

__MatrixMan__•12m ago
Would you pay to crystalize one of today's models in silicon so you can use it in 2028, or would you wait for another 6 months to see how models improve before pulling the trigger on that kind of commitment?
jolt42•5m ago
Even dumber question: What is new or novel about this openTPU?

Mistral Large 4

https://mistral.ai/news/mistral-large-4/\
943•Philpax•4h ago•634 comments

AI is now capable of developing its own inference hardware

https://github.com/FeSens/openTPU
83•fsbonetto•1h ago•44 comments

Nobel Prize in Physics goes to Francis Halzen

https://www.nobelprize.org/prizes/physics/2026/
359•solarist•7h ago•109 comments

Release of Polars 2.0

https://pola.rs/posts/release-polars-2/
278•simicd•5h ago•46 comments

The Early History of Smalltalk (1993)

https://worrydream.com/EarlyHistoryOfSmalltalk/
50•_reza•2h ago•11 comments

Tapo (Rust/Python library) now speaks TP-Link's TPAP protocol

https://mihai.dinculescu.dev/posts/tapo-speaks-tpap/
83•faithraven•3h ago•26 comments

JetBrains reported a net financial loss first time in its tracked history

https://www.helgilibrary.com/companies/jetbrains
466•thw_9a83c•5h ago•448 comments

Benchmark in Milliseconds

https://matklad.github.io/2026/10/05/benchmark-milliseconds.html
65•surprisetalk•1d ago•14 comments

Meta's Muse Is an Adorable Privacy and Security Dumpster Fire

https://www.techdirt.com/2026/10/06/metas-muse-is-an-adorable-privacy-and-security-dumpster-fire/
265•beardyw•4h ago•167 comments

Gleam doesn't compile to Erlang source anymore

https://gleam.run/news/gleam-doesnt-compile-to-erlang-source-anymore/
209•ingve•9h ago•93 comments

Utah to let AI examine patients and prescribe medication without human oversight

https://www.techspot.com/news/114111-utah-become-first-state-ai-examine-patients-prescribe.html
23•healsdata•26m ago•14 comments

Mathematics of Geothermal Energy

https://www.ebsco.com/research-starters/power-and-energy/mathematics-geothermal-energy/
40•srameshc•4h ago•16 comments

Subquadratic 3SUM and Subcubic APSP

https://arxiv.org/abs/2610.06783
28•mauriziocalo•4h ago•9 comments

Show HN: I turned my iPhone and a $20 smart plug into an f-stop timer

https://peterszentkiralyi.eu/darkplug/
12•pentakkusu•3h ago•1 comments

Nature's capacity to 'bounce back' when species are lost is vastly overestimated

https://phys.org/news/2026-10-nature-capacity-species-lost-vastly.html
209•pseudolus•6h ago•96 comments

Beam: Reflection's 501B open-weight model

https://reflection.ai/blog/introducing-beam
513•Philpax•22h ago•163 comments

Show HN: Parseable, an open observability datalake, handles 100M time-series/min

https://www.parseable.com
40•yashdotrv•3h ago•7 comments

Find the flattest route between any two points in SF

https://flattensf.com/
272•ishan0102•19h ago•94 comments

Dust: Pretraining Transformers Without Backpropagation

https://qlabs.sh/research/dust
251•E-Reverance•20h ago•66 comments

World's First enhanced geothermal power plant completed in just 23 months

https://techcrunch.com/2026/10/01/worlds-first-enhanced-geothermal-power-plant-completed-in-just-...
111•hochmartinez•5h ago•44 comments

Friendship ended with Deno, now Node is my best friend

https://dbushell.com/2026/10/03/deno-to-node/
274•ibobev•18h ago•185 comments

Direct retinal projection display for smart glasses using a meta-optic mirror

https://www.tdk.com/en/news_center/press/20261002_01.html
90•bookofjoe•3d ago•40 comments

Opus 5.5 agents discover two room-temperature magnetic semiconductor candidates

https://www.vals.ai/blogs/room-temperature-magnetic-semiconductors
444•outlier99•20h ago•304 comments

Testing 12 different Zigbee temperature/humidity sensors

https://smarthomescene.com/reviews/best-selling-zigbee-temperature-sensors-tested/
186•walrus01•2d ago•105 comments

Competitive Programmer's Handbook (2018) [pdf]

https://cses.fi/book/book.pdf
268•vinhnx•3d ago•65 comments

Two ARM64-specific compiler optimization bugs, in GCC 15/16 and Rust, hit curl

https://mastodon.social/@bagder/117392573268225646
36•torutofu•4h ago•5 comments

The complement of true is true, except when it's false

https://dryperspective.github.io/posts/complement-of-true/
50•aw1621107•2d ago•18 comments

Example.com just launched the biggest redesign in decades

https://www.debugbear.com/blog/example-dot-com-redesign-history
302•jgx0•18h ago•207 comments

The lamps in my house

https://arslan.io/2026/10/05/the-lamps-in-my-house/
293•farslan•1d ago•112 comments

Resurrecting iChat Audio and Video Conferencing

https://blog.pipetogrep.org/2026/09/11/resurrecting-ichat-audio-and-video-conferencing/
99•thepipetogrep•13h ago•33 comments