frontpage.
newsnewestaskshowjobs

Open Source @Github

fp.

Open in hackernews

Show HN: The First Open Source Diffusion ASR Audio Model 15x Faster Than Whisper

https://arxiv.org/abs/2607.13013
2•khurdula•2h ago
We trained diffusion-gemma-asr, an open-source speech recognition model that is 15x faster than Whisper, based on DiffusionGemma and Whisper Small.

Instead of generating text one token at a time like Whisper, in diffusion-gemma-asr we transcribe by denoising an entire transcript in parallel.

To enable the model to learn, we don’t touch DiffusionGemma in training, rather we use a frozen whisper-small’s encoder for acoustic features, and train only a 42M parameter adapter, about 0.16% of the entire model.

We keep both the Whisper-small encoder and DiffusionGemma’s original weights frozen. And train a convolutional projector that compresses the Whisper’s 1,500 audio frames into 188 audio tokens, along with small LoRA adapters on DiffusionGemma’s attention layers so the model can learn to route those new tokens.

DiffusionGemma is different from most diffusion language models as it does not begin with a sequence of mask tokens. It begins with a fixed canvas filled with random vocabulary tokens. At each step it keeps the positions it is confident about and replaces the uncertain positions with fresh random tokens until the canvas settles into a transcript.

We pass the projected audio embeddings through DiffusionGemma's frozen language head and optimize them with a CTC loss. That gives the projector a learning signal without relying on attention. Once the embeddings became meaningful, attention naturally started using them, and the model actually started transcribing audio.

The CTC loss broke that loop where model produced repeated english stop words by supervising the projected audio directly through DiffusionGemma’s frozen language head, without requiring attention to use it first. And as those audio embeddings began predicting transcript tokens, attention naturally started using them and the model began transcribing. CTC is only a training scaffold and it is removed entirely during inference.

Upon completion of training on Librispeech, english WER dropped significantly on clean test set. We then warm started from that checkpoint and finetuned on FLEURS across English, German, French, Spanish, Hindi, and Mandarin, later mixing in VoxPopuli to improve ASR on accented speech.

On Librispeech test set for english, our model sees 6.6% WER, compared to 8.3% for Whisfusion and roughly 7% reported by TransFusion, while using a smaller encoder, making us SoTA for Diffusion-ASR .

We're still behind standard Whisper, but that's not surprising. As Whisper was trained on millions of hours of speech. Our Gemma variant has seen roughly 219 hours. The architecture seems to scale; we simply haven't trained it on enough audio yet.

The biggest takeaway from this project was that a frozen diffusion LLM can learn a completely new modality by training only a small adapter as long as you give that adapter a way to learn before the rest of the model decides to ignore it.

The adapter weights, code, inference script, and model card are all open source. We’re Interfaze, a research lab building deterministic AI models. You can try it here: https://huggingface.co/spaces/interfaze-ai/diffusion-gemma-a...

Building Service Topology at Scale: Architecture, Challenges, Lessons Learned

https://netflixtechblog.com/building-service-topology-at-scale-architecture-challenges-and-lesson...
1•jjtang1•1m ago•0 comments

Hugging Face discloses breach linked to autonomous AI agent

https://www.bleepingcomputer.com/news/security/hugging-face-breach-autonomous-ai-agent-system-int...
1•Brajeshwar•2m ago•0 comments

Show HN: Imagin Raw – A 9MB Open-Source Alternative to Adobe Bridge for Mac

https://github.com/cristibaluta/Imagin-Raw
2•cristi_baluta•3m ago•0 comments

Augustus raises $180M to build a clearing bank for the AI and stablecoin era

https://www.coindesk.com/business/2026/07/21/augustus-raises-usd180-million-to-build-a-clearing-b...
1•AnhTho_FR•4m ago•0 comments

Google vs. SerpApi: The Court Granted Our Motion to Dismiss

https://serpapi.com/blog/google-v-serpapi-the-court-granted-our-motion-to-dismiss/
2•SerpApi•5m ago•0 comments

Firefox Containers Preview

https://blog.mozilla.org/en/firefox/firefox-containers-preview/
1•twapi•5m ago•0 comments

Show HN: Quartz – A (much) better Voice Memos app (iPhone)

https://www.madebywindmill.com/quartz/
2•jscalo•8m ago•0 comments

Guard-AI – A security linter for AI-generated code

https://github.com/Karthikvk1899/guard-ai
1•karthikvk1899•9m ago•0 comments

Apple to launch 'upgrade' device leasing program with Klarna to spur sales

https://www.bloomberg.com/news/articles/2026-07-21/apple-to-launch-upgrade-device-leasing-program...
3•anigbrowl•9m ago•0 comments

Show HN: I built a command palette for the terminal – 6.2MB, pure Go, no fzf

https://github.com/matheuzgomes/decoreba
6•MarinhoD•10m ago•0 comments

World Labs acquires SceniX

https://twitter.com/drfeifei/status/2079597384510898377
2•dmarcos•11m ago•0 comments

Apache Spark 4.2: Making Your Data AI‑Developer Friendly

https://techstrong.it/featured/apache-spark-4-2-making-your-data-ai-developer-friendly/
4•CrankyBear•11m ago•0 comments

Build a Basic AI Agent from Scratch: Security II

https://www.ruxu.dev/articles/ai/build-an-ai-agent-security-2/
1•ruxudev•11m ago•0 comments

Soulless

https://sebas.fika.bar/soulless-01KVT6WJVSPF5ZTHWNZPZNEGJ9
3•smtx•12m ago•0 comments

SkyPilot Is Out of Stealth

https://skypilot.ai/blog/skypilot-the-company
9•shenli3514•13m ago•0 comments

Delaware judge dismisses UnitedHealth Group lawsuit against the Guardian

https://www.theguardian.com/media/2026/jul/21/unitedhealth-lawsuit-the-guardian-dismissed
3•speckx•13m ago•0 comments

Show HN: Word in Web – Near MS Word Parity Docx Editor in Web

https://word-in-web.com/
3•theRealestAEP•13m ago•0 comments

Hugging Face says it resorted to a Chinese AI model

https://fortune.com/2026/07/20/hugging-face-turns-to-chinese-open-source-ai-to-fend-off-autonomou...
3•danielmorozoff•13m ago•0 comments

Buzz community for folks to test it out

https://startups.communities.buzz.xyz/invite/eyJjIjoiYjllYjBmMWUtNWVlMi00YmZmLWFiNTQtNDhiYjBmZjA5...
2•pdcd•13m ago•0 comments

Time-Frequency Consistency Learning for Robust Speech Deepfake Detection

https://arxiv.org/abs/2607.17761
2•zhinit•14m ago•0 comments

From Hierarchy to Intelligence

https://block.xyz/inside/from-hierarchy-to-intelligence
1•andsoitis•14m ago•0 comments

CipherX – Microdot Tattoos

https://cipherx.tech/
1•MrAlex94•14m ago•0 comments

WhichSpace: macOS menu bar utility for viewing and switching Spaces

https://github.com/gechr/WhichSpace
2•gechr•15m ago•0 comments

Laguna S 2.1

https://poolside.ai/blog/introducing-laguna-s-2-1
5•rexledesma•16m ago•0 comments

San Jose residents file federal lawsuit challenging citys ALPR use [pdf]

https://ij.org/wp-content/uploads/2026/04/Doc.-1-Complaint-for-Declaratory-and-Injunctive-Relief.pdf
4•smalltorch•16m ago•0 comments

Paramount Forced to Delay Warner Bros Merger After California Antitrust Lawsuit

https://www.techdirt.com/2026/07/21/paramount-forced-to-delay-warner-bros-merger-after-california...
3•cdrnsf•18m ago•1 comments

Jack Dorsey launches Buzz to combine team chat, AI agents and Git hosting

https://runtimewire.com/article/jack-dorsey-block-buzz-team-chat-ai-agents-git
2•ryanmerket•19m ago•0 comments

Where people and agents work together

https://buzz.xyz
3•TimCTRL•19m ago•0 comments

Show HN: Ave, a behavioral classification standard for agentic AI

https://github.com/aveproject/ave
2•bawbel•19m ago•0 comments

Buzz

https://engineering.block.xyz/blog/buzz
3•handfuloflight•19m ago•0 comments