frontpage.
newsnewestaskshowjobs

Open Source @Github

fp.

Whistle: Speech to Text in 16.9 MB

https://cactuscompute.com/blog/whistle
87•gmays•1h ago•24 comments

Trump administration is suspending Microsoft from a green card program

https://apnews.com/article/h1b-visa-program-vance-microsoft-e7b3a407f822702b269ee277d21343ea
440•alephnerd•2h ago•687 comments

Beauty in DVD Menus

https://vale.rocks/posts/dvd-menus
170•speckx•4h ago•109 comments

“Math 2.0” will need to value mathematical progress more holistically

https://mathstodon.xyz/@tao/117395269325940185
552•ent101•12h ago•562 comments

Archaeologists Are Reconstructing the 'Invisible' Technologies of the Stone Age

https://www.smithsonianmag.com/science-nature/archaeologists-are-reconstructing-the-invisible-tec...
50•Hooke•20h ago•15 comments

Tell HN: I've been paying for a rural Tanzanian's education for 10 years

511•lukehandcool•3h ago•129 comments

4-hour battery storage is cheaper to install than gas turbines all across globe

https://www.solarpowerworldonline.com/2026/10/4-hour-battery-storage-is-cheaper-to-install-than-g...
117•01-_-•1h ago•45 comments

The Slow Formation of Durable Software

https://newsletter.dancohen.org/archive/the-slow-formation-of-durable-software/
201•benbreen•2d ago•68 comments

Step 5 Preview, a 1M-context MoE from StepFun, shows up on OpenRouter

https://openrouter.ai/stepfun/step-5-preview
21•AnneWodell•1h ago•8 comments

Sub-1-Bit LLM Compression via Latent Factorization

https://github.com/SamsungLabs/LittleBit
59•brainless•4h ago•8 comments

OpenAI annualised revenues $20B less than previously signalled

https://www.ft.com/content/b66a9858-f8fb-46cb-b506-44bfe26fca2a
69•mfiguiere•1h ago•12 comments

I gave Opus 5.5 one prompt and six hours to visualize Invisible Cities

https://quesma.com/blog/invisible-cities-one-shot/
260•stared•6h ago•139 comments

Claude Haiku 5.5

https://www.anthropic.com/claude-haiku-5-5
1012•sfkgtbor•1d ago•474 comments

Telnet BBS Guide

https://www.telnetbbsguide.com/
84•kmstout•5h ago•28 comments

15-year search for a band that charted once and vanished

https://shahidhussain.com/writing/search-for-salvage/
312•shahidhussain•3d ago•114 comments

The Deeply Impersonal Personalized Recruiter Mail

https://blog.pentlander.com/the-deeply-impersonal-personalized-recruiter-mail/
11•speckx•45m ago•1 comments

OpenAI withdraws three mathematical results

https://twitter.com/danintheory/status/2108065033070789090
157•sashank_1509•10h ago•462 comments

Time Travel in Braid (2015)

https://qntm.org/braid
80•Ariarule•3d ago•34 comments

Living off-grid: Hundred Rabbits

https://100r.ca/site/home.html
342•Muhammad523•2d ago•149 comments

Float and integer arithmetic follow two different paradigms

https://blog.pkh.me/p/49-float-and-integer-arithmetic-follow-two-different-paradigms.html
48•ibobev•3d ago•10 comments

2027 Web Platform Feature Ranking

https://interop-rank.fxdx.dev/
38•jaffathecake•4h ago•9 comments

'Jonathan' is the oldest land animal on Earth

https://www.404media.co/oldest-living-land-animal-jonathan-the-tortoise/
249•gumby•22h ago•139 comments

Cleo (Mathematician)

https://en.wikipedia.org/wiki/Cleo_(mathematician)
240•djoldman•1d ago•49 comments

The Mathocalypse

https://scottaaronson.blog/?p=10169
350•6bitquant•22h ago•366 comments

Teams in Vienna and Beijing have built the first thorium nuclear clocks

https://www.nytimes.com/2026/10/07/science/first-nuclear-clocks-thorium-229.html
165•gumby•1d ago•47 comments

Show HN: Bigwords.page – Turn any screen into a sign. The URL is the app

https://bigwords.page/
650•SpeakingOfBrad•1d ago•151 comments

VECOS – A windows-like operating system for the Vectrex for the UVMC2 [video]

https://www.youtube.com/watch?v=9ranfp_vz30
49•CharlesW•1d ago•11 comments

Sharing AI progress in mathematics

https://openai.com/index/sharing-ai-progress-in-mathematics/
1312•OfficialTurkey•1d ago•1485 comments

Anne Carson wins Nobel Prize in literature 2026

https://www.theguardian.com/books/2026/oct/08/wins-the-nobel-prize-in-literature-2026
128•sonabinu•6h ago•27 comments

Docker Agent

https://github.com/docker/docker-agent
287•saikatsg•1d ago•130 comments
Open in hackernews

Whistle: Speech to Text in 16.9 MB

https://cactuscompute.com/blog/whistle
85•gmays•1h ago

Comments

andy_ppp•36m ago
Wow certainly in English this is incredibly accurate I tried to break it and it understood me perfectly!

I know it's slightly off topic but surely it must be easy by now to train a spell checker that doesn't annoy the crap out of everyone using it (looking at you here Apple)!

tecleandor•36m ago
Spanish is not good (seems to write non existing words and/or with terrible typos...) but English seem to work good even with my (Spanish) accent...
kaoD•34m ago
Spanish from where? Here (Castilian Spanish) it seemed to work fine.
chilicuil•8m ago
Mexican and venezuelan aren't detected correctly
tecleandor•6m ago
Madrid. But it will only work properly if I'm clearly dictating with a very regular rhythm (ViaVoice dictation, if anyone remembers...). If I use a more natural/conversational rhythm (no slang, no abbreviations...) it easily confuses words.
saturn8601•31m ago
Initial tests make this feel just like iPhone's terrible text to speech. It is the one thing I utterly hate about iPhone. Ive tried apps that try to embed themselves into the iPhone keyboard and they always don't work out well. Hopefully this gets better and we can somehow get it into the iPhone more seamlessly.
MayeulC•3m ago
Speech to text I assume? Maybe it has to do with your a accent or pronunciation? You could contribute a bit to Mozilla's Common voice, if that's the case. I assume it is part of every STT training corpus.
INTPenis•31m ago
I don't think the challenge with speech to text was size of the binary. In my experience the challenge is understanding my 84 year old Croatian father with a sagging mouth after a stroke, when he's trying to write his autobiography.

I just setup Windows speech to text for him last week and it's great to see how he can write an entire page in 10 minutes, it would take him days using the keyboard.

But every single sound he makes with his mouth ends up on the page too.

ComputerGuru•25m ago
Sorry about your father. He needs a dictation model, not a general purpose speech-to-text model. They ignore umms and ahhs, change things like “an elephant, no a monkey, went up the tree” to “a monkey went up the tree,” support saying punctuation aloud sometimes, etc.

Gemini team just released Gemini 3.5 Transcribe that’s supposed to be good at this; it’s available via api: https://blog.google/innovation-and-ai/models-and-research/ge...

cgbur•3m ago
For essentially infinite and fast dictation I use https://github.com/cjpais/Handy on Parakeet streaming (cohere is far better, but slower and has a token output limit so you cant ramble for many minutes). And then just do a cleanup pass with a cheap LLM, it will in my experience, do far better than trying to voice control to go edit a sentence or change words. I just weave instructions into my writing. I understand this requires technical know-how, but for those with it, this is the best solution I have found to long form writing without my hands.
boplicity•8m ago
joewhale•28m ago
I initially read this as whistle to text, which would be way cooler.
mejutoco•12m ago
A good project for training an llm

https://en.wikipedia.org/wiki/Silbo_Gomero

jasonwatkinspdx•2m ago
I've met folks that descend from the Zapotec in southern Mexico, and they still use whistling language to talk to each other across mountain valleys.
charv•7m ago
Just like Marvel's Yondu!
armcat•20m ago
Those are insane benchmarks at this size. Well done!
mrkn1•19m ago
love seeing more sub-20MB, CPU-first models. if anyone wants a CLI built on the same ethos (no GPU, no cloud), been using yapsnap streaming Zipformer ASR, plus diarization and timestamps all on CPU! It supports 10 languages.
mrkn1•11m ago
https://github.com/kouhxp/yapsnap/
kamranjon•12m ago
Sooo I haven't really been super impressed with the needle models before, but this is very impressive. It transcribed multiple sentences I gave it with complex timing and words and in such a small footprint, I'm super impressed. Excited to see what types of things can be built with something like this, the performance seems very good.
aidotguru•11m ago
eager to see if working in android phones
rpdillon•8m ago
FUTO keyboard (open-source, free) runs entirely on-device and has extremely good STT accurary, especially with the 70M parameter model. I've used it for years now and love it.

https://futo.tech/

mo2art•3m ago
RuntimeError: audio limit is 30 s
There are many different challenges, each requiring their own solution. I, for one, really miss the old Google Assistant on my Android phone. It would very reliably play most songs that I wanted to hear on Spotify. Gemini fails at this almost every time, and is significantly slower. It's actually a difficult problem, as the songs people want to hear are regularly being released, are often associated with uncommon names, or have words in unusual orders, so normal LLM style tools just don't cut it.
testycool•7m ago
Unrelated: I love your username.
yymir•1m ago
i mean for something this small, it can be fit into a l3 cache on a cpu and be essentially always on various purposes