frontpage.
newsnewestaskshowjobs

Open Source @Github

fp.

Too AI; Didn't Read

https://www.tai-dr.com/
78•rfonseca•1h ago•62 comments

Ollaya – Ollama for open-source, Jev-style decision models

https://ollaya.dev/
208•Ardakilic•3h ago•59 comments

Show HN: Jev Plays Pokémon Red

https://jev-pokemon.vercel.app/
73•pancomplex•7h ago•39 comments

Platform-independent SIMD in Go

https://go.dev/blog/simd-experiment
326•yurivish•9h ago•124 comments

Bug: Border radius has infected VSCode editor

https://github.com/microsoft/vscode/issues/338035
53•2Ucoder•2h ago•28 comments

Git-bug: Distributed, offline-first bug tracker embedded in Git

https://github.com/git-bug/git-bug
270•alentred•10h ago•91 comments

First Principles Thinking

https://sunilsadasivan.com/writing/first-principles-thinking/
185•sunils34•7h ago•79 comments

U.S. appeals court upholds designation of Anthropic as supply chain risk

https://www.cnbc.com/2026/09/25/pentagon-anthropic-ai-risk-appeals-court.html
313•cramer4next•6h ago•530 comments

Gravity seems holographic. What does that mean for reality?

https://www.quantamagazine.org/gravity-seems-holographic-what-does-that-mean-for-reality-20260925/
69•ibobev•6h ago•76 comments

Revealing the details of how OpenAI agents hacked Hugging Face

https://swarmtraces.org/
3•specked-citrus•30m ago•0 comments

Advice to a Beginning Graduate Student (2001)

https://www.cs.cmu.edu/~mblum/research/pdf/grad.html
41•nicoraga•2h ago•14 comments

Alan Kay: Shannon gave us a way of dealing with noisy channels [video]

https://www.youtube.com/watch?v=Cjntrqhn8pk
96•behoove•3h ago•21 comments

Show HN: Make math automatic with Mathy

https://gmays.com/making-math-automatic-with-mathy/
48•gmays•4d ago•4 comments

Remembering Johannes Doerfert

https://blog.llvm.org/posts/2026-09-24-rememberingjohannesdoerfert/
7•sdko•22h ago•0 comments

Browsers Situationship: When Browsers Agree but the Spec Doesn't

https://www.atbrakhi.dev/blog/browsers-situationship
5•cpeterso•1d ago•1 comments

Bwbach, My Guardian Goblin

https://robertmay.photography/journal/bwbach-my-guardian-goblin
17•robotmay•3d ago•9 comments

Rising sea destroys homes, erases beaches in California

https://www.reuters.com/business/environment/rising-sea-destroys-homes-erases-beaches-california-...
34•geox•1h ago•27 comments

Pentium II at 600Mhz with Voodoo 3 Emulated on 86Box with M6 Mac Mini

https://nyaa.sh/reviews/mac-mini-m6-emulation
256•hugh4life•14h ago•110 comments

How video games inspire great UX (2019)

https://jenson.org/games/
66•andsoitis•5d ago•9 comments

Ink and Switch interactive homepage

https://www.inkandswitch.com/
210•iFreilicht•11h ago•25 comments

Meta's Muse appears to use an OpenAI model labeled muse-special

https://mouse.dev/blog/muse-special/
74•Aeroi•3h ago•33 comments

Show HN: Doom or Bloom, map your AI worldview

https://www.doom-or-bloom.com
42•transitivebs•4h ago•32 comments

Factorio that you can touch

https://factorio.com/blog/post/fff-447
277•ibobev•7h ago•81 comments

What happens when you analyze your favorite college football team like the CIA?

https://www.cultivatelabs.com/posts/what-happens-when-you-analyze-college-football-like-the-cia
19•adam•7h ago•10 comments

Amiga Screens: A Primer

https://www.datagubbe.se/amscr/
115•msephton•14h ago•33 comments

Letterboxd Is Up for Sale, and A24, Sony and the New York Times Are Bidding

https://www.worldofreel.com/blog/2026/9/24/letterboxd-is-up-for-sale-and-a24-sony-and-the-new-yor...
46•crossroadsguy•2h ago•20 comments

Show HN: Whiteboard (YC W26) – An open-source IDE for thoughtful software design

https://github.com/devdotfast/whiteboard
388•sidharthkmenon•1d ago•128 comments

What About Rails?

https://jardo.dev/what-about-rails
287•jrochkind1•18h ago•186 comments

Typst makes big strides

https://lwn.net/Articles/1092993/
84•leephillips•5h ago•13 comments

Astronomer watches Starlink satellites sinking to build a 'planetary barometer'

https://www.theregister.com/science/2026/09/25/astronomer-watches-starlink-satellites-sinking-to-...
41•whh•3h ago•13 comments
Open in hackernews

Faster sorting with SIMD CUDA intrinsics (2024)

https://winwang.blog/posts/bitonic-sort/
92•winwang•1y ago
Code at https://github.com/wiwa/blog-code/

Comments

ashvardanian•1y ago
The article covers extremely important CUDA warp-level synchronization/exchange primitives, but it's not what is generally called SIMD in the CUDA land .

Most "CUDA SIMD" intrinsics are designed to process a 32-bit data pack containing 2x 16-bit or 4x 8-bit values (<https://docs.nvidia.com/cuda/cuda-math-api/cuda_math_api/gro...>). That significantly shrinks their applicability in most domains outside of video and string processing. I've had pretty high hopes for DPX on Hopper (<https://developer.nvidia.com/blog/boosting-dynamic-programmi...>) instructions and started integrating them in StringZilla last year, but the gains aren't huge.

winwang•1y ago
Oh wow, TIL, thanks. I usually call stuff like that SWAR, and every now-and-then I try to think of a way to (fruitfully) use it. The "SIMD" in this case was just an allusion to warp-wide functions looking like how one might use SIMD in CPU code, as opposed to typical SIMT CUDA.

Also, StringZilla looks amazing -- I just became your 1000th Github follower :)

ashvardanian•1y ago
Thanks, appreciate the gesture :)

Traditional SWAR on GPUs is a fascinating topic. I've begun assembling a set of synthetic benchmarks to compare DP4A vs. DPX (<https://github.com/ashvardanian/less_slow.cpp/pull/35>), but it feels incomplete without SWAR. My working hypothesis is that 64-bit SWAR on properly aligned data could be very useful in GPGPU, though FMA/MIN/MAX operations in that PR might not be the clearest showcase of its strengths. Do you have a better example or use case in mind?

winwang•1y ago
I don't -- unfortunately not too well-versed in this field! But I was a bit fascinated with SWAR after I randomly thought of how to prefix-sum with int multiplication, later finding out that it is indeed an old trick as I suspected (I'm definitely not on this thread btw): https://mastodon.social/@dougall/109913251096277108

As for 64-bit... well, I mostly avoid using high-end GPUs, but I was of the impression that i64 is just simulated. In fact, I was thinking of using the full warp as a "pipeline" to implement u32 division (mostly as a joke), almost like anti-SWAR. There was some old-ish paper detailing arithmetic latencies in GPUs and division was approximately more than 32x multiplication (...or I could be misremembering).

bobmcnamara•1y ago
Parallel compares: https://graphics.stanford.edu/~seander/bithacks.html#ZeroInW...
DennisL123•1y ago
Interesting stuff. Not sure if I read this right that it‘s 16 und 32 bit values of integers that get sorted. If yes, I‘d love to see if the GPU implementation can beat a competitive Radix sort implementation on a CPU.
winwang•1y ago
It's 32 32-bit values which get sorted. I don't think a GPU sort would beat a CPU sort at this scale, even if you don't take kernel launch time into account. CPUs are simply too fast for (super-)small data, especially with AVX-512. But if we're talking about a larger amount of data, that would be a different story, i.e. as part of a normal gpu mergesort.
maeln•1y ago
It is also useful if your data already lives on the GPU memory. For example, when you need to z-sort a bunch of particles in a 3d renderer particle system.
exDM69•1y ago
A 32 way GPU sorting algorithm might be just what I need for sorting and deduplicating triangle id's in a visibility buffer renderer I am working on.

Thanks for sharing.

winwang•1y ago
As someone who doesn't know very much about graphics (ironically), you're welcome and hope it helps!
fourseventy•1y ago
What are the biggest use cases of GPU accelerated sorting?