frontpage.
newsnewestaskshowjobs

Open Source @Github

fp.

Who's afraid of Chinese models?

https://stratechery.com/2026/whos-afraid-of-chinese-models/
195•mfiguiere•13h ago•133 comments

Kimi Work

https://www.kimi.com/products/kimi-work
375•ms7892•7h ago•173 comments

Human mathematicians are being outcounterexampled

https://xenaproject.wordpress.com/2026/07/20/human-mathematicians-are-being-outcounterexampled/
184•artninja1988•5h ago•67 comments

Jelly UI: Soft-body physics for native HTML form controls

https://jelly-ui.com/
323•baldvinmar•7h ago•132 comments

Hacker wipes Romania's land registry database

https://news.risky.biz/risky-bulletin-hacker-wipes-romanias-entire-land-registry-database/
562•speckx•11h ago•317 comments

Nativ: Run frontier open models locally on your Mac

https://blaizzy.github.io/nativ/
161•aratahikaru5•6h ago•67 comments

Show HN: Immersive Gaussian Splat tour of grace cathedral, San Francisco

https://vincentwoo.com/3d/grace_cathedral/
64•akanet•4h ago•12 comments

Airport Simulator

https://airport.apunen.com/
690•apunen•14h ago•139 comments

The Psychology of Software Teams

https://www.routledge.com/The-Psychology-of-Software-Teams/Hicks/p/book/9781032963389
12•dcre•5d ago•2 comments

I wrote an bash enumerator because I was sick of xargs

https://numerlab.org/2025/07/20/bashumerate-enumerator/
48•wallach-game•4h ago•41 comments

Agent swarms and the new model economics

https://cursor.com/blog/agent-swarm-model-economics
109•jlaneve•6h ago•47 comments

China’s open-weights AI strategy is winning

https://werd.io/american-ai-is-locked-down-and-proprietary-its-losing/
964•benwerd•10h ago•770 comments

The Power of Awareness: Overcoming Surveillance Capitalism

https://www.scottrlarson.com/presentations/overcoming-surveillance-capitalism-with-awareness/
49•trinsic2•4h ago•5 comments

Launch HN: Bloomy (YC S26) – AI-powered mastery learning for K-12

57•alexsouthmayd•8h ago•73 comments

LEDs’ potential to save our night skies

https://spectrum.ieee.org/led-light-pollution
207•defrost•11h ago•157 comments

My two year old taught me constraint solving

https://thecomputersciencebook.com/posts/how-my-2yo-taught-me-constraint-solving/
25•bambataa•1w ago•12 comments

Corners Don't Look Like That: Regarding Screenspace Ambient Occlusion (2012)

https://nothings.org/gamedev/ssao/
146•firephox•9h ago•63 comments

How we measured AI writing across arXiv, and where the measurement breaks

https://unslop.run/blog/measuring-ai-writing-on-arxiv
191•dopamine_daddy•8h ago•138 comments

Perfection is not over-engineering

https://var0.xyz/posts/perfection-is-not-over-engineering.html
190•var0xyz•10h ago•86 comments

Shinjuku Station in 3D

https://satoshi7190.github.io/Shinjuku-indoor-threejs-demo/
145•Gecko4072•11h ago•28 comments

85.3 GFlops: Optimizing FP32 Matrix Multiplication on a Single AMD Zen 3 Core

https://github.com/houslast3/85.30-GFLOPS-Single-Core-FP32-Matrix-Multiplication-on-AMD-Zen-3
42•houslast•3d ago•15 comments

The drivers behind software delivery inefficiency

https://dl.acm.org/doi/10.1145/3800646.3800650
58•qaprof•6h ago•29 comments

You only need the frontier model for one single edit

https://stencil.so/blog/prewalk
63•jxmorris12•5d ago•15 comments

Kimi K3, Qwen 3.8, and Anthropic's (Potential) Unravelling

https://www.emergingtrajectories.com/lh/frontier-lab-economics/
278•cl42•9h ago•287 comments

Hyprland 0.55 announced the switch to Lua for its config files

https://hypr.land/news/update55/
123•matesz•7h ago•104 comments

The Voice of Google

https://www.newyorker.com/culture/the-weekend-essay/the-voice-of-google
164•littlexsparkee•9h ago•84 comments

Claude Fable produced a counterexample to the Jacobian Conjecture

https://xcancel.com/__alpoge__/status/2079028340955197566
659•loubbrad•22h ago•423 comments

Show HN: OTP Inspired actor supervisor based full stack templates

https://shipstacks.tech
8•annrap1d•11h ago•1 comments

Cross sectioning insects in an electron microscope with a femtosecond laser [video]

https://www.youtube.com/watch?v=NwhVJ7cv9B4
78•surprisetalk•9h ago•4 comments

Show HN: The0 – self-hosted runtime for trading bots, bring your own language

https://github.com/alexanderwanyoike/the0
5•mercutio93•11h ago•1 comments
Open in hackernews

1-Bit LLM in the Browser

https://huggingface.co/spaces/webml-community/bonsai-webgpu
111•simonebrunozzi•4d ago

Comments

stfurkan•16h ago
I'm currently working on an open-source engine [1] exactly for this purpose. If anyone wants to try or has any suggestions, I'm happy to listen :)

1. https://github.com/stfurkan/bitgpu

2. https://aidekin.com --> this is one of my projects that's currently using bitgpu engine

mgoetzke•16h ago
Interesting. But trying the demo with 27G model and simple 8+5 did not work.

Guessing there is some work to do still

``` calculate({"expression":"g"}) → error: only numbers, + - * / and parentheses are allowed error: No user query found in messages. What is 8+5? Use the calculator get_time({}) → Mon Jul 20 2026 10:36:28 GMT+0200 (Central European Summer Time) error: No user query found in messages. hi Hello! 2 tokens · 6.2 tok/s · prefill 4563 ms · confidence 87% which tools do you know about ? I'myepic know about the following are two tools I am aware of course, I currently i know about I know about me know I am aware of course I know I know I am aware of the following

calculate({"expression":" I know I am aware of know I am aware of I am aware I know I know know "}) → error: only numbers, + - * / and parentheses are allowed error: No user query found in messages.```

stfurkan•16h ago
Thank you for trying it :) I just added the 27B support 1 hour ago. Let me check and try to fix it. All other models should work as expected but if you have issues with them too, I can take a look.
alex7o•16h ago
Hay man, I really want a bonsai style embeddings or re-ranking model for you think sth like that is feasible?
stfurkan•16h ago
Hi :) There are some smaller embedding and re-ranking models (I am currently using some of them on my aidekin project) but having it like bonsai models will be great to have.
luca-3•13h ago
This is great for high-privacy use cases. If you could add an agentic loop that can call tools within the website using the user’s session, that would be great. Then the user can chat with an AI about sensitive information in the most secure way possible.

Also, do you have data on the technical requirements for this? If someone uses an old phone, e.g., does it still work?

stfurkan•12h ago
Hi, the engine itself already has tool calling capabilities but if you're asking for the aidekin I didn't add tool calling there on purpose.

It needs webgpu support but I haven't had a chance to test with different devices.

EGreg•12h ago
Yes, I’m very interested in exploring this stuff.

Want to connect and discuss it? https://calendly.com/safebotsai

stfurkan•12h ago
Hello, I sent you a LinkedIn connection request. We can chat first and schedule a meeting. Thanks
espetro•11h ago
Cool stuff! I see you're targeting two verticals here (runtime on top of ONNX + TS library), I'm curious did you consider simply creating a AI SDK Provider (https://ai-sdk.dev/providers/community-providers/custom-prov...) instead of crafting a separate library?

Besides it, how are you dealing with the harness the model can use on web? I'm currently building https://github.com/bolojs/bolo to extend the base harness agents can use on web and I'm interested in hearing from your experience on crafting tools for bitgpu

stfurkan•7h ago
Thank you :) I wanted to create it from zero myself so I can control it better. e.g. I started with ONNX support but now it support gguf models too. It might be a good idea to add some popular connectors in the future. Good luck with your project :)
lnenad•16h ago
Cool demo, but even the smallest model does 2.5 tps on my 7900xtx which is not great.
k3liutZu•16h ago
I get 9-10 tps on my MacBook Pro. Though the small model seem really limited
whstl•15h ago
Interestingly, I get 44 tps on Safari, MacBook Pro base model.
Tade0•14h ago
60tps in Chrome here - 14" MBP M4. Safari loads 3% of the model and then stops - repeatedly.

Basically zero using a 7700S in Chrome on Ubuntu though. Must be not working at all, as the performance difference can't be that large.

christkv•15h ago
Hmm I get 44,7 tokens/s on the smallest 1.7b one in the browser on my m1 max.
LorenDB•9h ago
18.8 tps on an 11th gen mobile i7 here.
dim13•15h ago
Fails "car wash" test miserably. Ask it:

> I want to wash my car. The car wash is 50 meters away. Should I walk or drive?

Ref: https://opper.ai/blog/car-wash-test

SwellJoe•15h ago
Even very large models failed tests like this up until a year or two ago...this is a very, very, small model.
woadwarrior01•14h ago
Works with the Bonsai-27B-1bit-mlx model on my mac.

Edit: this is running with MLX, and not with WebGPU on the browser, also it is a fail.

"""

You should *walk*.

Here’s why:

- *Distance*: 50 meters is very short — about 0.3 miles or 0.5 kilometers.

- *Time*: Walking takes roughly 2–3 minutes. Driving would take longer due to parking, starting the engine, and maneuvering.

- *Effort*: Walking is light exercise and avoids parking hassle.

- *Safety*: Less risk of misjudging parking space or damaging the car.

Unless you have a disability that makes walking difficult, walking is the most practical and efficient choice.

"""

geon•14h ago
How is that not a fail?
whstl•12h ago
I think GP also failed the car wash test.
Xalutiono
indianmouse•14h ago
Could not get any usable output from it for any real usecase.

Can't see any real value in using 1 bit llms except and apart for the sake of running it for demo or feel good factor of running a higher param model locally and in limited hardware!

Since use as it is very limited, could someone throw some light on any actual use case?

Lwerewolf•14h ago
Pretty sure this might be a duplicate. Regardless, tried the 1bit bonsai 27b gguf three different ways - their llama.cpp fork (prism, was it) on 2 machines (1255u/16g, m5 max/128g) and the web (on the 1255u). With llama.cpp it works. on the web it emits the same thing over and over. Prompt was "do you see anything wrong with this code: <C code, compound literal passed as a pointer to a function in an if statement>", result was literally "Yes, let'ssomestruct_struct_tsomestruct<a bit of similar garbage>structstructstruct<forever>.

Locally, it's definitely not the full 3.6 27b, but for ~6G with the context, it's quite impressive, and I like what this implies for larger models. Speeds on the 1255u are abysmal, though (~1TPS or less) - granted, CPU-only.

kees99•12h ago
Qwen is quite slow in general.

I'm getting ~0.5 tps from [qwen], and ~10 tps from [gemma] - roughly the same size, same quant, same hardware (8-core CPU) same software (llama.cpp).

[qwen] Qwen3.6-27B-Q4_K_M.gguf

[gemma] gemma-4-26B-A4B-it-qat-UD-Q4_K_XL.gguf

christkv•12h ago
Quite different models though Qwen 3.6 model has 27B active parameters. gemma-4 has 26B parameters with 4B active at any given time. You should compare with Qwen 3.6 35B A3B
DerNico•13h ago
Does not work Error: Unsupported device: "webgpu". Should be one of: wasm.
einpoklum•12h ago
I got 1.
estebarb•12h ago
I tried the smallest ome, the only available, and got an error:

``` Could not load.

Error: failed to call OrtRun(). ERROR_CODE: 1, ERROR_MESSAGE: Non-zero status code returned while running GroupQueryAttention node. Name:'/model/layers.0/attn/GroupQueryAttention' Status Message: Failed to create a WebGPU compute pipeline: A valid external Instance reference no longer exists. ```

leumon•9h ago
you need to enable webgpu
PeterSmit•11h ago
Cool idea. Crashes on iOS though, after downloading the 1.7B model
lelanthran•9h ago
1-bit is not much though. Here's what I got:

Me:

> Describe the process of pasteurisation

Response:

> pasteurization is a, which a, the which which is, the and process, the the past, the and the, and and and past, and past, and and and and the, the and past, process is the is is is is, and the, past, and and future, process, and and and and and and and the, past, process, process, the the, the the, is and the, the the, and and and, and process, and and the, past, and and and past, and the, the, and the past, the the, process, process, and past, the past, past, the and, the past, and and and and and and and and and the, and and and and and and and, the the, the the, the or and and, the the, the and the, which the, and the, the the, past, and n the process, and and and and, past, and, and and, the past, and, the the, past,, and the, the the, the is the, past, and and and and and, and and and past, the and the, the the, the the, are and past, and which the, the and n, n the, the n past, past, n the, and n, the the process, which past, the the, the n, the the, the is past, the the, is, past, the the, past, and past, process, the the, the the, the and and the, and which past, the

(and that basically just goes on and on like that)

wongarsu•9h ago
I got a decent response with the same prompt

But yes, this will happen. 1.7B is already a tiny model, running that at 1-bit is going to lead to some behavioral issues and will need tuning of temperature/topP/repeatPenalty/etc, and probably a software fallback that detects pathological cases like this and simply retries. Just like we all had to do in the times of GPT 2 and GPT 3

OutOfHere•9h ago
I confirm I got an extremely detailed and long response for "Describe the process of pasteurization." with 1.7B. I had to tune nothing.
nikhizzle•9h ago
leumon•9h ago
i found this space to work better/at all: https://huggingface.co/spaces/webml-community/bonsai-webgpu-...
spuz•9h ago
Unfortunately, this doesn't work for me. After loading for the first time, my first prompt had it generate an infinite series of exclamation marks. My subsequent queries just had it return nothing. I have plenty of RAM and VRAM. This is with Chrome on Linux Mint with 32GB RAM and a 12GB 3060.
om8•9h ago
This project needs webgpu -- I did it on cpu about a year ago.

My demo uses 2 bit quantization to run llama3 models on any device with enough ram.

https://galqiwi.github.io/aqlm-rs/

OutOfHere•9h ago
I run Bonsai-27B Q1_0 on my Android phone via the PocketPal app. It's slow but good. I wouldn't want to run fewer B. By any measure, my phone has less memory than my laptop browser.
aavci•9h ago
Other than privacy and security and simplified development, what are other valid reasons you'd prefer to run this instead of a backend LLM?
totallygeeky•5h ago
Impressive, I asked it how many r's in strawberry and it was immediately wrong with 2. any further questions I asked it proceeded to re-explain why there were 2 r's in strawberry, despite asking the silly "walk or drive to get my car washed" question. I struggle to see what these sorts of models could possibly be good at.
•
14h ago
Thats absolutly not the point of these models.

They are here to route things, do tool calls, ask experts.

OutOfHere•9h ago
That is completely false. It is a reasonable type of question and a reasonable expectation. You can't route correctly if you can't do reasoning.

Imagine a prompt:

```

INSTRUCTION: You're a customer service agent that routes customer service requests.

CUSTOMER MESSAGE: I haven't received my order and want a refund.

CONTEXT: A signature is required at delivery. No one was available to sign.

You can route to one of:

* Billing & refunds

* Shipping & tracking

* Order cancelation

* Level II customer service

```

The dumb model, e.g. 1.7B, wrongly routes to Billing. The smart model, e.g. GPT-5.6-Medium, correctly routes to Shipping. It matters. In the real world, the CONTEXT will even be 10-100x larger and noisier.

IMHO from guy working with heavily quantized models and trying to get results. I do think quantization at less than 4-bit requires more advanced mathematics than is usually used. See Google Scholar Babak Hassibi (founder of Prism ML). Generally am amazed about the resiliency of these new data structures to approximation.
cubefox•9h ago
There was also a NeurIPS paper which successfully trained 1-bit models directly without floating point weights or any quantization: https://arxiv.org/abs/2405.16339

I didn't see any further work in this direction however. The paper seems to have been mostly overlooked.

adrian_b•5h ago
While what can be done with less bits per parameter may be interesting, this is not really the important number for an LLM.

Much more important is the total amount of memory needed by an LLM, which can be decreased either by reducing the number of bits per parameter or the number of parameters. Arithmetic circuits are already cheap so for the inference cost the amount of memory transfers is more important.

It is not clear yet which number of bits per parameter allows the minimization of the total memory requirement, at a given LLM quality.

cubefox•4h ago
I thought the multiplications during training (backpropagation) are quite computationally expensive. The method from the paper above doesn't need any multiplications.