frontpage.
newsnewestaskshowjobs

Open Source @Github

fp.

How Can We Enjoy the Necessary Practice to Master Valuable New Skills?

https://createadaptablelife.com/2026/09/how-can-we-enjoy-the-necessary-practice-to-master-valuabl...
1•mooreds•1m ago•0 comments

McMurtry Spéirling [video]

https://www.youtube.com/watch?v=xkXvL8YCtIE
1•manlymuppet•1m ago•1 comments

Talk to any self-hosted AI from your iPhone, Watch, Mac and car (open source)"

https://github.com/GigaDuckAI/conduck
1•pkrueck•1m ago•0 comments

Explore DDD Conference

https://exploreddd.com/
1•mooreds•1m ago•0 comments

Global Privacy Control

https://support.mozilla.org/en-US/kb/global-privacy-control
1•ahamez•3m ago•0 comments

Math Can't Go on Like This

https://www.theatlantic.com/technology/2026/09/math-crisis-openai-millennium-prize/688631/
2•taiwandongsuan•3m ago•0 comments

I built my own filing cabinet

https://www.alexstephen.me:443/writing/filing-cabinet/
2•rambleraptor•4m ago•0 comments

Building Xbox games with Three.js [video]

https://www.youtube.com/watch?v=ETfSiUHehY4
1•arbayi•5m ago•0 comments

NHTSA orders Tesla to prove its Cybercab is legal to sell, under oath

https://electrek.co/2026/09/15/nhtsa-tesla-cybercab-special-order-fmvss-certification/
1•Tomte•7m ago•0 comments

Rune raises $40M to deploy off-grid compute capacity at renewable energy sites

https://fastcompany.com/91607727/what-if-ai-data-centers-didnt-need-new-power-plants
1•utiiiD•8m ago•0 comments

Git Worktree Gotchas

https://www.olafalders.com/2026/09/16/git-worktree-gotchas/
1•oalders•8m ago•0 comments

Show HN: Free WhatsApp MCP (+UI) – Give Your AI Agents Access to WhatsApp

3•fabian_shipamax•9m ago•0 comments

Django: Simple Streaming / SSE with Mercure

https://blog.tmk.name/2026/09/16/django-simple-streaming-sse-with-mercure/
1•tristanmk•10m ago•0 comments

The smallest possible Linux distribution

https://distrowatch.com/weekly.php?issue=20260914
1•utiiiD•11m ago•0 comments

Show HN: Agentbox – Teleport your repo into sandboxes with no worktree juggling

https://github.com/madarco/agentbox
2•madarco•12m ago•0 comments

Show HN: Gokudo Wiki – a database and tools site for a new Roblox game

https://gokudowiki.com/
1•richardharmer•13m ago•0 comments

LLM based CI pipeline code generator for DSCI

1•melezhik•13m ago•0 comments

Show HN: Wenlan – a living wiki AI keeps current without overwriting your edits

https://github.com/7xuanlu/wenlan
1•h164654156465•15m ago•0 comments

Mayfly Chat: Transient Chat for Agents

https://blog.exe.dev/mayfly-chat
1•indigodaddy•17m ago•0 comments

Show HN: Plutus – click a release, see what moved on your cloud bill

https://demo.plutus-cloud.com
1•thechenderson•17m ago•0 comments

Estimating the Cost of Combat Operations Against Iran [pdf]

https://www.cbo.gov/system/files/2026-09/62756-Iran.pdf
2•amarcheschi•17m ago•0 comments

Banana Mode

https://hatchet.run/bananas
1•noleary•20m ago•1 comments

There is a channel to 900M weekly users. What goes in it?

https://www.lesswrong.com/posts/gJJ9YHzuBvwAXrthW/there-is-a-channel-to-900m-weekly-users-what-go...
1•ddp26•20m ago•0 comments

Archive Builder

https://apps.microsoft.com/detail/9n5hjq99k6wq?hl=en-US&gl=US
1•RENiXTech•21m ago•0 comments

Self-evolving agents need pain, reflection, and sleep

https://medium.com/@robertindie2016/self-evolving-agents-need-pain-reflection-and-sleep-2c63e6bd6630
1•aaronrobert•22m ago•0 comments

Ask HN: Has anyone measured how often agents use the skills you ship?

1•sohaibtariq•23m ago•0 comments

Mojo compiler is open for open source contributions [Mojo]

https://forum.modular.com/t/mojo-compiler-is-open-for-open-source-contributions/3499
1•ivell•24m ago•0 comments

Predictive database benchmarks vs. RF, AutoML, Elastic etc., up to 10M scale

https://aito.ai/docs/api/v2/benchmarks/
1•arauhala•24m ago•0 comments

TypeScript team chose Go over Rust

https://www.thetrueengineer.com/p/typescript-team-chose-go-over-rust
1•adletbalzhanov•25m ago•0 comments

U.S. has deployed space-control weapons in orbit, Air Force secretary says

https://spacenews.com/u-s-has-deployed-space-control-weapons-in-orbit-air-force-secretary-says/
2•alakra•25m ago•0 comments
Open in hackernews

Show HN: Swift-Qwen3.8-27B, -58.3% thinking, x1.95 speed, accuracy of xhigh

https://huggingface.co/ukisai/Swift-Qwen3.8-27b
7•kisjovan•49m ago
Hi everybody, we post-trained Qwen 3.8 27B to be more efficient by figuring out which tokens were linked to overthinking and penalizing them without "attacking" the reasoning length directly then fixed the accuracy with a bit of secret sauce (hint On-Policy Distillation) and achieved great results (-58% length, 1.95x speed up, <1% accuracy loss) so we wanted to open-source it and hear the feedback of the community.

This is the link to the model: https://huggingface.co/ukisai/Swift-Qwen3.8-27b / It's at 80k downloads in 3 days with independent evals here: https://www.reddit.com/r/LocalLLaMA/comments/1wg7dd5/ukisai_...

We are also providing a Free Research Purpose API (OpenAI compatible), courtesy of Nvidia who were kind enough to provide us with the GPUs. It's limited at 5RPM. https://ukisai.com/api/swift/v1/models

We also made a GGUF (Q1-Q8): https://huggingface.co/ukisai/Swift-Qwen3.8-27B-GGUF and there's also a few nice community quants with even lower/higher precision (Bartowski: https://huggingface.co/bartowski/ukisai_Swift-Qwen3.8-27b-GG...). The community also created amazing MLX, NVFP4, W4A16 and Uncensored versions you can find on Huggingface.

IMPORTANT: Our training approach is not a replacement for the reasoning effort settings, chat templates or token caps but is complementary and targets a completely separate issue (overthinking and "anxiety-like" reasoning loops prior seen in PTQ, but as far as we identified also prominent in BF16 of this size class LLMs as well). Contrary to popular belief, these specific patterns do not contribute to answer quality when properly targeted. (our thesis being: reasoning length IS extremely important and should NOT be shortened by force, but rather optimized). This is also demonstrated bellow in our xhigh vs medium effort benchmark. The goal is to keep xhigh accuracy while reducing only the unnecessary part of thinking.

The TLDR of our thought process, research, training and a link to the Meta paper that inspired us is in the first comment.

The benchmarks:

Swift-27B vs Qwen3.8-27B (BF16, all benchmarks ran x5, thinking effort xhigh)

GPQA-Diamond: 88.4% -> 88.3%, 58% fewer median tokens

LiveCodeBench v6: 76.8% -> 81.6% (+4.8pp, due to default truncation in LCB it is not performance gain), 46% fewer median thinking tokens

Terminal-Bench 2.1: 66.7% -> 65.8%, 39% fewer median tokens

MMLU-Pro: 85.5% -> 85.0%, 28% fewer median tokens

C-Eval: 90.0% -> 90.6%, 19% fewer median tokens

IFBench: 73.5% -> 71.8%, 51% fewer median tokens

AIME 2026: 98.7% -> 94.0%, 50% fewer median tokens

HMMT (Nov 2025): 99.3% -> 96.0%, 46% fewer median tokens

ERQA (vision): 67.5% -> 66.3%, 55% fewer median tokens

Token savings hold at every reasoning effort (mean thinking reduction): xhigh 41%, medium 23%, low 26% (albeit with accuracy loses of 1-4% on medium and 1-2% on low which we need further testing for)

Swift at xhigh vs the base's own effort settings on GPQA-Diamond (198 questions x 5 seeds):

Base xhigh: 88.4%, 6,642 median tokens

Swift xhigh: 88.3%, 2,771 median tokens

Base medium: 84.1%, 1,753 median tokens

So Swift keeps xhigh accuracy at under half the tokens, and beats base-medium by 4pp at roughly 1.6x its tokens.

End note:

While we are keen on complete open-source, we still need to keep a part of our training and data private, being a new lab. The license is not Apache 2.0, but it only affects companies >$1M. We hope this does not pose a problem for the community and are open to feedback on it.

We want to contribute as much as possible to the community and would really appreciate feedback on our work, quantization or Swift model requests. We are working on Swift 3.8 Flash Next and have so far gotten up to -53% thinking token usage. We have strong indicators our methodology is reproducible on other model families as well and are asking the community which ones you want us to optimize next.

Comments

kisjovan•46m ago
I will TLDR you on our thought process, research, training and benchmarks.

1. When running our quantized Qwen 3.8 27B instances we were very annoyed by random reasoning loops (in the paper bellow refered to as "overthinking errors". These random loops were persistent throughout medium and low reasoning settings.

2. We found a paper by Meta that's supposed to target this phenomenon in PTQ, but when used straight out of the box got mixed results. https://arxiv.org/abs/2606.00206

3. We figured to try if it's a matter of the targeting the right keywords and tuning the parameters, so we used our 8xH100 box and and generated a large amount of different (ofc out of distribution) domain (coding, language, vision, agentic) traces.

4. We then grouped the ones with overthinking and found "common denominator" tokens between them and targeted the most prominent ones.

5. We then built an inference-time penalizer of those tokens as seen in the paper with the hopes of simply generating traces and doing cross-entropy SFT over them.

6. Did not work at all, but the penalizer seemed to work much better than the tokens provided in the paper and not only for lower precision models but for bf16 as well. Hence we kept experimenting with it. We built a loss function using the tokens we identified and ran LoRa SFT over the traces prev generated and reasoning seemed to be falling off significantly but the accuracy seemed to follow. The reasoning reduction seemed to be generalizing.

7. After a significant amount of tinkering (literally since the day of Qwen 3.8 27B release) we were satisfied with the reasoning reduction. After that we searched for ways of restoring the accuracy. We experimented with several methods, including RL(GSPO), On-Policy Distillation and using the ThinkingCap 3.6 27B adapter chunks until we were satisfied with our accuracy loss. We managed to restore it to <1% loss on almost all of our OOD in house tests

8. We then performed intensive intensive benchmarks, across several reasoning efforts, precision variants etc. We ran into a few problems, one of which is that to get a reliable score we needed to run each benchmark 10x (5x on base + 5x with our adapter, this being the standard procedure on the Qwen 3.6 27B model card on Terminal Bench which we followed). After running it, the performance converged to 40-60% token reduction with <1% accuracy loss across GPQA, MMLU, Terminal Bench 2.1, LiveCodeBench v6, ERQA, C-Eval, IFBench, HMMT25, with an exception being AIME26 with an accuracy loss of 4.6%, which we later linked to a bug during training with a specific token relevant for math-related reasoning being penalized and are planning to fix it in an updated release.

danilotodorovic•43m ago
I've been using this and it's quite amazing for the amount I've used it. Thanks for the hard work.
kisjovan•42m ago
Thank you so much!! It would be great if you could share some numbers with the community :)
founderjoeNY•41m ago
How does it perform on real long horizon tasks?
kisjovan•28m ago
You should check out our TerminalBench2.1 score for that, the tasks there can run up to 4h! We got almost no loss (we got Base 66.74% vs Swift 65.84%) with -38.7% thinking token reduction. There's also quite a few independent evals on the reddit link as well.

Please let me know when you try the model and if I can help you set it up :)