frontpage.
newsnewestaskshowjobs

Open Source @Github

fp.

Open in hackernews

Beating GPT-5.6 Sol on retrieval with 100x cheaper open models

https://neon.com/blog/how-castform-neon-beats-frontier-models-on-price-and-efficiency
43•moonikakiss•1h ago

Comments

mrinterweb•22m ago
There is so much opportunity for purpose built models like this. Ideally a harness should spin up a subagent to offload to targeted models for specific tasks like this. I know this is not a novel idea. Claude code does some of this by handing off the "explore" agent work to haiku. I just love seeing that specialized LLMs are being developed.
foota•21m ago
I feel like the future is people building applications with tightly integrated LLMs that work hand in hand with the application's own lifecycle and code.

I also didn't realize that people were using agentic harnesses for search, it's an interesting idea. If the context length is short enough it should be fairly cheap compared to running "normal" agentic coding workloads where you have O(100k) context length for doing almost anything.

Malp•14m ago
There are! Chroma has Context1, SID has SID-1, and you'd actually be surprised at how easy it is to post-train your own with pretty good pass@ recall@ ndcg@ etc.

There's also Hornet who have shared some interesting talks & blogs lately. I don't know that I'd exclusively use agents for retrieval the way Neon outlines here as well. I think distillation similar to what ZeroEntropy has done for bespoke retrieval & reranking with _some_ agent manipulation on top-k results works better (IME).

devolving-dev•4m ago
Models keep on improving though, so doesn't fine tuning become an ongoing task with ongoing maintenance burden?
ramon156•14m ago
Bit unrelated, I realized that z.ai gives you access to deepseek 4 flash. It's incredible how well it performs when given a detailed spec. I'm not sure I've seen a model one-shot like that, and I was already impressed by gemma 4's speed and efficiency.
aliljet•11m ago
There is a more serious question in here that's not being answered. How effective is the retrieval in finding buried needles in larger and larger haystacks. And there's a correlary question, how effective could you be in finding paired needles in that haystack where you need to hold a needle to unlock finding another needle.
richwater•5m ago
One thing that plagues [insert current FAANG] is the large amount of corpus knowledge that is outdated/misleading or just plain wrong. I'm curious how this addresses that if it's deriving the reward function from the corpus itself.
JCharante•3m ago
I have done my own testing and found that smaller models can beat their larger siblings on fact retrieval from documents. I haven’t investigated it in depth with a large enough dataset but my guess is that larger models overthink it while smaller ones just do it. I would like if they compared this with 5.6 Luna instead.

Price changes in consumer goods and services in the United States

https://ourworldindata.org/grapher/price-changes-consumer-goods-services-united-states
1•skadamat•32s ago•0 comments

Meta Muse Spark 1.2

https://developer.meta.com/ai/models/muse-spark/
1•wojciem•37s ago•0 comments

Unified Representation for Continuous-Latent Diffusion Language Modeling

https://arxiv.org/abs/2608.02602
1•E-Reverance•39s ago•0 comments

How to Survive in a Louisiana Swamp

https://unherd.com/2026/08/how-to-survive-on-a-louisiana-swamp/
1•bookofjoe•1m ago•0 comments

Ban the Throbber

https://banthethrobber.neocities.org/
1•kyledrake•2m ago•0 comments

TikTok Enters the Multichannel Fulfillment Race with FBT MCF

https://www.geekseller.com/blog/tiktok-enters-the-multichannel-fulfillment-race-with-fbt-mcf/
1•kull•3m ago•0 comments

PEP 842 – Module Exports

https://peps.python.org/pep-0842/
1•Ravencentric•4m ago•0 comments

Microsoft makes OpenAI GPT-5.6 Sol default in GitHub Copilot for staff

https://www.cnbc.com/2026/08/05/microsoft-makes-openai-gpt-5point6-sol-default-in-github-copilot-...
1•kjhughes•4m ago•0 comments

My Son's Internet

https://www.gordonmclean.co.uk/2026/08/04/my-sons-internet-2/
1•speckx•5m ago•0 comments

Tesla, Inc. vs. Angstrom Automotive Group, LLC

https://www.courtlistener.com/docket/73664213/1/tesla-inc-v-angstrom-automotive-group-llc/
2•hnburnsy•6m ago•1 comments

Microsoft's BitNet 2B on a 4 GB Raspberry Pi 5, in one 1.2 GB file

https://github.com/geisten/geistlib
1•geisten•6m ago•0 comments

SparklingTree: 30-40% faster specdec than DSpark by combining DDTree and DSpark

https://jwlabs.vercel.app/post/sparklingtree
1•shreybirmiwal•7m ago•0 comments

People prefer stories written by AI when told they're written by a human

https://techxplore.com/news/2026-08-people-stories-written-ai-told.html
1•wjSgoWPm5bWAhXB•7m ago•0 comments

Show HN: ShiftGrid – open-source transparent prompt-engine for pentests

https://github.com/BuFuuu/shiftgrid/blob/main/demo.gif
1•bufuu•8m ago•0 comments

Unpacking ChatGPT Work: The Agent for a Billion Users

https://www.latent.space/p/unpacking-chatgpt-work
1•haritha1313•8m ago•0 comments

Show HN: My receipt printer prints an original artwork every morning

https://github.com/matt-w-horn/morningprint
1•spectraldrift•8m ago•0 comments

Muse Code and Muse Spark 1.2

https://research.meta.ai/blog/introducing-muse-code-and-muse-spark-1-2
3•paulkrush•9m ago•1 comments

Releasing Muse Code in beta today, and Muse Spark 1.2

https://twitter.com/finkd/status/2085080750034940201
2•pmxi•12m ago•0 comments

DeepSeek V4-Flash-0731 is 12 pts more censored than preview (selectively)

https://www.ctgt.ai/research/v4-flash-0731-drift
3•cgorlla•12m ago•0 comments

Why AI agents aren't adopted widely

https://invertedpassion.com/why-ai-agents-arent-adopted-widely/
2•twapi•13m ago•0 comments

A new kind of rental scam strikes SF: He lost $18,000 to a duped Zillow listing

https://sfstandard.com/2026/08/05/san-francisco-rental-listing-scam/
2•randycupertino•15m ago•1 comments

How to Cross-Compile Rust for Windows from Linux

https://www.rustfaq.org/en/how-to-cross-compile-rust-for-windows-from-linux/
2•auraham•16m ago•0 comments

Show HN: DrakeAI – expense tracker you log by voice or text, no bank sync

https://drakeai.app/
2•a_protsyuk•17m ago•0 comments

Linux's Staging Area to Now Reject LLM Patches, Except for Real Security Fixes

https://www.phoronix.com/news/Linux-Staging-Reject-LLMs
2•Bender•20m ago•1 comments

FSF job opportunity: Engineering and Certification Manager

https://www.fsf.org/news/2026-job-opportunity-fsf-engineering-and-certification-manager
2•infognu•20m ago•0 comments

Show HN: Greenlight – Preflight Scanner for App Store and Google Play

https://github.com/RevylAI/greenlight
2•ethanzhoucool•20m ago•0 comments

Nvidia Becomes a Premier Sponsor of LVFS / Fwupd

https://www.phoronix.com/news/NVIDIA-Premier-Sponsor-LVFS
2•Bender•21m ago•0 comments

"AI" will never become conscious

https://mattbee.mataroa.blog/p/no-ai-will-never-become-conscious/
4•speckx•21m ago•0 comments

TikTok A/B testing withheld safety feature from ~10% of US users, lawsuit claims

https://www.bloomberg.com/news/features/2026-08-04/confidential-tiktok-report-shows-algorithm-saf...
2•anigbrowl•22m ago•0 comments

Cloudflare Announces Open-Source Cloudflare OS as AI "Operating System"

https://www.phoronix.com/news/Cloudflare-OS
2•Bender•25m ago•1 comments