frontpage.
newsnewestaskshowjobs

Open Source @Github

fp.

Open in hackernews

Show HN: Fast and Quality Code Chunking with Chonkie

1•snyy•1y ago
Hi HN,

We’re Chonkie (https://github.com/chonkie-inc/chonkie) — we build open source tools that help split documents into meaningful chunks for use with AI models.

When you use LLMs over large documents or codebases, you often need to break them into smaller parts to fit the model’s context window. Our chunkers do this in a smart way: they preserve structure and meaning, so only the most relevant pieces are passed into the model. This reduces hallucinations, avoids confusion, and improves performance and accuracy.

Today we’re launching our Code Chunker — a fast, structure-aware way to break down source code into high-quality, token-aware chunks.

How it works:

(See the code: https://github.com/chonkie-inc/chonkie/blob/main/src/chonkie...)

Code Chunker uses tree-sitter (https://tree-sitter.github.io/tree-sitter/) to parse your code into an abstract syntax tree (AST). It then recursively merges and groups nodes in a way that respects both code structure and token limits.

It supports all languages that tree-sitter supports, and is designed to preserve formatting and semantics. Large functions or class definitions won’t be split in the middle of a block — instead, we dive recursively into the AST to produce clean, coherent chunks that fit your configured token budget.

What it’s useful for:

  - Embedding-based code search

  - RAG (retrieval-augmented generation) over codebases

  - Long-context analysis of code

  - Preparing repos for fine-tuning or pretraining
Try it out:

  - Open source package: https://docs.chonkie.ai/chunkers/code-chunker

  - Hosted playground (free with account): https://cloud.chonkie.ai
Happy Chonking!

Windows on Itanium also provided for hot-patching, in an even simpler way

https://devblogs.microsoft.com/oldnewthing/20261001-57/?p=112747/
1•zdw•46s ago•0 comments

Stable Matching Problem

https://en.wikipedia.org/wiki/Stable_matching_problem
1•azhenley•1m ago•0 comments

Open source SDK to build Muse gadgets

https://github.com/facebookincubator/muse-gadget-sdk
1•kristianpaul•1m ago•0 comments

What if AI worked at 1.000.000 tokens per seconds?

https://www.echohive.ai/one-million-tokens-per-second
1•echohive42•3m ago•0 comments

Why It's Time to Rename Cybersecurity Awareness Month

https://semgrep.dev/blog/2026/rename-cybersecurity-awareness-month/
1•Garbage•3m ago•0 comments

Pi Durable, as Statecharts

https://scxmljs.tinyactors.dev/demos/pi-durable/
1•handfuloflight•4m ago•0 comments

Backend programming languages ranked by LLM recommendations

https://preseason.ai/rankings/devtools/backend-language
1•thedreammachine•4m ago•0 comments

Why I started (and stopped) making games (2003)

https://austinhenley.com/blog/makinggames.html
1•azhenley•5m ago•0 comments

We ported the original Doom to SQL

https://cedardb.com/blog/sqldoom/
1•shscs911•7m ago•0 comments

New Mexico wants Meta to pay $40B in penalties after data privacy trial

https://finance.yahoo.com/media-advertising/articles/mexico-wants-meta-pay-40-231305611.html
1•smurda•12m ago•0 comments

Canada, EU plan to link next-gen payment systems, easing transactions

https://www.theglobeandmail.com/politics/article-canada-eu-link-payment-systems-joint-statement-s...
2•MadrasTh0rn•15m ago•0 comments

Should you give the scarcest things to the fastest coders?

https://digitalseams.com/blog/should-you-give-the-scarcest-things-to-the-fastest-coders
1•bobbiechen•24m ago•0 comments

NVX: An Ultra-Light Micro-VM Sandbox from Microsoft

https://github.com/microsoft/nvx
1•jeswin•27m ago•0 comments

Vim-pi-chat: make Vim a pi-powered editor

https://github.com/galeone/vim-pi-chat
1•me2too•33m ago•0 comments

SWC stops accepting external PRs due to AI-generated content

https://twitter.com/swc_rs/status/2105505428893544847
4•csmantle•39m ago•1 comments

Everything Got Worse After 2020

https://www.youtube.com/watch?v=Do8PuebNSyY
3•cable2600•42m ago•0 comments

Code Must Be Faster

https://www.monaddle.com/blog/your-code-must-be-faster
2•roflc0ptic•43m ago•0 comments

Radio signal detected for first time from a planet outside our solar system

https://www.cnn.com/2026/10/02/science/radio-signal-detection-exoplanet-beta-pictoris-b
2•slater•51m ago•0 comments

Mars' North Polar Ice Is Much Cleaner Than Scientists Thought

https://scitechdaily.com/mars-north-polar-ice-is-much-cleaner-than-scientists-thought/
1•snarky-comments•1h ago•0 comments

Extra Big Ass Intelligence

https://www.extrabigassintelligence.com/
2•34679•1h ago•0 comments

Cloudflare Ohttp Gateway

https://blog.cloudflare.com/announcing-cloudflare-ohttp-gateway/
7•est•1h ago•0 comments

Show HN: Claude Scrolls TikToks for Me

https://tryrevline.com/
2•ZuraMakaradzeHe•1h ago•3 comments

Delta Is the Only Big Four Airline That Won't Use Starlink–Elon Musk Is Furious

https://www.wsj.com/business/airlines/delta-is-the-only-big-four-airline-that-wont-use-starlinkan...
2•doener•1h ago•0 comments

Lean Game Server: A repo of learning games for Lean

https://adam.math.hhu.de/
2•crescit_eundo•1h ago•1 comments

Hanami, Why?: Bits and Bobs

https://aaronmallen.me/writing/hanami-why-bits-bobs
1•thunderbong•1h ago•0 comments

Ask HN: Will source code become expensive if developers stop using GitHub?

4•debamitro•1h ago•2 comments

The Void: From Zero-Byte Responses to Continuation Control

https://zenodo.org/records/23070524
2•rayanpal_•1h ago•0 comments

Show HN: Fakeflac-go – A tool for quickly finding "fake" .flac files

https://github.com/drichline/fakeflac-go
1•DakotaR•1h ago•0 comments

Show HN: Hall Monitor. See all your agents across machines on your Mac's Notch

https://github.com/hiteshbandhu/hallmonitor
2•HiteshBandhu•1h ago•1 comments

SoftServe: A Scalable Quasi-Newton Method for Deep Learning

https://arxiv.org/abs/2610.02182
1•E-Reverance•1h ago•0 comments