frontpage.
newsnewestaskshowjobs

Open Source @Github

fp.

Open in hackernews

Show HN: Neither HTML nor Markdown is enough: a way out of the AI doc dilemma

1•xiongjy2104•50m ago
A recent post from Anthropic's Claude Code team ("The Unreasonable Effectiveness of HTML", https://thariqs.github.io/html-effectiveness/) argued that with today's context windows the bottleneck is human attention — rich HTML for decision surfaces, Markdown for durable records. That split makes sense to me. What it leaves open, in my experience, is the durable record itself: what happens when that is the thing an agent must iteratively co-author? You end up picking one and paying for the other — a surface people will read, or a record you can keep — and whichever you pick, the other drifts.

Markdown treats the document as one long flat string — an agent has to quote context back just to say where an edit goes. The case against HTML I don't need to make: that post's own thread ran 274 comments (https://news.ycombinator.com/item?id=48071940). Mine is narrower — HTML is generated, not maintained. Markdown fails at addressing; HTML fails at being the record.

That gap is what GEML is for: a plain-text format that reads like clean Markdown to humans, but acts like an ID-addressable map of typed blocks to agents — every block carries a stable #id, so an agent reads or replaces one block instead of the whole file. Here's what pushed me into building it.

Recently I asked Claude Code a plain question:

> "Claude, taking your work editing the previous README and this article as an example, how do you actually edit sections? What I saw was you first asking me for various command permissions, grepping for keywords, etc. Maybe you can explain your workflow, and based on that we can discuss whether there are areas where GEML could improve things."

Here's what it did before making its very first edit on 8 comparison docs:

- `ls` -> didn't know where the files were - `grep -rniEl "comparison|对比"` -> scanned the entire repo - `git log -- docs/comparisons/` and `git log -- spec/` -> calculated commit timestamp diffs (the only way it could guess what went stale) - read `COMPARISON.md` in full (218 lines) - `grep -n "^#{1,3} "` to get a table of contents -> `sed -n 220,450p` -> `sed -n 520,620p` -> `sed -n 618,660p` -> `sed -n 660,690p` (sliced the spec in four chunks because it didn't know which sections were needed) - `git diff` on the spec + heading diffs -> finally discovered that sections 3.2 and 3.3 were added

Read-to-write ratio: roughly 20:1. The actual edits were ~30 small string replacements, but to know what to replace, it read the spec and 4 whole documents end to end.

`geml get file.geml '#id'` returns just that block; `geml set` swaps just that block and refuses the write if it would break document integrity.

You don't pick a side: `--to html` renders the rich view people read (the playground below is exactly that), `--to md` projects back to GitHub-Flavored Markdown. Both are outputs; the source is what you edit.

I measured the actual cost across 47 real edits on 4 documents (giving Markdown a realistic 46-line grep context window):

- Addressing cost ("saying where an edit goes"): Markdown costs 21x more tokens because the agent must quote text back until unique, whereas GEML writes an address. - Bytes read: 3.1x lower at the median in GEML (and GEML was never the more expensive of the two on any single edit). - In a full working day replay of mixed edits, reading came out 3.65x lower and addressing 7x lower. (Both benchmarks are 1-command reproducible in the repo.)

The rest follows from having an addressable, typed document model:

- Bound charts: charts bind directly to tables by ID, so data exists once and numbers cannot drift. - Strict validation: a dangling or cross-document reference is a compiler-style build error (`geml check`, non-zero exit — in the playground, hit "Break a reference" and watch it go red). - Block-level history: `geml history` keeps micro-revisions in a plain-text .gemlhistory sidecar — roll back a single block offline without polluting git commit logs.

Two things I'll pre-empt:

1. "Why not just extend Markdown?" — Pandoc and kramdown bolt on {#id}, but every extension creates another incompatible dialect, and the same .md already parses differently under CommonMark/GFM/Pandoc. Get/set-by-ID, bound charts, and reference checking require a unified document model and a build step, not ad-hoc syntax hacks. GEML has one grammar, one normative spec, and a conformance suite that a second parser — written from the spec alone, importing nothing from the reference implementation — reproduces case for case.

2. "Am I locked in?" — `--to md` takes it back out: prose, tables, notes, footnotes, code and math come back intact; block IDs and bound charts drop, and the tool names each one it dropped rather than hiding it, because Markdown has no syntax for them. Your prose is never trapped.

It's deliberately small: 1.0 spec (stable), MIT code / CC-BY spec, no adoption numbers to invent. If token cost, drifting numbers, and mangled text have bitten you during AI editing, give it a spin; if Markdown already works for your flow, Markdown is genuinely fine.

Playground: https://geml-spec.github.io/geml/playground/ Repo & Spec: https://github.com/geml-spec/geml CLI: `npm i -g @geml/geml`

That opening transcript was just my agent answering an honest question about its own workflow. So close the loop the same way — don't take my word for it: install the CLI and ask yours,

> "If these documents were GEML and you had geml list / find / get / set, what would those same edits have looked like?"

Better yet, have it actually do one. Post what it says — especially where the answer is "GEML wouldn't have helped here." That's the feedback I want most.

Comments

xiongjy2104•33m ago
look forward to hearing some comments to this idea.

Show HN: Speech Generation – Free text-to-speech in 18 languages

https://speechgeneration.net/
1•littlepp•29s ago•0 comments

Show HN: A skill for uploading full-resolution images to ChatGPT mobile runtime

https://github.com/Byte-Naut/send-lossless-images-skill
1•Byte-Naut•1m ago•0 comments

In defense of two-state theme toggles

https://joshcollinsworth.com/blog/in-defense-of-two-state-theme-toggles
1•speckx•1m ago•0 comments

A P2P Distributed OS Core Protocol in 40 Implementations

https://github.com/EntityChurch/entity-core-keystone
1•billatbillslab•4m ago•0 comments

Show HN: ReadPlusOne – Spanish stories built around the words you're learning

https://readplusone.com/
1•brandonc7•5m ago•0 comments

France's tax agency got hacked (in French)

https://www.cybernetica.fr/piratage-des-impots-comment-en-est-on-arrive-la/
4•zakxxi•8m ago•0 comments

Hot Chips 2026: Micron warns HBM wafer penalty is widening with every generation

https://www.tomshardware.com/tech-industry/semiconductors/micron-says-the-silicon-gap-between-hbm...
1•rbanffy•9m ago•0 comments

Let the Bond Market Speak

https://www.wsj.com/opinion/let-the-bond-market-speak-81529d74
1•jcfrei•9m ago•2 comments

Buyer beware: Those mummified remains might carry toxic spores

https://arstechnica.com/science/2026/08/modern-trade-of-mummified-remains-may-carry-its-own-mummy...
1•vintagedave•9m ago•0 comments

NASA Live Space Walk W Sophie Adenot and Anil Menon

https://www.youtube.com/watch?v=I0j7as4MLHk
1•katagaminator•10m ago•0 comments

Third drone and suspected military explosives found near Leipzig airport

https://www.theguardian.com/world/2026/aug/25/germany-drone-suspected-military-explosives-found-n...
2•osivertsson•10m ago•0 comments

Re: "Handwriting but not typewriting leads to widespread brain connectivity"

https://www.frontiersin.org/journals/psychology/articles/10.3389/fpsyg.2024.1517235/full
2•CharlesW•10m ago•0 comments

Geometry of quantum advantage in a guessing game

https://sciencespectrumu.com/beating-the-odds-in-a-guessing-game-with-a-single-quantum-box-f85c58...
1•armchairquantum•11m ago•0 comments

ZeroSSL Acme Endpoint Downtime

https://status.zerossl.com/
2•sigio•11m ago•1 comments

The Great Recession

https://nicholasdecker.substack.com/p/the-great-recession
1•surprisetalk•11m ago•0 comments

Why Claude Code garbles when you rotate your phone

https://getxtend.com/blog/terminal-state-you-cant-replay.html
2•mendidou•11m ago•1 comments

The History of Postgres Sharding

https://planetscale.com/blog/the-history-of-postgres-sharding
2•hollylawly•14m ago•1 comments

My experience with LLM-assisted tools in software development

https://ounapuu.ee/posts/2026/06/08/llm/
2•speckx•15m ago•0 comments

What I Don't See When I'm Envious

https://www.mooreds.com/wordpress/archives/3741
1•mooreds•15m ago•0 comments

The Hundred-and-First Child: axioms for non-compensating moral bookkeeping

https://kdc-apps.web.app/zingeving/the-hundred-and-first-child
1•kareldecherf•16m ago•0 comments

Show HN: Diet Cola themed Chrome Extension to track you Claude usage

https://dietclaude.com/
3•raghavtoshniwal•16m ago•0 comments

Rails: The Sharp Parts. The Block Is Not the Transaction

https://baweaver.com/writing/2026/08/22/rails-sharp-parts-the-block-is-not-the-transaction/
1•mooreds•16m ago•0 comments

Copyrightability of LLM-generated code: Can we license "vibe code" into FS?

https://fsfe.org/news/2026/news-20260825-01.de.html
1•mkesper•17m ago•0 comments

What the 100 biggest GitHub repos put in their AGENTS.md files

https://www.coldtea.ai/blog/agents-md-field-study
2•ohans•17m ago•2 comments

Quantization-Aware Healing:a compressed 4-bit model that outperf full-precision

https://huggingface.co/blog/MultiverseComputingCAI/quantization-aware-healing
2•grigio•19m ago•0 comments

OpenAI restores 5-hour Codex and Work limits for ChatGPT Plus users

https://9to5mac.com/2026/08/24/openai-restores-5-hour-codex-and-work-limits-for-chatgpt-plus-users/
4•MC995•19m ago•1 comments

Bots passed our captcha and solved the quiz, but would not pay $4

https://paragraph.com/@1lastpx1/a-bot-farm-ate-our-airdrop-and-got-nothing
1•LastPixel•20m ago•0 comments

Show HN: "Flying Toasters!" reverse-engineered 1:1, in the browser

https://backslasher.github.io/flying-toasters-reverse-engineered/
3•Backslasher•20m ago•1 comments

Show HN: TextToolsStudio – text tools that never leave the browser

https://www.texttoolsstudio.com/
1•harsh_patel14•20m ago•0 comments

What's Within Walking Distance

https://strado.info/
1•bookofjoe•21m ago•0 comments