frontpage.
newsnewestaskshowjobs

Open Source @Github

fp.

Open in hackernews

Launch HN: Vespper (YC F24) – SOTA Docx MCP

https://www.vespper.com/blog/launching-vespper-docx-mcp
8•topaztee•1h ago
Hey HN! We're Dudu and Topaz from Vespper (https://vespper.com). Vespper is an MCP that lets AI agents efficiently edit Word documents, powered by our fine-tuned model. It's currently 3× faster, 2× cheaper and more accurate than the closest alternative. Check out an overview of how the product works here: https://youtu.be/odKxsgPjzzw

We came to work on this problem after spending a year building an AI document editor for pharma companies. Before that, Topaz(myself) was a senior SWE at Snyk, working on distributed systems, and Dudu was a deep learning engineer at Viz.ai, building computer vision models for stroke detection. Our editor helped pharma companies generate regulatory documents (e.g. CSRs) to speed up their submissions. Initially, the output was Markdown, displayed in a WYSIWYG editor. However, users preferred working with their own Word templates. That's when the problems began.

AI agents aren't great at editing Word documents. A Word document is a zip file of verbose XML files following the OOXML spec. Even "small" changes require backflips, for example: adding a numbered list requires creating an entry in numbering.xml with a fresh ID and linking it back in document.xml, bolding a sentence requires splitting it into 3+ run elements. The list goes on.

This makes editing the zip directly (unzip + grep + sed) a bad idea for agents because they burn a lot of time + tokens on these mechanics. In practice, today's tooling falls into roughly three categories. You can let the agent write code against low-level libraries like python-docx or the Open XML SDK, you can give it an MCP with opinionated editing tools (SuperDoc, Office CLI, Adeu, etc), or you can round-trip the file through Markdown/HTML with something like pandoc/mammoth.js. None of them really work. The first two categories still burn the agent's context on Word mechanics instead of the task at hand (MCPs also introduce a new DSL to learn), and the third is very lossy (pandoc/mammoth.js/etc don't preserve enough fidelity).

From firsthand experience, these problems hurt performance in downstream tasks.

When we tried having our agent fill large documents, things broke quickly. The context window was already packed with customer data (files, user context, global rules, etc.), and the agent burned tokens + time on exploring the document and debugging failed edits. Filling a single CSR (clinical study report) took ~50 minutes, and the result was bad (missed fields/sections, broken styling, etc.). Harvey.ai's team reached similar conclusions: https://www.harvey.ai/blog/building-an-agent-for-complex-doc...

That's when we shifted our focus. We designed an MCP that lets agents edit Word docs as if they were editing HTML. The agent receives HTML, makes find-and-replace edits, and we reconcile those edits back into the original .docx file. We picked HTML over Markdown because it's structurally much closer to OOXML and because CSS associates styles with elements roughly the way OOXML does. We had to write our own DOCX→HTML converter, since pandoc and mammoth.js didn't preserve enough fidelity. To be clear, our DOCX → HTML conversion is lossy too. That's fine though, because we never convert the HTML back to DOCX. The HTML is just a projection for the agent, so it only needs enough fidelity for the agent to understand the structure and styling of what it's editing. The original file stays the source of truth, and we mutate it in place.

This also means the agent doesn't need to learn a new DSL. Editing a Word document feels just like editing an HTML file on the file system, something agents are already great at. A lot of DOCX MCPs hand agents dozens or even hundreds of tools to figure out on the fly. Our MCP exposes just three tools (read, search, edit). The Word document is completely abstracted.

After an agent sends us an edit request (an "old_html" and "new_html" pair), we reconcile it to the original .docx file. The reconciliation is powered by our fine-tuned model, a 3-8B base with a LoRA adapter. It takes the HTML diff as input along with the original localized OOXML block and emits the new OOXML. The "localization" is done deterministically: we take the anchors the agent specifies and we try to find their XML twins, so the reconciler model has a single responsibility. On this narrow task a small model reaches the level of a frontier, well prompt-tuned model, while being small and fast enough to sit in the hot path.

This project turned into months of work, but we're happy to release our v1. Our internal benchmark shows it's more accurate than the DOCX skill and raw python-docx while being ~2x cheaper and ~3x faster, mostly because it takes 3 tool calls (p50) per task whereas the DOCX skill takes 10 and Office CLI takes 13.

Things aren't perfect yet. For example, we don't support manipulating images or comments at the moment. That said, we're already seeing people use our MCP in various ways:

- Legal tech companies powering their live-editing flow in Office.js.

- AI startup optimizing people’s resumes and applying on their behalf.

- Govtech who need to draft policy memos.

- A life sciences startup using long-running agents to complete regulatory forms.

A note on privacy: our MCP runs in the cloud, so users send us their .docx files. We don't train on user data, and teams can opt for ZDR or self-hosting.

There's a free tier with 500 edits a month, and we’d love you all to try. We want to bump that later, but we're a small team and running a fine-tuned model isn't cheap :(

We'd love to hear your ideas and comments about docx editing in general! We'll be in the comments for the next few hours to respond

Comments

david1542•1h ago
Btw we open sourced a Word add-in project that shows how to build a Word add-in with Vespper MCP: https://github.com/vespperhq/examples/tree/main/word-add-in
topaztee•1h ago
sign up for free to try: https://app.vespper.com/

Definitely not Windows (Win 11 parody)

https://definitelynotwindows.com/
89•jjbinx007•50m ago•27 comments

Pirating the Pirates

https://mubi.com/en/notebook/posts/pirating-the-pirates
167•piotrgrabowski•2h ago•51 comments

So long Google, and thanks for all the nudes

https://lecaro.me/20260921-google-less.html
33•speckx•37m ago•8 comments

Claude Sonnet 5.5

https://www.anthropic.com/claude-sonnet-5-5
148•D2OQZG8l5BI1S06•43m ago•103 comments

Hijacking the PS5's RTMP Stream

https://yashgarg.dev/posts/hijacking-ps5-rtmp-stream/
87•ibobev•3h ago•11 comments

Show HN: HN.watch – Videos of all Hacker News posts

https://hn.watch/
57•mrborgen•3h ago•15 comments

Parley: Federated, decentralised chat that speaks plain IRC

https://git.mills.io/prologic/parley
248•davidcollantes•8h ago•122 comments

Who Wrote Elizabeth I's Most Scathing Letters?

https://www.smithsonianmag.com/history/who-wrote-elizabeth-is-most-scathing-letters-new-research-...
7•benbreen•54m ago•1 comments

When did Google get so weird?

https://sancho.bearblog.dev/google-weird/
1700•sancho-panza•22h ago•935 comments

The Teen Portraits That Captivated Sofia Coppola

https://www.newyorker.com/culture/photo-booth/the-teen-portraits-that-captivated-sofia-coppola
7•prismatic•1h ago•1 comments

Launch HN: Vespper (YC F24) – SOTA Docx MCP

https://www.vespper.com/blog/launching-vespper-docx-mcp
9•topaztee•1h ago•2 comments

OpenAI still doesn't seem to have a handle on all of its rogue AI activity

https://techcrunch.com/2026/09/28/openai-still-doesnt-seem-to-have-a-handle-on-all-of-its-rogue-a...
48•mikelgan•1h ago•43 comments

What Heraldry and Mon Can Teach Us About Building Visual-Identity Generators

https://benovermyer.com/blog/2026/09/japanese-vs-western-heraldry/
41•bovermyer•3h ago•13 comments

GrapheneOS – When an App Is Slow

https://blog.wirelessmoves.com/2026/09/grapheneos-when-an-app-is-slow.html
3•speckx•20m ago•0 comments

MongoDB CEO resigns to join Meta

https://www.reuters.com/technology/mongodb-ceo-desai-steps-down-lead-metas-enterprise-platform-20...
185•diek•3h ago•173 comments

Cf: The Agentic CLI for the Cloudflare API

https://blog.cloudflare.com/cloudflare-cf-cli-launch/
36•macleos•3h ago•12 comments

I made a visual workspace for AI Automations

https://www.biom.dev/
7•jduhking•48m ago•2 comments

37,500 border drawings: a map of the world as people remember it

https://www.habibicode.org/thedrawnworld
145•nicocarsui•10h ago•31 comments

Kids turned low-traffic NPR Spotify comments into a secret group chat

https://www.thisamericanlife.org/897/transcript
134•simonpure•3h ago•89 comments

What Would a Serious AI Product Look Like?

https://blog.glyph.im/2026/09/serious-ai-product.html
73•lumpa•7h ago•23 comments

Coding Is Not Solved

https://blog.alexewerlof.com/p/coding-is-not-solved
336•firstSpeaker•4h ago•341 comments

The problem is not AI code, but not knowing about system architecture or intent

https://www.ssp.sh/brain/the-problem-is-not-the-ai-code-but-nobody-knows-anything-anymore/
279•zazuke•2h ago•189 comments

Show HN: PaperMono, e-ink fridge magnet shopping list with mobile web page

https://github.com/seamusc/papermono-shopping-list
95•seamus_c•8h ago•45 comments

Solving a corn puzzle with CP-SAT

https://thill.me/2026/07/16/corn-puzzle-sat-solver.html
10•luu•21h ago•4 comments

Footguns with Postgres “at time zone 'UTC'”

https://bookofrevenue.com/blog/6ab81e9a97a13f0001f7e4e1/postgres-at-time-zone-u-does-not-do-what-...
146•birdculture•1d ago•82 comments

Show HN: Destroy Any Website with Stickman

https://destroy.spritefusion.com/
12•HugoDz•2h ago•5 comments

Owed a billion dollars in Nvidia stock

https://colo.to/nvidia-stock-narrative.html
1016•Eric_Gullichsen•16h ago•426 comments

Show HN: Free alternative to graphics design giants

https://scissor.studio/
66•sia_xi•9h ago•30 comments

Ember-1

https://fireworks.ai/blog/ember-1
566•gmays•1d ago•242 comments

Nissan's third generation e-POWER powertrain

https://www.nissan-global.com/EN/INNOVATION/TECHNOLOGY/ARCHIVE/E_POWER_GEN3/
144•mroche•16h ago•329 comments