frontpage.
newsnewestaskshowjobs

Made with ♥ by @iamnishanth

Open Source @Github

Open in hackernews

Show HN: We built a document AI, you only pay for good data

https://undatas.io/
1•jojogh•2d ago
Hi HN, I’m one of the creators of Undatas.io. We've been working with document AI for a while and got tired of the standard model: upload sensitive files to a third-party server, send them to a black-box API, and pay for every page, even if the output is garbage. We decided to build a platform that fixes this from first principles. Today we're launching V3, which is a major overhaul focused on data control, verifiable results, and a fair pricing model. Here’s a breakdown of the technical approach: 1. Securely Process from Your Cloud Storage (S3, Box, Dropbox, GCS, Azure): Instead of uploading files to us, you connect your own cloud storage. The architecture works via native integrations. For major providers like AWS, GCP, and Azure, you grant our system secure, temporary access using scoped-down credentials (e.g., an IAM role you control). For services like Box and Dropbox, you connect your account via a standard OAuth 2.0 flow, granting read-only permissions. In all cases, the principle is the same: our workers fetch the object for in-memory processing and it's immediately discarded. Your documents never land on our persistent storage. 2. "Glass Box" Visual Validation: To solve the black-box problem, we built an interactive workspace. The backend returns JSON with detailed bounding box coordinates for every extracted token, line, and table cell. Our frontend uses these coordinates to map the structured data back to the source document image, allowing you to click on any JSON element and see it instantly highlighted. You can see a live, no-signup demo of the UI here: https://undatas.io/ 3. State-of-the-Art Table Extraction: This was our biggest R&D effort. Most tools fail at complex tables (merged cells, nested headers, no borders). Our model moves beyond simple heuristics. It uses a hybrid approach, combining a vision transformer (ViT) to understand the visual layout with graph neural networks (GNNs) to reconstruct the logical cell-to-cell relationships. This allows it to correctly parse table structures that would otherwise be ambiguous. 4. "Pay for Quality" API: This is built into our API workflow. When you process a document, the results for each page enter a "pending" state. You use the visual validator (or an approval webhook) to review them. Only when you explicitly "accept" a page's results is the transaction committed and your credits are used. A "discard" call costs nothing. We're trying to be as open as possible. While the core engine is proprietary, we are open-sourcing our client SDKs and other tools. Link to try the full platform: https://undatas.io/ A signup is needed for the full platform to manage API keys and credits. To make it easy for the HN community to test everything, we’re giving everyone 5,000 free credits (good for ~1,000 pages) and a 7-day trial of all features, including the private cloud connections. As a thank you to early adopters from this community, the first 50 people who subscribe to a paid plan will get a 25% lifetime discount. We’re here all day to answer your questions. Thanks for checking it out.

It might seem a bit silly, but this is my "philosopher's stone" right now

https://twitter.com/henophilia/status/1954215929098711249
1•sigalor•1m ago•0 comments

Cloudflare recommends migrating from Workers to Pages

https://developers.cloudflare.com/workers/static-assets/migration-guides/migrate-from-pages/
1•mmarian•1m ago•0 comments

War Has Changed: Foreign Influence Networks and the Art of Strategic Deflection

https://vasily.cc/blog/war-has-changed/
1•nabla9•1m ago•0 comments

Russia Has an Arsenal of New AI Drones Built with Smuggled Nvidia Chips

https://www.forbes.com/sites/davidhambling/2025/08/08/russia-has-an-arsenal-of-new-ai-drones-built-with-smuggled-us-chips/
1•sharpshadow•3m ago•0 comments

Someday Is Already Here

https://pieces.app/blog/someday-is-already-here
1•thunderbong•3m ago•0 comments

Trump announces 100% tariff on computer chip

https://www.usatoday.com/story/money/2025/08/08/trump-tariff-chip-semiconductor-consumer-prices-impact/85562097007/
1•taimurkazmi•10m ago•0 comments

After User Backlash, OpenAI Is Bringing Back Older ChatGPT Models

https://www.cnet.com/tech/services-and-software/after-user-backlash-openai-is-bringing-back-older-chatgpt-models/
2•pera•10m ago•0 comments

Dyson Sphere Could Bring Humans Back from the Dead

https://www.popularmechanics.com/science/a65615574/dyson-sphere-digital-resurrection-human-immortality/
2•Bluestein•11m ago•0 comments

Labubu AI

https://labubuai.net
1•MintNow•13m ago•0 comments

Classification of the Approaches to the Technological Resurrection

https://www.academia.edu/36998733/Classification_of_the_approaches_to_the_technological_resurrection
1•Bluestein•16m ago•0 comments

LLM advises to delete the Linux dynamic linker during a troubleshooting session

https://old.reddit.com/r/linux4noobs/comments/1mlveoo/help/
2•Santosh83•30m ago•0 comments

The Most Nihilistic Conflict on Earth

https://www.theatlantic.com/magazine/archive/2025/09/sudan-civil-war-humanitarian-crisis/683563/
1•YeGoblynQueenne•36m ago•1 comments

Show HN: I'm trying to quit vape and hoping someone could join me

https://www.iquitvape.com/
1•jayqinohboi•40m ago•0 comments

What Declarative Languages Are

https://semantic-domain.blogspot.com/2013/07/what-declarative-languages-are.html
1•fanf2•46m ago•0 comments

I cancelled my Chat GPT subscription today

6•dontlike2chat•49m ago•0 comments

Onion: Stack Language Compiled to Lua

https://github.com/yumaikas/onion
2•Bogdanp•54m ago•0 comments

A Fully Automatic Morse Code Teaching Machine (1977)

https://c2.com/morse/
1•austinallegro•58m ago•0 comments

'It's missing something': AGI, superintelligence and a race for the future

https://www.theguardian.com/technology/2025/aug/09/its-missing-something-agi-superintelligence-and-a-race-for-the-future
1•nhojb•58m ago•0 comments

Workers whose jobs AI can do less likely than other workers to be unemployed

https://eig.org/ai-and-jobs-the-final-word/
1•JumpCrisscross•59m ago•0 comments

Hospital Shift Scheduling with OR-Tools

https://barkeywolf.consulting/posts/hospital-scheduling/
1•jjhbarkeywolf•1h ago•0 comments

Show HN: The "Firebase" for MCP Servers – Build, test, and deploy MCP servers

https://www.contexaai.com/
1•rupesh_raj29•1h ago•0 comments

Ask HN: Which Do you know any open source games?

2•Forgret•1h ago•1 comments

The new era of house music

https://open.spotify.com/playlist/2sCu2R0XnUTw9na0ofT4vb
1•playlsd•1h ago•0 comments

Logarithmic mean energy optimization a metaheuristic algorithm

https://www.nature.com/articles/s41598-025-00594-2
1•bryanrasmussen•1h ago•1 comments

Money Habits That Separate Successful Traders from the Rest

https://propfirmfx.com/
1•malavika_manoj•1h ago•1 comments

Systemic Racism and Memetics

https://medium.com/luminasticity/on-systemic-racism-f708ac2efe51
2•bryanrasmussen•1h ago•0 comments

Spent $510 on cursor in the last 30d – AMA

1•xucian•1h ago•1 comments

Globe TV: Free Live TV Worldwide

https://globetv.app/
1•thunderbong•1h ago•0 comments

Why Deep Learning Works Unreasonably Well

https://www.youtube.com/watch?v=qx7hirqgfuU
1•phildawes•1h ago•0 comments

CMakeDependencyDiagram – Interactive target dependency visualization for CMake

https://github.com/renn0xtek9/CMakeDependencyDiagram
1•renn0xtek9•1h ago•1 comments