When PaddleOCR-VL-1.6 dropped, independent benchmarks put it at the top of document parsing models. I had to try it. I needed a provider, but there simply isn't one ready for production that I would trust.
So i set one up myself. I assumed that even after getting it running, serving a vision-language model would be expensive.
It turns out the opposite is true. Once I had it running properly, the cost was absurdly low. At proper GPU utilization, the cost is only around $1 per 1,000 pages.
The nearest competitors are either much lower quality (Azure Read) or absurdly expensive (Extend or Reducto). Even Mistral OCR 4 which is really good and pretty cheap is still 4x more expensive.
So I had to share it. I made the endpoitns public, and vibe-coded a simple dashboard.
Let me know what you think!
c4kar•1h ago