frontpage.
newsnewestaskshowjobs

Open Source @Github

fp.

I'm a CTO. We took PE money 6 months ago. AMA

4•scottndecker•1h ago•8 comments

Ask HN: What Did Anthropic Bill Me For?

3•OhMeadhbh•1h ago•1 comments

Ask HN: Good large format (>20 inches) touchscreen E-Paper display options?

24•foota•11h ago•7 comments

Ask HN: Should GitHub downtimes impact considering Golang for a new project?

2•gchamonlive•1h ago•0 comments

Ask HN: Those making $500/month on side projects in 2026 – Show and tell

65•kaladan•1d ago•83 comments

How I Read Books

12•nomilk•17h ago•10 comments

A technology that counts reps and scores muscle failure from the Apple Watch

3•baraa_bilal•21h ago•0 comments

Talk Like Claude Day

19•KenPainter•1d ago•6 comments

Ask HN: What AI companies provide human support?

2•kbrannigan•21h ago•2 comments

Tell HN: HN Algolia Search not showing new results from the past 24 hours

4•infinitifall•21h ago•1 comments

I write my own Markdown

2•KenPainter•22h ago•0 comments

Do teams still use scrum and sprints?

5•peter_retief•23h ago•12 comments

Ask HN: Trust Between Agent's Companies

2•mathieu_aithos•23h ago•0 comments

Ask HN: Is there any way to use workflow to control Harness?

2•juntz•23h ago•4 comments

Ask HN: Why is my innovative IDEA is always behind the others?

3•juntz•1d ago•5 comments

You Don't Need AI to Generate Code

4•Lozybug•1d ago•4 comments

Harpoon for Sublime Text 4

3•tomoekn•1d ago•0 comments

Ask HN: Is Codex's 5-hour usage limit back?

3•linzhangrun•1d ago•2 comments

Ask HN: Claude Code randomly burning through 10% of the weekly limit?

6•terabytest•1d ago•2 comments

I've tested some local LLMs on prosumer hardware, here are some findings

6•felineflock•1d ago•0 comments

I built a tool to generate Cornell notes from YouTube videos

6•cristyg0101•2d ago•4 comments

Ask HN: Why do corporate failures always seem to punish the wrong people?

120•mittermayr•1d ago•121 comments

Ask HN: Anyone set up ways to easily obtain and read transcripts from Ted, YT?

5•MollyRealized•2d ago•9 comments

You've reached the end!

Open in hackernews

Ask HN: Best on device LLM tooling for PDFs?

4•martinald•1y ago
I've got very used to using the "big" LLMs for analysing PDFs

Now llama.cpp has vision support; I tried out PDFs with it locally (via LM Studio) but the results weren't as good as I hoped for. One time it insisted it couldn't do "OCR", but gave me an example of what the data _could_ look like - which was the data.

The other major problem is sometimes PDFs are actually made up of images; and it got super confused on those as well.

Given this is so new I'm struggling to find any tools which make this easier.

Comments

raymond_goo•1y ago
Try something like this

  !pip install pytesseract pdf2image pillow
  !apt install poppler-utils
  #!apt install tesseract-ocr
  from pdf2image import convert_from_path
  import pytesseract

  pages = convert_from_path('k.pdf', dpi=300)

  all_text = ""
  for page_num, img in enumerate(pages, start=1):
      text = pytesseract.image_to_string(img)
      all_text += f"\n--- Page {page_num} ---\n{text}"

  print(all_text)
constantinum•1y ago
give https://pg.llmwhisperer.unstract.com/ a try