frontpage.
newsnewestaskshowjobs

Open Source @Github

fp.

Is it just me, or has Claude Opus gotten worse recently?

4•de6u99er•1h ago•7 comments

Ask HN: What would happen if your company stopped using all AI tomorrow?

33•jc_811•20h ago•65 comments

Dropbox Data Breach

28•hmate9•14h ago•5 comments

Claude Code now appends a link to a Claude session in every commit

9•codexon•12h ago•1 comments

Tell HN: Remember to Change Your Filters

3•surprisetalk•8h ago•2 comments

Old.reddit.com no longer works for logged out users

25•kradeelav•11h ago•8 comments

Hotline Telegram Support and Sales Hub (hotline.tg)

2•hotline•8h ago•0 comments

Ask HN: What do your job interviews look like?

26•Hixon10•1d ago•8 comments

Tell HN: PayPal blocks GrapheneOS

514•leumon•5d ago•325 comments

Ask HN: Why is Founder Mode not working for Airbnb?

6•jorisboris•1d ago•7 comments

Claude 20x usage is only for the 5 hour window, not for the weekly limit

15•vmg12•20h ago•3 comments

Native Apps Why?

6•dmvjs•22h ago•15 comments

Ask HN: Which self-hosted Docker UIs support rootless mode?

3•vsilent•22h ago•0 comments

Ask HN: What is the next Burning Man in NA?

3•johnnyApplePRNG•1d ago•6 comments

Ask HN: What to do when a vendor doesn't respond to security issues?

3•cudder•1d ago•2 comments

Ask HN: Are Prompt Injections "Malware"?

2•razorbeamz•1d ago•9 comments

Ask HN: What are your biggest problems and fixes with multisession engineering?

5•top_rooster•1d ago•3 comments

Which AI Do You Think Will Have the Greatest Impact on the World?

5•SudilaDasun•1d ago•3 comments

You Are the Harness

6•learningstud•1d ago•0 comments

Is a One Person Company (OPC) just another type of Uber driver?

6•Bobby_Liu•1d ago•10 comments

Ask HN: AI for Home Lab Infra?

2•voakbasda•1d ago•3 comments

You've reached the end!

Open in hackernews

Ask HN: Best on device LLM tooling for PDFs?

4•martinald•1y ago
I've got very used to using the "big" LLMs for analysing PDFs

Now llama.cpp has vision support; I tried out PDFs with it locally (via LM Studio) but the results weren't as good as I hoped for. One time it insisted it couldn't do "OCR", but gave me an example of what the data _could_ look like - which was the data.

The other major problem is sometimes PDFs are actually made up of images; and it got super confused on those as well.

Given this is so new I'm struggling to find any tools which make this easier.

Comments

raymond_goo•1y ago
Try something like this

  !pip install pytesseract pdf2image pillow
  !apt install poppler-utils
  #!apt install tesseract-ocr
  from pdf2image import convert_from_path
  import pytesseract

  pages = convert_from_path('k.pdf', dpi=300)

  all_text = ""
  for page_num, img in enumerate(pages, start=1):
      text = pytesseract.image_to_string(img)
      all_text += f"\n--- Page {page_num} ---\n{text}"

  print(all_text)
constantinum•1y ago
give https://pg.llmwhisperer.unstract.com/ a try