frontpage.
newsnewestaskshowjobs

Open Source @Github

fp.

Ask HN: Are there AI models for generating sounds based on a text and reference?

22•onemiketwelve•1d ago•11 comments

Ask HN: Why is Ask HN only showing me 14 posts?

39•Gooblebrai•4h ago•35 comments

Ask HN: What do you think about Fractional CFOs

2•mtmosestn•38m ago•0 comments

Ask HN: What do you run on a $5 VPS that's worth keeping online 24/7?

9•mariocesar•2h ago•4 comments

Ask HN: What games do you religiously play on your phone browser?

6•saimiam•3h ago•8 comments

Ask HN: Do you still own a printer?

9•nunorbatista•8h ago•12 comments

Ask HN: How to deal with AI "true believer" leadership at work

5•microflash•9h ago•3 comments

Ask HN: How do you describe what your relationship brings you?

2•maxignol•6h ago•3 comments

Ask HN: What would you like to see in a new AI technology release?

2•ikishade•8h ago•2 comments

ChatGPT self-distances when admitting fault

3•chrisjj•3h ago•1 comments

Ask HN: Vibecoding Follies

4•vegnus•9h ago•2 comments

Ask HN: How come everyone is an LLM expert?

3•delis-thumbs-7e•10h ago•3 comments

Tell HN: GitHub refuses to remove cracked copies of my software after a month

54•IvanK_net•7h ago•52 comments

Ask HN: Show me your agentic SDE semantics

2•danielovichdk•12h ago•0 comments

Ask HN: Is observability broken for you?

2•tmach32•13h ago•0 comments

Ask HN: My Apple ID is locked for a week, Apple is a single point of failure

11•akg_67•14h ago•3 comments

Ask HN: Agent access to chat data should require participant consent?

2•maxwellito•16h ago•0 comments

Ask HN: Alternatives to NeurIPS, ICLR and ICML

3•john-titor•19h ago•0 comments

Ask HN: What brings you back to personal AI agents like Instinct and Muse?

3•sdrth•21h ago•1 comments

Orcah Studio: A local-first video agent that can search your videos

8•iliashad•1d ago•5 comments

Ask HN: Why does Astra compact context so frequently vs. Fable?

2•yesitcan•23h ago•0 comments

Ask HN: What do you think about AI generated slides for conferences?

4•Hixon10•23h ago•1 comments

Ask HN: How are AI budgets changing in your company?

5•bobby-cb•1d ago•2 comments

Tell HN: 2026 is the year of the Linux desktop, agent-adjusted

3•dvrp•1d ago•2 comments

Tell HN: Uceprotect is extorting website owners

101•goldenmember•1d ago•59 comments

Our digital privacy is being aggressively disbanded this week

12•Steaglsz•1d ago•1 comments

DOS Game Stunts Port to Linux,Windows,Browser,etc.

4•LowLevelMahn•2d ago•1 comments

Who is cleaning up all the garbage LLMs generate?

6•kbrannigan•2d ago•7 comments

Ask HN: Anyone accepted into OpenAI "Codex for open source"?

4•awb•1d ago•0 comments

Tell HN: OVH price increase for dedicated servers

11•esher•4d ago•6 comments
Open in hackernews

Ask HN: Best on device LLM tooling for PDFs?

4•martinald•1y ago
I've got very used to using the "big" LLMs for analysing PDFs

Now llama.cpp has vision support; I tried out PDFs with it locally (via LM Studio) but the results weren't as good as I hoped for. One time it insisted it couldn't do "OCR", but gave me an example of what the data _could_ look like - which was the data.

The other major problem is sometimes PDFs are actually made up of images; and it got super confused on those as well.

Given this is so new I'm struggling to find any tools which make this easier.

Comments

raymond_goo•1y ago
Try something like this

  !pip install pytesseract pdf2image pillow
  !apt install poppler-utils
  #!apt install tesseract-ocr
  from pdf2image import convert_from_path
  import pytesseract

  pages = convert_from_path('k.pdf', dpi=300)

  all_text = ""
  for page_num, img in enumerate(pages, start=1):
      text = pytesseract.image_to_string(img)
      all_text += f"\n--- Page {page_num} ---\n{text}"

  print(all_text)
constantinum•1y ago
give https://pg.llmwhisperer.unstract.com/ a try