frontpage.
newsnewestaskshowjobs

Open Source @Github

fp.

Open in hackernews

Notes on gotchas while migrating 35kb preprompts from Opus to self-hosted Ollama

https://patrickmccanna.net/notes-on-migrating-large-prompts-away-from-anthropic-openai-to-self-hosted-llms/
22•0o_MrPatrick_o0•1h ago

Comments

SyneRyder•13m ago
TLDR: Local models have a smaller context window, so your 35kB prompts that worked fine against a hosted 1 Million token window, crash out when you only have a 65K (!) token window locally.

I dislike being negative, but I was really hoping for more substance when reading this. It would have been an interesting topic.

0o_MrPatrick_o0•6m ago
Thanks for the feedback. I wanted to get into more detail, but I spent the whole weekend working these problems and then constructing this post.

Dario’s behavior this weekend made me feel like this just needed to get out quick. In the future, I’ll be sharing more details about some other things in the process and some ways I found to use automation to accelerate splitting prompts for use on local inference.

monegator•6m ago
Yes, please!
dell2024•4m ago
I had hoped to get some new information out of this topic, but unfortunately found the same local "dead-ends" that I explored myself.

It unfortunately feels like we will be stuck waiting for a burst bubble before local hardware can be reasonably acquired for personal LLM usage.

DiabloD3•3m ago
The article doesn't really describe the problem: if your prompt is 35kb, your prompt is confusing, unfocused, and doesn't work right on any LLM, and is needlessly bloating your context.

At this point in time, due to how most people and companies run their inference engine, regardless of the model (yes, this includes the newest from OpenAI and Anthropic and the Chinese Tigers and Dragons), you run out of useful context that the model can accurately attend to around the 250k mark no matter how much they advertise their context size is.

You need to cut your prompt up. If you believe LLMs work, have the LLM help you shape the overall plan, and then have multiple sessions run each step in the plan without being bloated with the context of previous successful steps.

I don't see LLMs being production-ready until the context rot and sampling problem is fixed forever. This has not occurred, and the big inference providers aren't even bothering to integrate any of the research on that subject.

If anything, many of the bigger companies are actively making inference quality worse just to extend their runway a tiny bit farther before they go bankrupt.

The only thing the article gets right is this: if you're serious about LLMs, abandon Big AI and infer locally only. This is the only way you have control over the quality of the output.

How to Write an Effective Software Design Document

https://refactoringenglish.com/excerpts/write-an-effective-design-doc/
118•fagnerbrack•2h ago•11 comments

Notes on gotchas while migrating 35kb preprompts from Opus to self-hosted Ollama

https://patrickmccanna.net/notes-on-migrating-large-prompts-away-from-anthropic-openai-to-self-ho...
22•0o_MrPatrick_o0•1h ago•5 comments

The AI job market in 2026

https://www.ilinmaks.com/blog/en/ai-jobs-market-2026
29•dxs•1h ago•1 comments

Graphic Rants: Nanite Tessellation

http://graphicrants.blogspot.com/2026/02/nanite-tessellation.html
9•ibobev•44m ago•1 comments

New $100K H-1B Visa Fee Pushes Tech Jobs Offshore

https://spectrum.ieee.org/h-1b-visa-us-government
87•rbanffy•1h ago•101 comments

RubyGems Open Source Supply Chain Security and OpenAI

https://rietta.com/blog/rubygems-supply-chain-openai/
12•rietta•32m ago•3 comments

Where has Construction Automation been successful?

https://www.construction-physics.com/p/where-has-construction-automation
14•ltononro•57m ago•2 comments

EuroBirdPortal – Live bird movements across Europe

https://www.eurobirdportal.org/ebp/en/
184•NKosmatos•6h ago•52 comments

Jabber/XMPP: How Do We Gain Traction?

https://gultsch.de/posts/how-do-we-gain-traction/
13•inputmice•33m ago•5 comments

A 386 PC for Your RP2350

https://github.com/rh1tech/frank-386
149•SamuraiLion•6h ago•42 comments

An atlas of periodic solutions to the three-body problem

https://www.threebodyorbits.com/
179•danielmorozoff•2d ago•38 comments

Devil's Arrows: Ancient builders hauled 55k-lb stones 11 miles for UK stone row

https://www.sciencedaily.com/releases/2026/09/260909005152.htm
26•bookofjoe•3d ago•17 comments

Apple's Siri AI Can Be Swapped Out for Claude, ChatGPT, Code Shows

https://www.macrumors.com/2026/09/14/siri-can-be-swapped-out-for-chatgpt-claude/
158•tosh•3h ago•80 comments

Understanding FlashAttention Pt 1: Personal Notes

https://chizkidd.github.io//2026/09/13/flashattention/
8•ibobev•1h ago•0 comments

Largest known Roman mosaic, beneath Baths of Trajan, opens to the public

https://www.theartnewspaper.com/2026/09/14/largest-roman-mosaic-opens-to-the-public
16•bookofjoe•2h ago•3 comments

Unsolved Problem by Fields Medalist Breached by Two High School Students

https://www.htx.com/en-in/news/internet-shocked-unsolved-problem-by-fields-medalist-breache-IBDgZ...
19•soltanov•2h ago•13 comments

SDR–; open source SDR with a patchable signal graph, Rust DSP, web UI

https://github.com/Newspicel/sdrminusminus
19•newspicel•2h ago•1 comments

Show HN: Pelican-bicycle alternatives (updated for 2026)

https://gally.net/temp/20260914pelican-alternatives/index.html
21•tkgally•1h ago•2 comments

Texas judge rules TikTok misled users on child safety feature

https://www.reuters.com/legal/litigation/texas-judge-rules-tiktok-misled-users-child-safety-featu...
88•1vuio0pswjnm7•2h ago•28 comments

Show HN: Kinesis – Control your Mac with the Meta Neural Band

https://github.com/callbacked/kinesis
81•callbacked•3h ago•26 comments

Ubuntu 26.10 completes transition to Rust-based coreutils

https://www.omgubuntu.co.uk/2026/09/ubuntu-2610-rust-coreutils-complete
49•theanonymousone•1h ago•30 comments

Trying to Make a Loop Auto-Vectorize

https://jsgroth.dev/blog/posts/trying-to-make-a-loop-auto-vectorize/
4•zdw•4d ago•0 comments

OpenArch – PyTorch implementations of modern LLM architectures

https://github.com/anuj0456/OpenArch
102•anuj0456•7h ago•23 comments

Spaceships (Reverse Asteroid)

https://spaceships.treybastian.com/
320•zdw•4d ago•64 comments

Apple's Dimensional Drawings

https://developer.apple.com/accessories/dimensional-drawings/
317•herbertl•15h ago•107 comments

Fable 5.1 Solves the Cyphral Distich, a 370-year-old cipher

https://www.vals.ai/blogs/fable-solves-cyphral-distich
1129•u1hcw9nx•18h ago•522 comments

Temporal raises $550M at a $12.55B valuation

https://temporal.io/blog/temporal-raises-usd550m-series-e-at-usd12-55b-valuation-ai
51•gk1•1h ago•32 comments

Optimizing Anamorphic Sculptures

https://tncardoso.com/blog/2026/09/optimizing-anamorphic-sculptures/
9•zbsc•3d ago•0 comments

The case against JPEG XL

https://giannirosato.com/blog/post/case-against-jxl/
232•contact9879•14h ago•297 comments

XCancel suspended "due to a new development in the ongoing legal proceedings"

https://xcancel.com/twitter
240•unfocso•3h ago•196 comments