frontpage.
newsnewestaskshowjobs

Open Source @Github

fp.

Open in hackernews

Ask HN: Thoughts on AMD Ryzen AI MAX+ 395 for Local AI?

2•mempirate•4h ago
Wondering if anyone has played around with the Strix Halo based systems with 128GB of unified memory for local AI, and what their experience has been so far in running different models: - Practical limit on number of parameters x quantization - What you've been using it for - Which models and inference engines combine well

Comments

Iolaum•4h ago
Own a framework desktop 128gb.

GPU bandwidth limits the amount of tokens/sec that you get.

I 've mostly been running QWEN-3.6-35B-A3B at Q8 and QWEN-3.5-122B-A10B at Q4 with Q8 kv-cache on llama.cpp (use vulkan version). Also use MTP GGUF's you get ~ 50% speed up when predicting 2 or 3 tokens.

Haven't tested the new laguna S 2.1 yet.

Another thing I liked was that I had enough ram to also run an embedding open model for LLM Wiki apps.

Overall I think it's cheaper and better (intelligence wise) to get a $20 codex sub and always run gpt-5.6-luna than getting a desktop for local inference. Unless you have explicit needs for it (whether it is just experimentation or you have personal things you 'd rather keep on your computers or w/e). OpenCode go at $10 is also a great option.

mempirate•1h ago
Thanks for the detailed response. My goal is to deploy small to medium-sized LLMs on it for things that I'd rather keep private and local, i.e. automation around smart home, but also tracking finances and portfolio, aggregating & summarizing news feeds and blog posts and so on. For these purposes, medium intelligence at low TPS seems good enough to me.

But a large part of it is also wanting to learn more about serving these models, and potentially some small-scale experiments with training and fine tuning.

Ask HN: How do you use Copilot in the CLI, App, and IDE?

2•EspressoGPT•45m ago•0 comments

Why I prefer Opus 5 to Fable 5

13•novlrdotcom•3h ago•5 comments

Tell HN: Our paid Claude AI subscription unavailable >1 week and no support

36•KellyCriterion•6h ago•16 comments

Can anti-fraud make large-scale attacks unprofitable?

3•jezzwar•3h ago•1 comments

Ask HN: How to rewrite `Claude.md` and install the skill for Opus5 and Fable5

5•hyhmrright•7h ago•5 comments

Ask HN: Thoughts on AMD Ryzen AI MAX+ 395 for Local AI?

2•mempirate•4h ago•2 comments

Ask HN: Is there legal risk in AI memory?

2•sarjann•4h ago•0 comments

BMWs shows in-car ads for Spiderman

33•bigmattystyles•19h ago•15 comments

Ask HN: What do you use local models for?

3•tjomk•6h ago•1 comments

Internet is no longer accessible?

37•cute_boi•12h ago•23 comments

Ask HN: How to deal with security implications of running/installing projects?

12•johng•17h ago•10 comments

Goodbye POSIX, Hello Tea Room: Inside MinervaOS Architecture

10•xentnex•22h ago•6 comments

Claude's code comments – too much or just enough?

8•lalaleslieeeee•11h ago•5 comments

Ask HN: What apps are you building?

13•totaldude87•11h ago•15 comments

Ask HN: Why is every company encorporating AI everywhere?

24•vanessa1211•1d ago•38 comments

Ask HN: What are the most promising RL fields for a new master student?

34•zecice•2d ago•14 comments

Claude Code getting "API Error: 529 Overloaded"

5•croemer•1d ago•2 comments

Messaging platform I can self-host for my agents to communicate

2•giribisthebest•16h ago•3 comments

Ask HN: How do you figure out what you're passionate about?

11•charliebwrites•1d ago•20 comments

Ask HN: Any New Computer Ideas?

8•robalni•1d ago•13 comments

Ask HN: What livestream do you keep open in a tab?

53•missmoss•2d ago•18 comments

Tell HN: Namecheap gave my account to an unverified third party

495•Thrashed•4d ago•177 comments

Opus Cost = Minimum Wage in South Africa

4•bridgettegraham•1d ago•2 comments

Ask HN: Where do those LLM-generated outreach emails come from?

4•nicbou•1d ago•2 comments

You've reached the end!