frontpage.
newsnewestaskshowjobs

Open Source @Github

fp.

DeepSeek Harness

https://github.com/deepseek-ai/deepseek-harness
101•bjin•1h ago•41 comments

Choosing an AI model: one prompt, 11 models, different results

https://www.netlify.com/blog/one-prompt-11-models-very-different-results/
58•toddmorey•1h ago•33 comments

Deutsche Bank becomes first foreign yuan clearing bank in Europe

https://tradersunion.com/news/central-banks/show/2973571-deutsche-bank-becomes/
191•Markoff•2h ago•206 comments

My Rules for Using Spreadsheets

https://leancrew.com/all-this/2026/08/my-rules-for-using-spreadsheets/
18•surprisetalk•1h ago•15 comments

Picking berries is my meditation

https://www.tsoon.com/posts/picking-berries-meditation/
62•mooreds•4d ago•41 comments

ChatGPT Desktop (Codex Desktop) for Linux

https://openai.com/codex/
312•allanrbo•9h ago•212 comments

Better Gaussian Splatting in Julia

https://pxl-th.github.io/blog/better-gs-julia/
37•pxl-th•3d ago•5 comments

Show HN: MCP Memory – Fast Agent Memory Using Google's OKF and SQLite FTS5

https://github.com/fellowgeek/mcp-memory
4•pcbmaker20•17m ago•0 comments

The lattice of sets of natural numbers is rich (2021)

https://jdh.hamkins.org/the-lattice-of-sets-of-natural-numbers-is-rich/
72•benmandrew•3d ago•11 comments

ATG (YC F25) Is Hiring Member of Technical Staff (Data Platform)

https://atg.science/careers
1•dkobran•2h ago

Show HN: I told Claude Code never to reveal my secrets. It sent 3 of 4 anyway

https://github.com/crp4222/PrivAiTe
6•crp4222•19m ago•0 comments

DeepSeek V4 Pro 0813

https://openrouter.ai/deepseek/deepseek-v4-pro-0813
993•explosion-s•22h ago•426 comments

Tracking down the 16-year-old WAL-reset SQLite bug

https://tailscale.com/blog/sqlite-wal-reset-bug
1125•ropbear•23h ago•212 comments

I Built a 500k-Domain Search Engine for Makers in a Weekend for $10

https://alexmorleyfinch.github.io/marlin/history/v1/article/the_birth.html
3•dreamforever•39m ago•1 comments

Qwen3.8-2.4T

https://huggingface.co/Qwen/Qwen3.8-2.4T-A95B
685•Philpax•23h ago•159 comments

Delta

https://zed.dev/blog/introducing-delta
618•khy•19h ago•224 comments

Principia Mathematica is modern and insightful

https://okmij.org/ftp/Computation/Impressions/PrincipiaMathematica.html
229•matt_d•14h ago•102 comments

DeepSeek API Pricing Update

https://twitter.com/deepseek_ai/status/2087864589895798968
31•mfiguiere•1h ago•4 comments

What Garbage Collection Costs

https://shivanshuag.com/blog/what-garbage-collection-actually-costs/
6•shivanshuag•3d ago•3 comments

uBlock Origin Is Giving Up the Fight to Keep Ads Off Facebook

https://digitalescapetools.com/2026/08/ublock-origin-stops-chasing-facebook-ads.html
634•Markoff•1d ago•761 comments

Come for Eniac, Stay for Univac and Skeduflo

https://uniqueatpenn.wordpress.com/2026/08/05/come-for-eniac-stay-for-univac-and-skeduflo/
17•cainxinth•2d ago•2 comments

Antiqua–Fraktur dispute

https://en.wikipedia.org/wiki/Antiqua%E2%80%93Fraktur_dispute
128•buzzy_hacker•3d ago•54 comments

Mushroom behind 'tiny people' hallucinations identified

https://phys.org/news/2026-08-qa-mushroom-tiny-people-hallucinations.html
230•wglb•5d ago•190 comments

2026 Eclipse Webcams

https://jonty.github.io/2026_eclipse_webcams/
502•zoenolan•1d ago•135 comments

Flutter 3.47

https://flutter.dev/blog/whats-new-in-flutter-3-47
163•gumby271•14h ago•167 comments

PBS loses 50TB of data; 70 years of TV history after cloud vendor defunct

https://www.tomshardware.com/software/cloud-storage/nine-pbs-loses-access-to-70-years-of-data-aft...
18•vinayakborkar•1h ago•4 comments

Rapid warming may tip AMOC at 2°C, slower warming may avert collapse

https://phys.org/news/2026-08-rapid-atlantic-circulation-2c-slower.html
58•maxboone•3h ago•45 comments

The punched card tabulator

https://www.ibm.com/history/punched-card-tabulator
32•Bluestein•5d ago•12 comments

Grok 4.6

https://x.ai/news/grok-4-6
605•iLuddite•22h ago•555 comments

HTML over WebSockets: real-time SPAs with barely any JavaScript

https://en.andros.dev/blog/ef4968f5/html-over-websockets-real-time-spas-with-barely-any-javascript/
231•redbell•21h ago•169 comments
Open in hackernews

Choosing an AI model: one prompt, 11 models, different results

https://www.netlify.com/blog/one-prompt-11-models-very-different-results/
57•toddmorey•1h ago

Comments

isqueiros•48m ago
> Build a one-page site for a neighbourhood coffee shop: opening hours, the address, a short menu and a photo. Nothing on it changes unless I edit it myself.

If that's the entire prompt, it's quite depressing how much alike these all look. I appreciate some of the details from the Opus 5 version, but I can't help but strongly feel the AI vibes emanating from that design.

hombre_fatal•32m ago
Without more knowledge about the technical aspect, it could be a good thing that they're all so similar. If you tell humans to go from their kitchen to their bedroom, they all stand up and walk the same way. Nobody decides to crab walk. Maybe you'd have to cripple the model in some way to do that.

On the other hand, if you want something different with LLMs, all it takes is a few more words of creative flair in the prompt.

skinfaxi•16m ago
This is an interesting point. But even then everyone has their own style. Some might walk with a bit of a swagger, some with a limp, some might have to get in a wheelchair to go over. Would model temperature be another knob to turn beyond a more creative prompt?
mym1990•4m ago
The models are giving you very average, middle of the road responses. The average human gait doesn't have swagger, limps, or extensive accessibility needs. If you want the swagger, you need to steer it that way.
andy99•30m ago
It’s always been interesting to see how similar output is across ostensibly very different models. I remember testing short story writing in the early days and having the models all choose the same niche topic across e.g GPT, Llama, Phi, Claude, etc
tantalor•30m ago
It's also kind of pointless. Why does a coffee shop need a website anyway? Nobody is saying, "man I would love to go to this coffee shop but I can't find their website".

If all you want is "opening hours, the address, a short menu and a photo" there are easier and cheaper ways to do that.

miyoji•26m ago
> Nobody is saying, "man I would love to go to this coffee shop but I can't find their website".

Me, I'm saying that, and I've skipped going to coffee shops and restaurants because they don't have a website, just a fucking Facebook page. I don't use Meta products and can't see their page if I'm not logged into an account I don't have, so I do what the business owner intended: I go fuck myself and get coffee somewhere else.

altmanaltman•18m ago
But... why do you want to look up a coffee shop online before going there? Honestly have never heard anyone say this in my life before
skinfaxi•14m ago
To see the hours, get a sense of the menu. There are many matcha shops by me for instance and my wife likes to check the specials before choosing which one to go to.
altmanaltman•6m ago
Doesn't most of this show up on google anyway?
kifler•43m ago
I always enjoy these comparisons between models, especially when they demonstrate the actual costs in addition to the outputs.
arjie•42m ago
Building ad-hoc evals is trivial these days. And you can place an LLM judge in front to disambiguate between output. e.g. I have Frigate gating a video feed so that an LLM can watch our home cameras and the labeling task takes a few minutes and then the evals run at fairly low cost https://wiki.roshangeorge.dev/w/images/4/4e/Screenshot_-_Eva...

This means that generic benchmarks and evals are sort of passé. There's no need to have them write "make a coffee shop page" or whatever. Instead, simply focus on the real problem you have and evaluate whether they solve that problem. You want to move along the price-performance frontier alone and the SOTA models are decidedly at the top of the performance curve but very expensive so they serve very well to produce the gold standard and to judge.

The change from prior to today is that benchmaxxing and per-token pricing allowing for evaluation point to the same direction: do not proxy results. Instead deploy the highest end model you have as a judge for others until you have statistical confidence in discrimination and move along the task-specific frontier.

epolanski•19m ago
> Building ad-hoc evals is trivial these days.

I highly doubt that you have solved it. Writing proper evals and benchmarks for real-world scenarios is far from trivial.

Benchmarking an agent essentially means freezing, at the very minimum:

- the model

- the model's configuration (e.g. effort, permissions, provider)

- the dataset (e.g. a git repository at a specific sha)

- the code running the agent itself (you can build your own harness, trivial, but you still need to ship it as a single executable, froze in time. benchmarking against closed source runtime like claude code is quite useless, they change too frequently and in ways you cannot directly inspect).

- the tools at agent's disposal. Even a slightly different implementation of tool X (e.g. grep or readfile or sed) has an impact. In general this implies also freezing a very specific container image. In my personal benchmarks I provide a specific list of tools that come with the executable, there's no possibility of interacting with the outside world besides the provided apis, the agent bundles its own tools.

And even then: there's significant noise coming from the LLM providers themselves which noticeably change the models behaviour, I don't know whether this is because they optimize some settings or change the inference over time, etc.

And, last but not least, the output of LLMs is non deterministic.

Also, the LLM as judge presents essentially the same non-deterministic problems, has to be benchmarked itself thoroughly, and writing quality rubrics or "golden answers/outputs" is just difficult. One of the metrics I consistently try to emphasize is to avoid the "shotgun vomit dump" of information. So answers that get right to the point in plain terms avoiding dumps of information filled of jargon on top of the user are rated differently.

In short: its far from trivial to benchmark models on real-world agentic work taken from your personal or professional projects.

And even creating the test cases themselves is hard. No: you cannot take the output of some "sota" and use it as gold standard. This is a very crap approach. It's the sloppiest solution to the problem, in the very sense of slop: plausible, average, lacking any creativity or out of the box thinking, the things that make the real difference in complex software development.

The very point of creating these benchmarks is to find which configuration/model/tools/harness (skills/mcps/documentation/agents.md, etc) works better.

And it only works if you create these benchmarks yourself from genuinely difficult non-trivial work and find a solution that is better by most metrics implementation-wise, albeit you could settle on the implementation solving a series of cases and edge cases.

plumb_samji•41m ago
Interesting exploratory comparison, but I be cautious about treating it as a model benchmark With only three runs per model, the results are highly sensitive to randomness
neom•36m ago
I'd be curious to see Terra xhigh vs Sol low, only in that the visual languages are actually kinda different, that Terra work was a lot less "llm" feeling than a lot of the other designs, wonder if pumping up the effort would result in a more in-depth design but within that style.FWIW I enjoyed reading this way more than any usual benchmark posts we see.
jwr•35m ago
I've been doing a lot of benchmarking for a long time now with a number of local models for the purposes of spam filtering. The major observation is that there is a lot of variance in model performance. This should not be surprising, as these are probabilistic machines based on random numbers, so your performance will vary from run to run. But this also means that any sort of evaluation of benchmark with a sample size of 1 is essentially worthless for the purposes of model comparison.

In my benchmarks, I started insisting on having at least 5 runs.

This also makes me very suspicious of some of the posted benchmarks that compare models. If the results don't say how many benchmark runs were performed and what the variance was, there is really no basis for comparison. You are just guessing.

toddmorey•22m ago
Yeah the difference in token usage across difference models they found in the article was broader than I expected, but I was sort of floored by the variance in token usage for the same prompt (run multiple times) with the same model.
Systemerror7A69•32m ago
Am I wrong or are these evaluations, while interesting, not really meaningful for anyone doing serious development work?

I'm asking because I personally only use AI with specific and detailed instructions, building my projects piece-by-piece.

I mostly don't look at the low level code and some of it I don't understand as much as I'd like, but I very much give much more technical instructions than a simple, two sentence prompt.

So, it seems to me this "oneshot from a simple prompt" eval is fairly meaningless when it comes to model evaluation itself, as it is in no way representative of real world application.

This would be more something for "vibe coders", people with little to no programming background wanting a website?

217•27m ago
I'm repeatedly noticing that people working at big ai and tech companies are surprisingly not that... good... at using ai? It's like theyre doing a plausible thing to get something done and calling it a day
toddmorey•25m ago
As someone who's done web development for 20+ years, I find the model personalities pretty dang fascinating, especially how they develop (and evolve) design sensibilities.

I'm interested to know how the classic AI "purple preference" emerged (organically?) and if the beige wave came out of specific training efforts to combat it?

To your point on development work (the code itself), I was talking to some friends on the Google Chrome team about any research understanding the model's preferences around framework ergonomics and abilities to properly implement core web standards for given tasks. I think that would be super fascinating.

runtime_terror•20m ago
I'd assume it comes from Tailwind boilerplate/template sites as it seemed almost all of them were purple at the time
horsawlarway•30m ago
Approaching this from the perspective of a potential customer and not a designer - I find that I actually like the smaller/open model output quite a bit more.

Kimi, GLM, and Deepseek all absolutely run away with the "Can I quickly read the menu and find the address" challenge.

Most of the rest of the pages are stylistic, but hard to parse.

If I were in a car on a mobile phone trying to find the address of the place to meet a friend for coffee... I don't want a bunch of fluff and stylistic design that makes it hard to parse the information on the site.

jannishan•30m ago
I think we should establish a professional evaluation organization for models; otherwise, it will be difficult for informal evaluations to form standards.
pedrosbmartins•30m ago
Pretty interesting how DeepSeek V4 Flash 0731 has such disparate (and cheap!) results. I would never guess they come from the same model and prompt.
chrisjj•22m ago
> Vector graphics actually require a lot of work from the models

How so? Surely they can just steal such generic graphics off existing web sites.

dataviz1000•6m ago
I'm going to scream at the top of my lungs, 3 runs on a model doing a complex task is nowhere near enough to prove anything! Posts like this don't tell us much.

Sorry for linking these again, but they are a solid rebuttal to every "we compared 11 different models" post that reaches the top of Hacker News.

1. Massive variance on identical tasks: Sonnet ranges from 8,163 to 17,334 tokens solving the exact same lambda calculus problem over 5 runs [0]. Five runs isn't even enough for a full comparison, but it clearly demonstrates how wildly variable model output is. The flame graphs help visualize the thinking process to show why that variance happens.

2. With enough runs of the same prompt, models become extremely predictable. Although not deterministic, understand this will help us use the models more effectively. Scroll down to the blue "Walk it from the start" button on this page [1]. It maps the probability of solving a problem as it gets progressively more complex, helping visualize the probability distribution of reasoning models so it stops being a complete black box. (Note: this required thousands of runs and serious compute).

Single-prompt or 3-run "evals" are just noise without statistical rigor.

If you scroll a little further there is a clean comparison of capacity -- opposed to capability which is using SVG to paint a pelican on a bicycle -- of Qwen3-4B vs. Phi-4-reasoning over thousand of runs on the same 144 problems.

[0] https://adamsohn.com/lambda-variance/

[1] https://adamsohn.com/reasoning-grid/

seamlessdev•3m ago
Definitely. It was already a trend before LLMs exploded. And the whole "all websites look the same" has been a thing since at least Bootstrap times.
pistoriusp•23m ago
I checked out the development benchmarks for agents, and was suprized to find that they're not really representative of many of my own day-to-day development workflows.

It appears that they're mostly testing the ability to make business tasks autonomous, with ~20% associated to development tasks (ssh here, install this, etc.), but not actual programming.

epolanski•15m ago
You touch a point I quickly skimmed in another comment.

Yes, the most valuable benchmarks and evaluations you can write are those that resemble your work.

The evaluations are extremely hard to write and test.

And yes, virtually all benchmarks are E2E one shots, they do not reflect multi turn processes or how most people interact with LLMs.

Which is why every Opus after 4.6 looks better on benchmarks, but is hard to work with interactively.

matheusmoreira•13m ago
I've been using code review as my benchmark. Launched a massive parallel Fable/max code review on my lone lisp codebase and recorded all the results and transcripts. Switched to OpenAI and am now running the exact same code review with Sol/max.

It's still a work in progress but preliminary results reveal that Sol is able to reproduce 70%-90% of Fable's performance. This is a very meaningful result for me because code review is what I use AI for.