frontpage.
newsnewestaskshowjobs

Open Source @Github

fp.

Open in hackernews

Show HN: Assay – a QA CLI that finds bugs with no LLM and no tests written

https://github.com/awss1i/assay
1•awss1i•1h ago
I've always found it odd there was no deterministic QA tools with zero baseline needed, especially after so long. It's always been either (1) playwright or cypress that require written tests, dependant on the tests YOU write, or (2), asking the LLM to find and test bugs in your code, which has 2 cons, its costly (sometimes), and its non-deterministic.

That's why i built assay. Its only dependencies is playwright & chromium, and you point it at any folder with a webpage and it runs tests on its own by driving the webpage itself in a browser and clicking every control.

It then flags anywhere the webpage contradicted/disagreed with itself; it doesn't know right from wrong NOR what the webpage is about, but universally, a page disagreeing with itself is almost always a bug.

An example of this is on a paint app, if you draw two strokes on a canvas and click undo twice, you expect them both to undo each stroke sequentially; in one of 225 benchmarks, clicking the undo the first time didn't remove the first stroke, and only the 2nd click worked. assay drove that and caught it.

As for benchmarks, there are 2 variants that are all reproducible in the repo. The first is 225 generated web pages that were first checked by hand, then had assay run on them. There was a total of 20 bugs across all programs, and assay caught 15 with the right reasoning, and had 0 false failures. The second was a harder one where an LLM planted 5 bugs across 10 working original webpages. assay found 12 of the 50 and had 0 false failures, which may seem low, but assays superpower is that its cheap and quick.

assays median runtime is 14s, and it plugs into 15 agentic coding harnesses with a skill, plus claude code and deepseek harness have their own plugin that automatically runs it with a stop hook every time the LLM touches a webpage.

My favourite feature is that it groups flagged tests together if their of the same bug. More details are in the repo about this since this is getting a little long, but the links it makes have its own benchmarks, 51/51 links it got right with 0 false links. This serves as additional context to you, if you use it as a raw cli, or to your agent if your using it with an LLM harness.

pip install assay-ui. All alternative methods to install are on the github, including harness support.

Show HN: I'm a dermatologist and I vibe coded a 3D biophysical skin model

https://www.drmagnuslynch.com/skin
1•sungam•40s ago•0 comments

EnigmaForge – an LLM benchmark where the question is hidden in the story

https://arxiv.org/abs/2609.30144
1•robottwo•1m ago•0 comments

Buying a New Computer (1993)

https://archive.org/details/BuyingaN
1•doener•1m ago•0 comments

It Was the Harness, Not the Model

https://blog.herlein.com/post/harness-not-model/
2•aray07•3m ago•0 comments

How Pew Research Center is – and is not – using AI in our work

https://www.pewresearch.org/decoded/2026/09/28/how-pew-research-center-is-and-is-not-using-ai-in-...
3•hn_acker•5m ago•0 comments

The Untold Origins of Trump's Plan to Sharply Restrict Mail-In Voting

https://www.propublica.org/article/trumps-plan-to-restrict-mail-voting
3•hn_acker•5m ago•0 comments

Parasocial media: why influencers aren't your friend

https://www.baldurbjarnason.com/2026/04-parasocial-relationships/
1•PaulHoule•7m ago•0 comments

Stanford Encyclopedia of Philosophy

https://plato.stanford.edu/
1•akeck•8m ago•0 comments

Dodge's Fire-Preventing Ejecting Battery Patent

https://www.jalopnik.com/2269595/dodge-fire-preventing-ejecting-battery-patent/
2•cf100clunk•8m ago•0 comments

Manus 2.0

https://manus.im/blog/introducing-manus-2-0
3•canberkys•8m ago•0 comments

Ask HN: Finetuning strategies

1•moka_labs•9m ago•1 comments

Small Decisions: Engineering a Leading Model

https://brooker.co.za/blog/2026/09/28/engineering-system-one.html
2•emschwartz•9m ago•0 comments

The Download: rogue agent liability and the AI Hype Index

https://www.technologyreview.com/2026/09/28/1145202/the-download-rogue-agent-liability-and-the-ai...
2•joozio•11m ago•0 comments

Singer: The Downfall of a Great American Manufacturer

https://www.worseonpurpose.com/p/who-makes-singer-sewing-machines
2•cainxinth•12m ago•0 comments

Firefox 157

https://www.neowin.net/news/firefoxs-biggest-update-in-years-launches-tomorrow-heres-how-to-get-i...
3•bundie•13m ago•0 comments

How private equity is killing public access to hospitals and emergency care

https://www.washingtonpost.com/ripple/2026/09/25/private-equity-hospital-closures/
3•toomuchtodo•13m ago•2 comments

Fifty Years of Semicolons

https://jordanzimmerman.com/fifty-years-of-semicolons.html
2•adrianfcole•14m ago•0 comments

Driver Ticketed for No Insurance Just Because Flock (YC 2017) Said She Didn't

https://www.techdirt.com/2026/09/28/driver-ticketed-for-no-insurance-despite-having-insurance-jus...
2•HotGarbage•15m ago•0 comments

The Artificial Analysis Cyber Index

https://artificialanalysis.ai/articles/artificial-analysis-cyber-index
1•AnodicElegy•16m ago•0 comments

Florida asks court to bar OpenAI from new models as part of child harm lawsuit

https://www.reuters.com/world/florida-asks-court-bar-openai-developing-new-models-part-child-harm...
2•theanonymousone•17m ago•0 comments

Synthetic Sagas

https://www.scattered-thoughts.net/writing/synthetic-sagas/
1•ffin•17m ago•0 comments

We Can't Let Enormous Weirdos Regulate AI

https://little-flying-robots.ghost.io/why-we-cant-let-enormous-weirdos-regulate-ai/
1•SLHamlet•17m ago•0 comments

AI tutoring outperforms in-class active learning

https://www.nature.com/articles/s41598-025-97652-6
1•theanonymousone•19m ago•0 comments

An AlphaGo Moment for Inference?

https://int21.ai/insights/ai-generated-inference-engines/
1•antinucleon•22m ago•0 comments

How SLR cameras work: Nikon F3 (2023) [video]

https://www.youtube.com/watch?v=0AA-Le7cmH0
1•nayuki•22m ago•1 comments

Watch an AI agent try to prove the Riemann hypothesis on the cheap

https://www.zeyaddeeb.com/experiments/proofs
2•zdeeb•24m ago•0 comments

Voice Agents Can Just Do Things: Why voice is the next capability overhang

https://www.ignorance.ai/p/voice-agents-can-just-do-things
1•swolpers•25m ago•0 comments

Extrinsic World Modeling with Opus, Astra and Grok

https://all3d.ai/research/grok-spatial-reasoning
3•KaiserPister•25m ago•1 comments

UDP Broadcasting and the Brave New World of IPv6

https://hackaday.com/2026/09/24/udp-broadcasting-and-the-brave-new-world-of-ipv6/
1•topham•25m ago•1 comments

Using multiple Git remotes for true distributed version control

https://optimizedbyotto.com/post/multiple-git-remotes/
2•edward•25m ago•1 comments