frontpage.
newsnewestaskshowjobs

Made with ♥ by @iamnishanth

Open Source @Github

fp.

Show HN: Textclip.sh

https://textclip.sh
1•everlier•37s ago•0 comments

EMF Disrupt Mitochondrial Function and Trigger Cellular Dehydration

https://peteranthonycowan.substack.com/p/how-low-level-electromagnetic-fields
1•nateb2022•51s ago•0 comments

What Happens When We Insist on Optimizing Fun?

https://www.bloomberg.com/news/features/2025-12-26/how-ai-is-changing-the-games-we-play-from-poke...
1•safe_duck7727•1m ago•0 comments

Show HN: Local Code reviews for introverts

https://github.com/turingmindai/turingmind-code-review
1•vinkupa•2m ago•0 comments

Existential Risk and Growth [pdf]

https://philiptrammell.com/static/Existential_Risk_and_Growth.pdf
1•gmays•3m ago•0 comments

Claude Shannon Watch

https://www.danecjensen.com/2026/01/08/claude-shannon-watch.html
2•danecjensen•4m ago•0 comments

Lucid, Nuro, and Uber Unveil Global Robotaxi at CES, Announce On-Road Testing

https://investor.uber.com/news-events/news/press-release-details/2026/Lucid-Nuro-and-Uber-Unveil-...
1•gfortaine•4m ago•0 comments

African Union backs campaign to replace Mercator map (2025)

https://www.npr.org/2025/08/21/nx-s1-5508358/mercator-map-africa
1•divbzero•5m ago•0 comments

What if Linux ran Windows and meant it? Meet Loss32

https://www.theregister.com/2026/01/06/loss32_crazy_or_inspired/
1•losgehts•6m ago•0 comments

Point of no returns: researchers crossing a threshold in the fight for funding

https://www.nature.com/articles/d41586-025-04060-x
2•occamschainsaw•7m ago•0 comments

Can AI do your job? See the results from hundreds of tests

https://www.washingtonpost.com/technology/interactive/2026/ai-jobs-automation
1•pseudolus•9m ago•1 comments

Show HN: ShutterSnap – uses AI to find and clean up Photos.app similar photos

https://shutterslim.com/
1•wheels•10m ago•0 comments

Jellyfish and sea anemones rely on sleep to repair DNA damage in their neurons

https://www.nature.com/articles/s41467-025-67400-5
1•tguvot•10m ago•0 comments

French Court Orders Google DNS to Block Pirate Sites

https://torrentfreak.com/french-court-orders-google-dns-to-block-pirate-sites-dismisses-cloudflar...
2•HotGarbage•12m ago•0 comments

Musk's Grok AI Generated Undressed Images per Hour on X

https://www.bloomberg.com/news/articles/2026-01-07/musk-s-grok-ai-generated-thousands-of-undresse...
3•erhuve•12m ago•1 comments

Claude Code 2.1: The Pain Points? Fixed

https://paddo.dev/blog/claude-code-21-pain-points-addressed/
1•m4r71n•12m ago•0 comments

Impedance Matching

https://www.futilitycloset.com/2025/09/19/impedance-matching/
1•surprisetalk•13m ago•0 comments

Show HN: Wozz – open-source Kubernetes cost linter and cluster auditor

1•wozzio•13m ago•0 comments

MMAP Advent of Code 2025

https://mmapped.blog/posts/46-aoc-2025
2•surprisetalk•13m ago•0 comments

It's the Same Great Taste

https://brooke.substack.com/p/its-the-same-great-taste
1•surprisetalk•13m ago•0 comments

Compiler-Based Context- and Locking-Analysis for the Linux Kernel

https://www.phoronix.com/news/Linux-Compiler-Locking-Analysis
1•mco•15m ago•0 comments

Among the Agents

https://www.hyperdimensional.co/p/among-the-agents
1•jger15•15m ago•0 comments

Show HN: gsw – Switch Google Cloud configurations with session isolation

https://github.com/sakebook/gsw
1•sakebook•15m ago•0 comments

Exploiting deobfuscation in ImunifyAV for code execution (CVE-2025-65530)

https://blog.popovs.lv/imunifyav-code-execution/
1•aleksejs•16m ago•0 comments

Show HN: AI Vibe Coding Hackathon

https://vibe.devpost.com
1•abdibrokhim•17m ago•0 comments

Show HN: Interactive Periodic Table with Experiments

https://grundamnen.nu
1•ACS_Solver•17m ago•0 comments

List of Active Usenet Groups

https://newsgrouper.org/alt.fan.usenet/1767338943/1767497394
2•mghackerlady•18m ago•0 comments

Language Models as One-Time Teacher for Hierarchical Planning in Text Environs

https://arxiv.org/abs/2512.09897
2•JnBrymn•19m ago•0 comments

Show HN: 90% of GPU Cycles Are Waste. A New Computing Primitive for Physics AI

https://github.com/isaac-sim/IsaacSim/discussions/394
1•ZuoCen_Liu•21m ago•1 comments

Qwen3-VL-Embedding and Qwen3-VL-Reranker: Next Gen of Multimodal Retrieval

https://qwen.ai/blog?id=qwen3-vl-embedding
2•pretext•22m ago•0 comments
Open in hackernews

Show HN: KeelTest – AI-driven VS Code unit test generator with bug discovery

https://keelcode.dev/keeltest
28•bulba4aur•1d ago
I built this because Cursor, Claude Code and other agentic AI tools kept giving me tests that looked fine but failed when I ran them. Or worse - I'd ask the agent to run them and it would start looping: fix tests, those fail, then it starts "fixing" my code so tests pass, or just deletes assertions so they "pass".

Out of that frustration I built KeelTest - a VS Code extension that generates pytest tests and executes them, got hooked and decided to push this project forward... When tests fail, it tries to figure out why:

- Generation error: Attemps to fix it automatically, then tries again

- Bug in your source code: flags it and explains what's wrong

How it works:

- Static analysis to map dependencies, patterns, services to mock.

- Generate a plan for each function and what edge cases to cover

- Generate those tests

- Execute in "sandbox"

- Self-heal failures or flag source bugs

Python + pytest only for now. Alpha stage - not all codebases work reliably. But testing on personal projects and a few production apps at work, it's been consistently decent. Works best on simpler applications, sometimes glitches on monorepos setups. Supports Poetry/UV/plain pip setups.

Install from VS Code marketplace: https://marketplace.visualstudio.com/items?itemName=KeelCode...

More detailed writeup how it works: https://keelcode.dev/blog/introducing-keeltest

Free tier is 7 tests files/month (current limit is <=300 source LOC). To make it easier to try without signing up, giving away a few API keys (they have shared ~30 test files generation quota):

KEY-1: tgai_jHOEgOfpMJ_mrtNgSQ6iKKKXFm1RQ7FJOkI0a7LJiWg

KEY-2: tgai_NlSZN-4yRYZ15g5SAbDb0V0DRMfVw-bcEIOuzbycip0

KEY-3: tgai_kiiSIikrBZothZYqQ76V6zNbb2Qv-o6qiZjYZjeaczc

KEY-4: tgai_JBfSV_4w-87bZHpJYX0zLQ8kJfFrzas4dzj0vu31K5E

Would love your honest feedback where this could go next, and on which setups it failed, how it failed, it has quite verbose debug output at this stage!

Comments

ericyd•1d ago
I'd be curious to hear more about how it determines when a failure is a source code bug. In my experience it's very hard to encapsulate the "why" of a particular behavior in a way the agents will understand. How does this tool know that the test it wrote indicates an issue in the source vs an issue in the test?
bulba4aur•1d ago
Hey, thanks for the question.

So from my experience with the LLMs if you ask them directly "is this a bug or a feature" they might start hallucinating and assume stuff that isn't there.

I found in a few research/blog posts that if you ask the LLM to categorize (basically label) and provide score in which category this issue belongs it performs very very well.

So that's exactly what this tool does, when it sees the failing test it formulates the prompt in a following way:

## SOURCE CODE UNDER TEST: ## FAILED TEST CODE: ## PYTEST FAILURE FOR THIS TEST: ## PARSED FAILURE INFO: ## YOUR TASK: Perform a deep "Step-by-Step" analysis to determine if this failure is: 1. *hallucination*: The test expects behavior, parameters, or side effects that do NOT exist in the source code. 2. *source_bug*: The test is logically correct based on the requirements/signature, but the source code has a bug (e.g., missing await, wrong logic, typo). 3. *mock_issue*: The test is correct but the technical implementation of mocks (especially AsyncMock) is problematic. 4. *test_design_issue*: The test is too brittle, over-mocked, or has poor assertions.

Then it also assigns the "confidence" score to it's answer, based on that either full regeneration of the tests proceeds, commenting on the bug in the test, fixing mocks or full test redesign (if it's to brittle)

While this is not 100% bullet proof, i found this to be quite effective way - basically using LLM for the categorization.

Hope that answers your question!

bulba4aur•1d ago
To clarify, each failing test triggers "review" agent, to determine "why" the test fails, and again, it can be improved with better heuristics probably, more in depth static analysis than the source code, but it is how it works in the current version.
arthurstarlake•20h ago
i wonder if always having a design doc of some substance discussing the intended behavior of the whole app would help reduce instances of hallucination. The human developer should create it and let it be accessed by the AI
bulba4aur•20h ago
100% agree with that
hrimfaxi•1d ago
How exactly do credits work? Your pricing mentions files and functions but doesn't appear to give a true unit of measure.
bulba4aur•1d ago
Hey, thanks for the feedback, i will make sure to make it more visible/less confusing. So the model is actually quite simple.

1 credit - 1 file up to 15 functions. <-- only this tier is available in alpha, due to current limitations in the implementation, i tried generating on bigger files and it took quite a long time, so i am in the workings on solving this issue before enabling larger files support.

2 credits - 1 file up to 30 functions. 3 credits - 1 file 30-35 functions.

P.s if generated tests have <70% pass rate (at which point probably something went horribly wrong, your credits are refunded)

Hope this answer clears things up!

joshuaisaact•1d ago
I notice one of the things you don't really talk about in the blog post (or if you did, I missed it) is unnecessary tests, which is one of the key problems LLMs have with test writing.

In my experience, if you just ask an LLM to write tests, it'll write you a ton of boilerplate happy path tests that aren't wrong, per se, they're just pointless (one fun one in react is 'the component renders').

How do you plan to handle this?

bulba4aur•1d ago
I actually though about it multiple times over at this point.

You're right, this deserves more attention, and is a valid problem going forward with this app. And I had this problem when just started building, it either generated XSS tests for any user input validation method (even if it used other validators) or just 1 single test case.

For now I attempt to strictly limit the amount of tests for LLM to generate.

This is achieved with "Planner" that plans the tests for each function before any generation happens, that agent is instructed to generate a plan that follows the criteria:

- testCases.category MUST be one of "happy_path" | "edge_case" | "error_handling" | "boundary".

And it is asked to generate 2-3 tests for each category. While this may result in the unnecessary tests, it at least tries to limit the amount of them.

Going forward I believe the best approach would be to tune and tweak the requirements based on the language/framework it detects.

observationist•1d ago
Do a structured code review, with a few passes by Claude or Codex. Have it provide an annotated justification for each test, and flag tests with redundant, low, or no utility within the context of the rest of the tests. Anything that looks questionable to you, call it out on the next pass, and if it's not justified by the time you fully understand the tests, nuke it.

You could automate this, but you'll end up getting rid of useful tests and keeping weird useless ones until the AI gets better at nuance and large codebases.

OptionOfT•21h ago
What I see a lot is a generated test for something I prompt, and the test passes. Then I manually break the test and it fails for a different reason, not what I wanted to verify.

Guess I need to make it generate negative tests?

aleksiy123•18h ago
The automated version of this is mutation testing.

Which is actually probably a solid idea for this exact use case.

rcarmo•22h ago
Weird. Copilot knows what tests are and only "fixes" them after we've refactored the relevant code.

I really wonder if Claude Code and other agents keep track of these dependencies at all (I know that VS Code exposes its internal testing tools to agents, and use Anthropic and OpenAI tools with them).

bulba4aur•5h ago
Indeed, the Microsoft Copilot eco-system might be a bit more sophisticated these days.

It so just happens than people around me, including myself, don't use the copilot, we "left" for the next big thing when Cursor was release, and copilot was still a glorified auto-complete.

From your feedback it seems like they became quite good?