frontpage.
newsnewestaskshowjobs

Open Source @Github

fp.

The Missing Piece in Rust Error Handling

https://mcmah309.github.io/posts/the-missing-piece-in-rust-error-handling/
1•lukastyrychtr•1m ago•0 comments

Enzo's – AI home search that shows cap rate and flags overpriced homes

https://enzos.ai/
1•nicomontuschi•1m ago•0 comments

Is Anthropic a Good Business?

https://www.polymathinvestor.com/p/is-anthropic-a-good-business
1•rwmj•3m ago•0 comments

FDA may allow some toxic chemicals to be added to food without safety review

https://www.theguardian.com/us-news/2026/oct/10/fda-toxic-chemicals-food-analysis
1•NewJazz•4m ago•0 comments

I tested whether AI agents obey text you never see. 4 of 8 did, every time

https://smallprint.dev/blog/one-sentence-four-of-eight-agents
1•nickcosta•4m ago•0 comments

Show HN: Lunatic Fling – a cel-shaded 3D remake of After Dark's Lunatic Fringe

https://lunaticfling.com/
1•hoag•8m ago•0 comments

Safe SIMD in Rust, even on the inside

https://shnatsel.github.io/safe-simd-in-rust-even-on-the-inside/
1•ravenical•8m ago•0 comments

Ask HN: A way for HN to verify "founders" and give them extra voting power?

1•amichail•9m ago•0 comments

The idea that temperature sets when we sleep

https://www.bbc.com/future/article/20261007-how-does-temperature-affect-your-sleep
1•Noaidi•10m ago•0 comments

Ask HN: Can we expect LLM-based coding agents to become noticeably faster?

1•lysace•12m ago•0 comments

Show HN: Peer-to-peer instant messaging with proof-of-work spam protection

https://github.com/iljah/p2pIM
2•iljah•13m ago•0 comments

My personal AI agent posted my bank details on company Slack

https://www.businessinsider.com/personal-ai-agent-grok-bot-posted-bank-details-company-slack-2026-10
3•bhrlady•18m ago•3 comments

AI Is Supercharging a Scam Economy Bigger Than the Cocaine Trade

https://www.bloomberg.com/graphics/2026-ai-supercharges-scam-economy/
1•toomuchtodo•20m ago•1 comments

Neanderthal wooden tools from Spain found preserved in stone

https://arstechnica.com/science/2026/10/neanderthal-wooden-tools-from-spain-found-preserved-in-st...
2•Brajeshwar•21m ago•0 comments

Are we dead moths in a puddle

https://jessicar.substack.com/p/are-we-dead-moths-in-a-puddle
1•airhangerf15•21m ago•0 comments

Logs as UX

https://www.jvt.me/posts/2026/10/09/logs-interface/
1•Brajeshwar•21m ago•0 comments

FlowCE: Calculus on your TI-84 Plus CE

https://tommyb3ll.github.io/FlowCE/
1•jgx0•21m ago•0 comments

Google expands SynthID Detector for AI content

https://blog.google/innovation-and-ai/models-and-research/google-deepmind/synth-id-ai-content/
1•gmays•21m ago•0 comments

The Human Soul as a Manifestation of Quantum-Like Fields

https://www.researchgate.net/publication/374749062_The_Human_Soul_as_a_Manifestation_of_Quantum-L...
2•olvy0•29m ago•0 comments

Post Capitalism

https://geohot.github.io//blog/jekyll/update/2026/09/29/post-capitalism.html
2•7777777phil•29m ago•0 comments

Radiolab: Math vs Machine

https://radiolab.org/podcast/math-vs-machine
3•kohsuke•31m ago•1 comments

Bitwarden Dual License Model

https://community.bitwarden.com/t/published-version-update-in-app-stores/102750
38•Cider9986•32m ago•9 comments

Agents and the Illusion of Productivity

https://opeonikute.dev/posts/agents-and-the-illusion-of-productivity
2•porridgeraisin•32m ago•1 comments

Show HN: Simple Security Checker

https://github.com/yuyalapis/security-checker
1•rozenapp•34m ago•0 comments

Is This Machine Playing?

https://machineplaying.org/
1•vinhnx•36m ago•1 comments

World Models for Biomedicine

https://www.cell.com/cell/fulltext/S0092-8674(26)01005-6
1•brandonb•37m ago•0 comments

Unikernels were hard. key word: were

https://ghuntley.com/unikernels/
2•ghuntley•37m ago•0 comments

FBI Arrests Executive at Ransomware Negotiation Firm

https://krebsonsecurity.com/2026/10/fbi-arrests-founder-of-ransomware-negotiation-firm/
5•rafaelc•38m ago•0 comments

Valar Atomics and Day One Ventures

https://twitter.com/isaiah_p_taylor/status/2108679133467414855
1•babelfish•38m ago•0 comments

Nicolas Flamel and the Philosophers' Stone

https://www.bbc.com/future/article/20261008-the-true-story-of-alchemist-nicholas-flamel-and-the-p...
1•tosh•40m ago•0 comments
Open in hackernews

Show HN: Tessary – Find the AI agent failures your sampled evals miss

https://github.com/tessaryai/tessary
1•infinitetrooper•55m ago
Hey, I'm Akhil, co-founder of Tessary. An open-source agent reliability platform that monitors every production trace, detects issues that sampled evals miss, and investigates to find the root cause.

We started up about a year ago building synthetic users for usability testing of B2B software, we built the agents that would personify real users and use the software but, we were never able to make them work reliably enough for long running sessions. We had built a massive eval set to tune our agents but, then the evaluation itself got prohibitively costly - like 5x costlier than actually running the agent since we needed to grade across the many narrow intents of our agent. That's where Tessary originated, we wanted to make agent reliability both cover more ground and be simultaneously less costly to run.

Tessary works on 2 layers.

L1 classifiers - which look at every agent trace and flag potentially erroneous ones extremely cheaply. We are talking 2-3 orders of magnitude cheaper than running the actual agent.

L2 agents - which triage and find the root cause of the issue from the erroneous traces. These are run on SOTA models but, because we run them on already flagged traces, the time and $ spent finding the cause is much cheaper than running grading on even 1% of traces. In our testing on complex agents, at 1M traces, this approach was 5x cheaper than sampled evaluations and as your agent gets more reliable, the cost of reliability also goes down.

we are open source and can be self-hosted so, you don't need to send your traces to us. We support existing `gen_ai` spec and ingestion over OTLP so, you can add this as a new exporter to your existing setup. We have a cloud version, you can quickly try the features with a $10 credit.

We are looking to build more classifiers and enable people to build custom classifiers for their needs and would love to hear feedback on common failure modes people have built solutions or evaluation for

Repo: https://github.com/tessaryai/tessary Website: https://tessary.ai Try for free: https://app.tessary.ai