frontpage.
newsnewestaskshowjobs

Open Source @Github

fp.

Open in hackernews

Show HN: Trunchbull, run real models against any benchmark in your browser

https://trunchbull.dev
2•pepperpoppins•59m ago
Hi HN,

Today I'm showcasing Trunchbull, a benchmarking platform designed for authoring benchmarks and running them against different models. We have direct support for benchmarks that use the harbor authoring system, custom tool authoring via the vercel ai sdk and configuration limits.

We've also already imported terminalbench 2.0, as a sort of proof of concept that our harbor task orchestrator works, although you currently need a paid account as we are provisioning sandbox environments.

I've made several popular benchmarks publicly available for testing. You dont need an account or your credit card information to access these:

[Public benchmark demos](https://trunchbull.dev/sandboxes)

[GSM8K math reasoning](https://trunchbull.dev/try/gsm8k)

[SkateBench](https://trunchbull.dev/try/skatebench)

[ARC-Challenge](https://trunchbull.dev/try/arc-challenge)

[TruthfulQA](https://trunchbull.dev/try/truthfulqa-mc1)

[Medical AI Failure Atlas](https://trunchbull.dev/try/medical-ai-failure-atlas)

These demos let you pick from a preselected list of models, and will systematically test the selected models against their case scenarios.

Any and all feedback is welcome, but i'm particularly interested in: - knowing what kind of benchmark evidence youd like to expect - any improvements on our benchmark run page, anything that can provide clarity or better understanding of the benchmark u just ran. - what you'd prefer to see on the overview page. - improvements on our documentation - better configuration and spend limits.

C++26: Std:Indirect

https://www.sandordargo.com/blog/2026/08/12/cpp26-indirect
1•signa11•18s ago•0 comments

Decrypting my Flume water monitor's MQTT traffic with a passive relay

https://lithostech.com/2026/08/decrypting-flume-water-monitor-traffic/
1•stevecrozz•1m ago•0 comments

CodeRabbit raises a $143M Series C at a $1.5B valuation

https://www.coderabbit.ai/blog/introducing-agentic-change-management
1•juanpflores•1m ago•0 comments

We tested our new Voice Isolation across 10 STT engines. Word errors dropped 46%

https://krisp.ai/blog/voice-isolation-2-5/
1•davitb•2m ago•0 comments

Different SAST/DAST scanners only share %1.8 of common findings on same source

https://blog.audn.ai/posts/audn-vs-codex-vs-aikido-juice-shop-comparison
1•boratac•2m ago•0 comments

DeepSeek-V4-Pro-0813

1•uneven9434•3m ago•0 comments

Music notes are no longer the bottleneck (ca. 1600)

https://bsky.app/profile/jeff.auriemma.xyz/post/3msvhsitaf22m
1•jdauriemma•3m ago•0 comments

DosTips, Windows scripting knowledge trove, scraped to death

https://www.dostips.com/
1•iqml2568•3m ago•0 comments

RL Environments Explained: How AI Agents Learn Real-World Work [video]

https://www.youtube.com/watch?v=a00xIn5kwhM
1•ronfriedhaber•3m ago•0 comments

There is no escaping the solar system, not at speed anyway

https://disassociated.com/no-escaping-solar-system-not-speed/
1•Brajeshwar•4m ago•0 comments

Grok 4.6 (High) Intelligence, Performance and Price Analysis

https://artificialanalysis.ai/models/grok-4-6
2•theanonymousone•7m ago•0 comments

Footage Shows Dolphins Hunting with Shells Off the East Coast of Australia

https://www.sciencealert.com/dolphins-filmed-hunting-with-shells-and-possibly-passing-the-skill-t...
1•thunderbong•7m ago•0 comments

Where Has All the Testosterone Gone?

https://www.theatlantic.com/family/2026/08/testosterone-decline-theories/688247/
2•fortran77•7m ago•0 comments

Show HN: My pocket-sized Personal IoT device. Data-local by default

3•AJTSheppard•8m ago•0 comments

DeepSeek V4 Pro 0813

https://openrouter.ai/deepseek/deepseek-v4-pro-0813
4•explosion-s•8m ago•1 comments

IO Factory: Simulating AI-Enabled Influence Campaigns at Scale [pdf]

https://arxiv.org/abs/2608.10920
1•andrybak•9m ago•0 comments

YouTuber Hank Green says his AI usage is 'not healthy'

https://techcrunch.com/2026/08/01/youtuber-hank-green-says-his-ai-usage-is-not-healthy/
2•abnercoimbre•9m ago•2 comments

Show HN: Aphorio – 3,250,000 WebSocket sends in 0.002 seconds (Ryzen 9 9900X)

https://github.com/prettydiff/aphorio
1•austin-cheney•9m ago•0 comments

Math identity that applied at model-load time, makes MLA decode at 71,000 tok/s

https://efficientagent.substack.com/p/the-mla-decode-speedup-hiding-in
1•bnayak25•11m ago•0 comments

Putting Claude's Watermarking to the Test

https://dreasays.substack.com/p/putting-claudes-watermarking-to-the
1•gojkoa•11m ago•0 comments

Archer OS: draft spec for AI agents operating apps under OS authority

https://github.com/coachpato/archer-os
2•Billie_Archer•14m ago•0 comments

Remind HN: Perseid meteor shower peaks tonight

1•cjbarber•14m ago•0 comments

NIST RFI for modernizing the NVD in the wake of AI

https://www.federalregister.gov/documents/2026/08/12/2026-16371/request-for-information-rfi-on-mo...
2•MattSayar•14m ago•0 comments

Show HN: /show-me: agent skill for compact visual representations

https://www.humanlayer.com/blog/show-me-skill
1•dhorthy•15m ago•0 comments

Local LLM Hardware Calc

https://dubir.net/tools/local-llm-hardware-calculator/
1•delneg•15m ago•0 comments

QuestDB 10.0

https://questdb.com/blog/questdb-10-release/
1•tosh•16m ago•0 comments

Water Groups Push Washington for Cyber Rules After Hacking Spree

https://www.wsj.com/pro/cybersecurity/water-groups-push-washington-for-cyber-rules-after-hacking-...
1•toomuchtodo•17m ago•1 comments

SpaceXAI: Grok 4.6

https://openrouter.ai/x-ai/grok-4.6
3•theanonymousone•18m ago•0 comments

Samsung with Claude: "Design and verification shortened to 2 days from a month"

https://biz.chosun.com/it-science/ict/2026/08/12/XIEQWWZCDRFH7BJV5Z3DOY2RLQ/
1•TMWNN•19m ago•1 comments

Introducing Grok 4.6

https://cursor.com/blog/grok-4-6
4•stevefan1999•20m ago•0 comments