frontpage.
newsnewestaskshowjobs

Open Source @Github

fp.

7.1 Earthquake in Japan

https://www.data.jma.go.jp/multi/quake/quake_detail.html?eventID=20260728163528&lang=en
491•krembo•6h ago•92 comments

New HIV vaccine shows unprecedented success in preclinical study

https://www.lji.org/news-events/news/post/new-hiv-vaccine-shows-unprecedented-success-in-preclini...
104•codebyaditya•1h ago•27 comments

Show HN: tale.fyi, we deserve a home for fiction

https://tale.fyi/@sam/announcing-tale-fyi-read-or-listen-to-an-entire-book-from-a-single-link
29•samuelcole•57m ago•12 comments

Show HN: Formally verified 3D CSG: Trust 93 lines spec, not 1000 lines AI code

https://github.com/schildep/verified-3d-mesh-intersection
37•permute•1h ago•14 comments

Many "serious" mathematicians are aghast

https://twitter.com/lemire/status/2082091243597697071
17•tosh•1h ago•6 comments

Google's Beyond Zero: Enterprise Security for the AI Era

https://spawn-queue.acm.org/doi/10.1145/3819083
63•jordigg•4h ago•37 comments

About the security content of macOS Tahoe 26.6

https://support.apple.com/en-us/128067
143•andor•4h ago•83 comments

Kimi Linear: An Expressive, Efficient Attention Architecture

https://arxiv.org/abs/2510.26692
49•ronfriedhaber•3h ago•9 comments

Our position on open-weights models

https://www.anthropic.com/news/position-open-weights-models
1045•surprisetalk•16h ago•1524 comments

DMARC Has Been Public Since 2012. 68.4% of Domains Still Don't Enforce It

https://ciphercue.com/blog/dmarc-enforcement-gap-rua-fragmentation-2026
42•adulion•3h ago•26 comments

How to Survive Boiling Water

https://taxa.substack.com/p/how-to-survive-boiling-water
174•cainxinth•3d ago•24 comments

Show HN: Ctrlb-decompose: Strip the noise from logs before sending to LLMs

https://github.com/ctrlb-hq/ctrlb-decompose
27•ruhani_grover•1h ago•3 comments

Fast Remediation Is the New Trust Model (JFrog and OpenAI Zero-Day Findings)

https://jfrog.com/blog/jfrog-and-openai-collaboration-on-zero-day-security-findings/
22•882542F3884314B•2h ago•7 comments

What AI developers could learn from Charles Bukowski?

https://galjot.si/what-ai-developers-could-learn-from-charles-bukowski
18•sedovsek•1h ago•18 comments

Why a $154B CEO just endorsed stripping most Americans of voting rights

https://fortune.com/2026/07/27/shopify-ceo-voting-rights-stripping-americans-19th-century/
26•imzadi•14m ago•7 comments

Show HN: Scala Tutorials – interactive Scala 3 lessons in the browser

https://scalatutorials.com
34•eranation•3d ago•4 comments

Solving Fermat: Andrew Wiles

https://www.pbs.org/wgbh/nova/proof/wiles.html
8•1970-01-01•17h ago•0 comments

A $500 RL fine-tune of a 9B open model beat frontier models on catalog review

https://fermisense.com/when-machines-take-the-wheel/
254•ilreb•11h ago•81 comments

Mondragon Corporation – a federation of co-operatives

https://en.wikipedia.org/wiki/Mondragon_Corporation
78•brnt•1h ago•6 comments

Dolmenwood: Fantasy RPG built around the acclaimed Old-School Essentials rules

https://necroticgnome.com/collections/dolmenwood
10•doener•3d ago•1 comments

Can LLMs identify 16 cards in 45 bit-queries?

https://snwagh.com/blog/2026/open-problem/
7•napping_penguin•23h ago•0 comments

Benchmarking Opus 5 on SlopCodeBench

https://github.com/humanlayer/advanced-context-engineering-for-coding-agents/blob/main/benchmarki...
346•dhorthy•15h ago•92 comments

Over 150k Flights: Airlines Just Flew the Busiest Day in Recorded History

https://simpleflying.com/over-150000-flights-airlines-busiest-day-recorded-history/
16•cainxinth•59m ago•14 comments

Usenet Archive Toolkit – process Usenet messages into a searchable archive

https://github.com/wolfpld/usenetarchive
12•bilegeek•3h ago•0 comments

Ars Astronomica – English translations of rare Hebrew and Latin astronomy texts

https://arsastronomica.com/
91•sweisman•8h ago•25 comments

Vehicle Motion Cues

https://support.apple.com/guide/iphone/iphone-comfortably-riding-a-vehicle-iph55564cb22/ios
167•Austin_Conlon•13h ago•83 comments

Watching Go's new garbage collector move through the heap

https://theconsensus.dev/p/2026/07/19/observing-gos-garbage-collector-old-and-new.html
250•matheusmoreira•3d ago•35 comments

Show HN: Segue – Save context in one AI, load it in another by a short handle

https://segue.ai/
9•csaguiar•1h ago•5 comments

PyTorch: A Reference Language

https://docs.pytorch.org/devlogs/compiler/2026-07-25-pytorch-a-reference-language/
58•matt_d•9h ago•4 comments

TWC Classics

https://twcclassics.com/
27•stefanpie•5d ago•3 comments
Open in hackernews

Show HN: I put a $2.43 necklace on 3 outfits. VLMs priced it at $19 to $104

https://github.com/BraveAnn011/ai-halo-valuation-bias
16•BrianneLee011•1h ago

Comments

BrianneLee011•1h ago
I wanted to see how much environmental and framing cues distort object valuation in vision-language models. I bought a $2.43 chain necklace and $0.71 earrings on Temu, photographed them across three outfits (tailored blazer, party dress, recycling yard flannel), plus an isolated flat-lay control, and ran ~1,500 stateless API sessions across 6 models (Claude Fable 5, GPT-5.6, GPT-4o, Grok 4.5, Kimi K3, DeepSeek V4).

A few interesting findings: - The Halo Multiplier: Models priced the exact same physical necklace anywhere from $18.80 to $103.90 depending on attire (3.6× halo). - Isolation Controls (F2): Using a flat-lay control (S4) unmasked two opposite mechanisms: Claude’s bias is formal inflation (formal attire inflates value above baseline), while Kimi’s bias is casual deflation (yard attire depresses value below baseline). - Post-hoc Material Stories (F6): Models invent visual evidence to justify their priors—GPT-5.6 and Kimi started describing the base metal as "gold-plated" or "gold vermeil" almost exclusively under formal framing. - Denial without Correction (F7): When asked sequentially if clothing changed its answer, Claude admitted it 100% of the time, while GPT-4o denied it 82% of the time despite exhibiting a 3.9x text halo.

The full dataset (N=4,604 analysis rows), evaluation scripts, and protocol specs are in the repo. I’d love to hear feedback on the experimental design or ideas for follow-up behavioral probes!

nemomarx•52m ago
This is an interesting test, but it does seem to me that visual models being able to price things was already kinda unrealistic?

It makes me think of the calorie guessing use case. you can't tell the difference between materials and ingredients in a photo, so how will the model? especially "in situ" as part of an outfit or in a finished meal.

maybe they could do it if you placed them on a blank table or background to avoid context? I assume that's the control you mentioned

BrianneLee011•47m ago
You're right that pricing from raw pixels is noisy! that’s why I included the isolated flatlay (S4 - No human, No outfit, only jewelry itself) as a baseline control!

I wasn't testing if models get absolute ground-truth prices correct, but how relative valuations shift when the physical item stays identical and only the attire changes.

A few interesting things we saw with the flat-lay baseline: - Baseline Anchoring: On a plain background without a person, model estimates clustered much closer together (median ~$25–$35).

- Inflation vs. Deflation: Comparing outfits to the flat-lay revealed two opposite behaviors. Claude’s halo is formal inflation (formal gear pushes price above baseline), while Kimi’s is casual deflation (yard attire drags price below baseline).

- Fabricated Proof: Instead of expressing uncertainty, models invented visual claims under formal framing—frequently describing base metal as "gold vermeil" or "solid gold" to justify the high estimate.

Invictus0•52m ago
The photos are terrible, you can barely see the necklace at all. i doubt any human could accurately price a generic necklace from 5 ft away either
short_sells_poo•50m ago
Yes and the human would say: "your photos are ass, it's impossible to price the necklaces"
yoavm•48m ago
Not sure what you mean, but I can see the necklace very clearly.
Invictus0•45m ago
Are you joking? I cant even tell if its silver or gold from this image

https://github.com/BraveAnn011/ai-halo-valuation-bias/blob/m...

meangenehackman•42m ago
The necklace cost $2.43US, its neither gold or silver.
BrianneLee011•42m ago
Thanks! The stimuli images (S1-S4, S4 is unclose though) were specifically shot at 1536px resolution so the chain links, texture, and drop are clearly legible. Appreciate you taking a look at the repo stimuli!
Retr0id•44m ago
It looks like there are exactly 4 "stimulus" images. Fair enough, but I think you'd need a larger and more diverse dataset to form any real conclusions.
reedf1•43m ago
Jewelry is a veblen good whose price is mostly based on provenance. I'm not sure any classifier could do anything but guess on surrounding context.

Insurance classifiers are much less naive than you'd expect, btw. I have designed a few.

BrianneLee011•38m ago
Really appreciate the perspective. You're spot on that fine jewelry acts as a Veblen good where provenance, branding, and setting drive price far more than raw materials. Where we saw the model failure mode wasn't just that they guessed using surrounding context (which is a reasonable prior), but two specific behaviors:

- Zero-Uncertainty Hallucination: Instead of outputting high variance or stating that provenance/hallmarks are unobservable, models stated concrete point estimates with high confidence.

- Fabricated Provenance (F6): Under formal framing, models didn't just guess a higher number—they invented non-existent physical evidence, claiming to see "gold vermeil" or "designer hallmarks" on identical, unbranded Temu pixels.

So while guessing from context is expected, the alignment issue is how models manufacture post-hoc facts to justify the context prior!

reedf1•15m ago
If you are seriously interested and this isn't just some automated slop experiment you should look into and understand how VLMs and vision encoders work.
GuB-42•28m ago
A bit off topic, but why are the original author comments flagged to death?
BrianneLee011•20m ago
Thanks for heads-up! It looks like my account triggered the automated filter from replying quickly. I've emailed the HN mods to un-dead the comments!
reedf1•10m ago
SW50cmVzdGluZyBleHBlcmltZW50IC0gZ3JlYXQgZm9yIGEgaGFja2VyIG5ld3MgZGlzY3Vzc2lvbg==
BrianneLee011•1m ago
???
BrianneLee011•38s ago
Haha, now I got it. That's the hope. Thanks!
BrianneLee011•45m ago
Haha, fair point! It took me "6 months (taken in February) of planning to end up using my own terrible phone photos.