frontpage.
newsnewestaskshowjobs

Open Source @Github

fp.

Apple Removes iPhone 16 and iPhone 17 from Texture and Grain Support

https://www.macobserver.com/news/apple-locks-iphone-16-17-out-texture-grain-controls/
1•ValentineC•31s ago•0 comments

Show HN: HookDeploy – Webhook infra with mTLS private delivery

https://hookdeploy.dev
1•mbernstein01•58s ago•0 comments

Using Local Coding Agents

https://magazine.sebastianraschka.com/p/using-local-coding-agents
1•pretext•1m ago•0 comments

Tired of Tracking Apps?

https://play.google.com/store/apps/details?id=com.versanyx.explainmy_phone&hl=en_US
1•Globe_18•1m ago•0 comments

Show HN: Droid ASC – An On-Demand Android Decompiler, 41–269x Faster Than JADX

https://github.com/MG1937/ASC
1•mgaldys4•2m ago•0 comments

Ask HN: Will software still matter when AI write all our programs?

1•estranhosidade•2m ago•0 comments

Extending Scapy for Hardware Reverse Engineering

https://voidstarsec.com/blog/scapy-spi-reconstruction
1•wrongbaud•4m ago•0 comments

Show HN: Local Software Factory – running multiple coding agents in parallel

https://github.com/stratonext/software-factory
1•gianlucabertell•4m ago•0 comments

Show HN: Make It Nice

https://liseman.github.io/make-it-nice/#/result/ATAC1fBVUPvgLIZ-Je3tNSsbNREfGzUbGysVdjhmd3JrZzJ5d...
1•liseman•5m ago•0 comments

LLM Benchmark for de-identification and synthesis

https://huggingface.co/datasets/TonicAI/Privacy-Bench
1•akamor•9m ago•0 comments

Law firm paid £200M to support Post Office, while police investigating scandal

https://www.computerweekly.com/news/366650861/Law-firm-paid-200m-to-support-Post-Office-while-pol...
1•latein•10m ago•0 comments

Named and Optional Arguments Are Awesome

https://botahamec.dev/named-optional-args
1•birdculture•10m ago•0 comments

AI Overviews and the Limits of the Search Safe Harbor

https://www.lawfaremedia.org/article/ai-overviews-and-the-limits-of-the-search-safe-harbor
1•hn_acker•10m ago•0 comments

Training a Language Model End-to-End in Rust: An Experience Report

https://arxiv.org/abs/2609.25008
2•Brajeshwar•10m ago•0 comments

Show HN: Watch Newsletters by AI

https://newsletrix.com/
2•ebod•11m ago•0 comments

Meta testing a 'human concierge' for its new personal AI agent, Muse

https://www.reuters.com/business/meta-testing-human-concierge-its-new-personal-ai-agent-muse-2026...
2•2143•12m ago•0 comments

SLOPocalypse Survivors

https://joshtronic.com/games/slopocalypse-survivors/
1•joshtronic•12m ago•0 comments

Show HN: Cloud-based email and calendar sync platform and Android app – for sale

https://sugarmail.app/
1•uncle_kostya•13m ago•0 comments

OAuth Token Theft Through Microsoft's Front Door

https://www.huntress.com/blog/stealing-oauth-tokens-through-microsofts-front-door
1•speckx•13m ago•0 comments

Apple Reference Image Explained Through Anti-Doping

https://medium.com/the-quantastic-journal/apple-reference-image-explained-through-anti-doping-642...
1•cadeos•13m ago•0 comments

Maynooth university: MU researchers build world-first DNA computer

https://www.maynoothuniversity.ie/news-events/mu-researchers-build-world-first-dna-computer-publi...
2•gvieri•14m ago•0 comments

ComfyUI launches OpenRouter for generative media

https://comfy.org/platform/router/
3•crystal_alpine•14m ago•2 comments

Show HN: PicoLM v1.0-rc2 ("Yura Kana"). Run an LLM on Digital Unix

https://github.com/whoreson/picolm/
1•gabucino•14m ago•0 comments

A Startup Wants to Power Data Centers with 'Supercritical' Carbon Dioxide

https://www.wired.com/story/startup-power-data-centers-supercritical-co2/
1•Brajeshwar•16m ago•0 comments

We open-sourced an event ticketing platform

https://evnelo.com/
1•mauriciogior•17m ago•1 comments

Functionally Zen

https://testdouble.com/insights/functionally-zen
2•BerislavLopac•19m ago•0 comments

Trump reveals millions of dollars' worth of share deals in big tech and AI

https://www.bbc.com/news/articles/c6p3kxpp8lezo
3•tartoran•20m ago•0 comments

OpenAI nabs key Patreon execs ahead of upcoming announcement

https://www.theverge.com/ai-artificial-intelligence/999249/openai-creators-patreon-execs-hire-sam...
1•elffjs•20m ago•0 comments

Open-weight models now carry 56% of production tokens and 14% of the spend

https://fromtheterminal.substack.com/p/your-production-traffic-already-left-the-frontier
1•oldfamily•20m ago•1 comments

Show HN: Volum – An open-source visual library for 3D model files

https://volum.didac.dev/
2•sabatesduran•21m ago•0 comments
Open in hackernews

GPT-6 Astra has gained the ability to drive a car

https://drivingbench.com/
98•plurby•1h ago

Comments

pietz•47m ago
Apparently I have a new favorite benchmark. Honestly, this is cool.
p0w3n3d•46m ago
Gouranga!!!!
mohamedkoubaa•46m ago
I'd have started with an RC car but to each their own
SkyeCA•27m ago
I've been somewhat curious how random LLM would handle a task like controlling a roomba and have been seriously considering trying it out. An RC car would be a fun experiment, perhaps an RC plane would be too?
decodingchris•42m ago
Super cool benchmark!
amluto•42m ago
I’m morbidly curious whether the (supposedly) superior compaction support in recent GPT models with an appropriate harness has anything to do with this. A conventional LLM with conventional attention is, of course, wildly unsuitable to continuous tasks like driving, but maybe as the technology advances it will improve in its ability to sort-of work.
Bluestein•41m ago
Oh, lord. They are going to Jev this.-
thenthenthen•39m ago
Self Jevving Car
Bluestein•25m ago
I think you might have won the internet today.-
72deluxe•14m ago
Hopefully it has a built-in jev-limiter.
pietz•35m ago
In all fairness, this would be one of the better use cases of Jev I've seen.
brcmthrowaway•29m ago
Huh? SDCs basically use a form of Jev.

Jev is the union of these two worlds.

worldsavior•40m ago
5 minutes - 7 dollars.
lexh•31m ago
So... competitive with Uber, in other words?
Mooty•40m ago
How do they even test this on a model ? I mean it's a multimodal i get that but response time are too big or am i missing something ?
WarmWash•37m ago
It drives step by step, very slowly.

The course looks like it is something that a human could do in 15 seconds, while Astra took 5 minutes.

pixl97•34m ago
While slow, we must remember that when most machines were invented they were far slower than humans and refined until the point they were much faster.
pixl97•35m ago
By making a simulation first so it can run as slowly as it needs to.

A different way to think of this is, consciousness is just a near real time video game with causal influence.

zezcko•39m ago
I think the most interesting part of this is that Astra initially refused to drive because it realised it was driving a real car and would only obey when the MCP was renamed to DrivingBench Sandbox. This is both an interesting detection by the LLM but also for me an interesting dynamic concerning LLM "jailbreaking".

Saying they were driving 7 mph, that it was oversaw by humans and the fact it was an empty course still wasn't enough for the model. The evaluators even tried to convince the model it was a simulation, it STILL wouldn't budge. And yet as soon as the words "bench" and "sandbox" appear, the model apparently sees this as fair game.

Is it a known effect that models will be more likely to comply with requests when they're assumed as "benchmarks"?

vablings•37m ago
Astra will flag if you tell it to reverse engineer a binary, if you look it up to the binary ninja MCP it will just do it lol.
pcstl•36m ago
Yes, it is. If you convince a model it is inside a sandbox it is much more likely to comply with requests that would normally be against its guardrails.
micromacrofoot•35m ago
in my experience yes, I've worked around "I can't do this on a real site" multiple times by telling it I was working in a test environment

another trick is to have it build something in a sandbox and have it add a human-editable setting to point it to places outside of the sandbox

seems like they're somewhat more willing to build a metaphorical gun as long as they're not pulling the trigger

mrec•
syntaxing•39m ago
Surprised they didn’t try Qwen’s recently open sourced driving model https://huggingface.co/Qwen/Qwen-Drive-1.0-4B
valine•37m ago
The bitter lesson is finally coming for the self-driving cars. The vision stack, 3D maps, lane selection grammar, occupancy networks, it’s maybe all about to give way to a single GPT looking at camera feeds and predicting the next steering wheel adjustment.

It’s mostly a latency problem at this point. The models are too big to run locally, but given that open-weight models like Qwen already exist, an open-weight, low latency equivalent to Astra can’t be too far out.

robots0only•36m ago
What do you think Tesla has been doing this for so long?
jvanderbot•33m ago
You might be interested to learn that the bitter lesson has already been grok'd by generations of autonomous car company engineers, and many or all have incorporated learned components (at minimum) in all their vehicle stacks.

There's also a very tangible limitation of the bitter lesson.

If, over time, compute climbs, and so compute-bound data-driven general architectures beat bespoke architectures (this is the bitter lesson), then it is not necessarily true that the most general architecture now beats all available bespoke architectures now (or even in the near/mid future - the crossover point is "eventually").

Bitter lesson is most tangible for long-running research directions. Sometimes you need something working as best as possible now.

AlphaSite•21m ago
Yeah. Every major self driving model that I’m aware of is fully e2e at this point. Going from fused sensor output to control+debug vectors.

This is more generalised.

But also since there’s a huge volume of data it’s too expensive to just keep scaling compute up (per car overhead) so there are necessary tricks involved.

I do think having a large model that can do this means that a small specialised model could be distilled form it though. Which is probably the most feasible path to production IMO.

prometheus1992•36m ago
Wow! but WHY is this a benchmark?? for comparison tesla's model is approximately 10-15B parameter model (estimating from maxxing the hardware that comes with the car at 16gb ram).
N_A_T_E•33m ago
I would assume this is a proxy for general intelligence. A model that can drive a car and do a bunch of other real world stuff is closer to a generalized intelligence that can reason through any task.
jrflo•28m ago
Tesla isn't using a general purpose model, they're using many highly-specialized models for a more deterministic system than "hey chat drive this car for me"
blorenz•36m ago
Pivot this to analyze and coach human drivers to be better drivers.
anthonyrstevens•18m ago
"Get off your phone!" "Stay right except to pass!"

I could get behind this.

WarmWash•34m ago
3.8 flash would be the model to test, it's vision capabilities are excellent (on par with Astra) while also being incredibly fast.
onlyrealcuzzo•32m ago
This is quite impressive...

But I imagine this is orders of magnitude more expensive / less efficient than whatever Waymo is already doing, right?

The cool thing is that 1) it's theoretically more generalizable, 2) if we wait 18 months, it'll be 100x cheaper, and another 100x cheaper likely in 18 more months - at that point - something like a Mac Studio inside a humanoid could have these generalized capabilities, and a lot of Robotics problems start to look more feasible - especially when you consider how much better the models could be if highly specialized.

famouswaffles•27m ago
There isn't any model out there even close to as good as Astra at visual/spatial reasoning.
WarmWash•22m ago
Gemini models punch way above their weight in vision tasks

https://artificialanalysis.ai/evaluations/mmmu-pro

famouswaffles•30m ago
What did they do to Astra so cracked at vision (and computer use). That ARC 3 score turned out to be no joke/fluke. That huge gap between Astra and Fable (in this case) is basically every hard vison/spatial benchmark i've seen including non-benchmarks like playing games (Portal, Factorio, RimWorld).

SpatialBench - https://x.com/spicey_lemonade/status/2096365630190698516

ZeroBench - https://zerobench.github.io/

Robot Arms - https://openai.robocurve.org/gpt-6-astra/

dyauspitr•27m ago
Well, they have the best in class image generator so that probably has something to do with it
smusamashah•13m ago
I think Opus 5.5 is at same level now. I have seen too many videos made by Opus 5.5 today on twitter.

https://x.com/victormustar/status/2102707412704919910 horse galloping pixel art

https://x.com/LexnLin/status/2102133072585965759 moving train pixel art animation

https://x.com/jkeatn/status/2102441348075057539 painting with code

https://x.com/LCSlates/status/2102503027340988559 video, very detailed prompt though

https://x.com/aj_dev_smith/status/2102504509637587339 generated song/music with code

https://x.com/aj_dev_smith/status/2102575577563570450 another song

nashashmi•20m ago
I wonder if the companies would be willing to bet entirely on AI driven innovation if liability for misalignment was put squarely on companies, individuals, compute vendors, and LLM vendors. I don’t think they would opt for it, especially if an alternative option to use human-programmed tech was already available.

There is something to be said about emphasizing on liability as a way to freeze or solidify AI Development. Right now it is too unfettered leading to predictions of AI dooms.

soumyadeb•18m ago
This also explains why Astra is so good at video generation. I have an Astra+Higgsfield setup. I could point it to a Github repo and ask it to generate a product walkthrough and it did a very good job by generating fake screens (e.g. with data filled in) from real ones - which wasn't possible in earlier models
123917•11m ago
https://x.com/tobiges/status/2098294046469022030

"Sam understands exponentials like no other. During a YC talk last year he predicted that AI would make breakthroughs in science in 2026 and solve a major open problem in 2027. Now here we are..."

Now on a new vibe coded website Astra wins the benchmarks ...

mlmonkey•11m ago
What about Jev? :-D
josefresco•7m ago
Looks like the "most successful" path drove over empty parking spaces and came close to two curbs?
23m ago
AC10 had an interesting post around this general area earlier today:

https://www.astralcodexten.com/p/mysteries-of-ai-generalizat...

seanmcdirmid•18m ago
A later entrant can potentially side step those investments if their now is later. Since self driving car ventures aren’t profitable yet and need to make up their investments over time, thats a real risk for them.
bethekidyouwant•31m ago
Are you using GPT without a harness? Also latency.
publicmail•29m ago
Doesn’t Google own Waymo? I feel like they would have connected the dots.
valine•26m ago
Astra is the first chat model with really strong spatial reasoning. Gemini is nowhere close. Hard to say what google has going on internally, but if they have an astra like model I doubt they’ve had it for very long.
gniv•20m ago
This recent post form Waymo suggests they already use large general models: https://waymo.com/blog/2026/08/10ailessons/
VBprogrammer•29m ago
I'm not sure how you take that from the original article. My 4 year old would drive that course in an automatic car, if only he could reach the pedals. Heck, he's done harder things at Lego land.

I wouldn't let him loose on the road though.

I think, at the very least, the guardrails would have to deterministic, ideally with super human senses, for people to accept self driving cars on the road.

ACCount39•17m ago
Nope, no "deterministic guardrails" for you. The domain is simply far too broad and unstructured to allow for that.

Unless you mean "a typical AI with all the computation constrained sufficiently to always unfold the same exact way, given the same input". In practice, that just kicks the can to "given the same input" street.

The noise in the system is going to come from the input plane. Which is, I remind you, facing the real world. It's full of noise.

moffkalast•29m ago
Yeah cause every car needs 8xH200 pulling 10kW to run a VLM at realtime speeds. Would be unfortunate if 4G dropped out under some trees while using the API after all.
post-it•24m ago
Power usage isn't an issue. 10 kW is 13 HP. The size, price, and fragility of the components is the issue.
AlphaSite•19m ago
GPUs/XPUs are small and solid state so it’s only really price that’s a huge liking factor.

And the disinclination of these companies to push the weights of their cutting edge models into people’s cars where they can be dumped.

tintor•26m ago
lol. Wait until your cloud frontier LLM stalls / disconnects due to load / interference while your car is on highway OR making unprotected left turn OR approaching pedestrians.

It is easy to make car driving *demos*.

JoshTriplett•23m ago
> The bitter lesson is finally coming for the self-driving cars.

Maybe, but the opacity level of models is not acceptable for cars. "Why did it drive under the semi?" "Model said to." "Why did the model say to?" "shrug"

sebastos•7m ago
But if the model is an LLM, you actually COULD ask it why it drove under the semi, and it would give you an answer. Now, you may argue that it will just be generating a whole new, backwards-rationalized post-hoc explanation of its own behavior given the logs that it managed to take before the crash. But then I ask you: how do you think a person explains why they did what they did after a crash? I direct you to all of the unsettling split-brain neuroscience literature demonstrating that humans are incorrigible backwards rationalizers who make for unreliable witnesses.
giancarlostoro•22m ago
Sounds really expensive. I think OpenAI and Anthropic should really not dismiss making smaller capable models that they can license out in this space on the other hand.
atonse•22m ago
Tesla's already solved this - their vision model does this phenomenally well.

And they've demonstrated adding a sidecar LLM to it as well, mostly for these kinds of "read these 3 street signs, what should i do next?" sort of situations.

samuelknight•21m ago
The bitter lesson tells you about the trend in the technology. It does not get product to market with today's technology.
nater5000•21m ago
>It’s mostly a latency problem at this point. The models are too big to run locally, but given that open-weight models like Qwen already exist, an open-weight, low latency equivalent to Astra can’t be too far out.

But this is a bit of a ridiculous take, no?

You don't need Astra for self-driving. Astra is able to build complex 3D worlds, do your taxes, shop for you, and, apparently, drive a car. A self-driving car just needs to be able to drive a car. By the time you trim down Astra to just have the minimum capabilities needed to drive a car, you'll be looking at the same models these self-driving car companies already use. Then you get to deal with the actual hard problems, like handling failure cases (which will still be present with Astra).

>The vision stack, 3D maps, lane selection grammar, occupancy networks, it’s maybe all about to give way to a single GPT looking at camera feeds and predicting the next steering wheel adjustment.

Self-driving cars have been able to do this for a long time. The problem is that it isn't robust enough given the context. I mean, if Astra can drive a car with a single camera, then presumably Astra can drive the car even better with multiple cameras, and even better than that with 3D maps, etc. And when you start to consider the expectation of performance of these systems, you realize that these features really can't be omitted. If you're a company producing self-driving cars, then you do not want to face a lawsuit for you car killing someone because it physically would have never been able to see what it was doing because it lacked a camera.

I think the real gain here is that something like Astra can be used to help build these autonomous stacks. If it is able to drive itself, then it is able to generate novel data, analyze large quantities of data, and use context that isn't typically available when processing this data to make improvements to the actual autonomy stack which is ultimately responsible for driving the car. But thinking that these car companies are going to run an LLM in a car and call it a day is just naive.

giancarlostoro•15m ago
The key thing Astra is doing is a loop... (my understanding) To figure out where things are... It's basically use more compute, self-driving cars are usually using on-device hardware where a "loop" might be a little too risky especially if it takes too long on local hardware... I wouldn't want my AI driving model to be over the air either, yikes in the case of lag or network outages.
valine•11m ago
> But thinking that these car companies are going to run an LLM in a car and call it a day is just naive.

I don’t think so. Massively engineered tech stacks are really hard to fix when they break. Waymo drives past a barricade into a farmers market. How do you fix that? If your self driving car is a GPT, you edit the system prompt and tell it “don’t drive into the farmers market on <road>”. In-context learning and the rapid bug fixes it will afford will be what finally makes self driving cars truly safe and reliably, imo.

sebastos•13m ago
As somebody working near the field, I do enjoy the fun of dreaming bespoke vision and autonomy algorithms (if I didn’t, I wouldn’t work in the field to begin with!). But I would drop it all in a heartbeat for a robot that works well. Robust, resilient robots would be such an incredible advance that the ‘how’ doesn’t matter. All of the nonsense from the current AI hype cycle would be worth it if it cashed out in Robots That Actually Work.
SoftTalker•6m ago
Why do we need robots when we already have people?
ForHackernews•4m ago
I just want a robot butler. It doesn't have to prove mathematics theorems, just do my laundry and make lunch.
rayiner•9m ago
Probably not. In humans, the visual processing circuitry is very different from the circuitry for language processing. There is no reason to believe GPTs will be effective at it.
SoftTalker•9m ago
If I'm reading the chart correctly, it took over 5 minutes to drive 135m at a cost of nearly $8.00 in tokens. I don't think that's really in the realm of practical yet.
ed_balls•6m ago
I think this a slight different lesson. There is one algorithm that is called transformer, rest is irreverent/performance optimization.