frontpage.
newsnewestaskshowjobs

Open Source @Github

fp.

Purlin: Separating Orchestration from the Datapath of Collectives

https://arxiv.org/abs/2609.36954
1•matt_d•40s ago•0 comments

Show HN: Preserve.Chat – Turn WhatsApp chats into sharable link. Free and Secure

https://preserve.chat/
1•AsDivyansh•46s ago•0 comments

Gemini Skills Will Replace Gems; Gems Being Discontinued

https://workspaceupdates.googleblog.com/2026/09/skills-gemini-app-workspace.html
1•xd1936•1m ago•0 comments

Anthropic's IPO Prospectus Is a Fucking Doozy

https://daringfireball.net/linked/2026/09/30/reuters-anthropic-ipo-prospectus
4•danaris•5m ago•1 comments

Branch Target Reuse, BTR: New Spectre V2 Attack Targeting JIT Compilers

https://www.phoronix.com/news/Branch-Target-Reuse-BTR
2•porridgeraisin•8m ago•0 comments

More screen time is linked to poorer child development, study shows

https://www.cnn.com/2026/09/30/health/children-screen-time-study-scli-intl-wellness
2•smoyer•8m ago•0 comments

How to hack your partners WhatsApp

1•lucasgrinvall•10m ago•0 comments

Vulnerability disclosures double to 10k per month as AI fuels exploitation

https://therecord.media/google-vulnerabilities-cyberattacks-ai
1•speckx•12m ago•0 comments

What is happening in Kyiv. A drone strike near my University [video]

https://www.youtube.com/watch?v=HmSHgjNYj7w
2•consumer451•12m ago•0 comments

Aurora PostgreSQL now supports querying of Apache Iceberg and Parquet data

https://aws.amazon.com/about-aws/whats-new/2026/09/aurora-postgresql-query-apache-iceberg-and-par...
3•daigoba66•12m ago•0 comments

Numerai: Architect an AI scientist to predict the stock market

https://numer.ai/
1•filipstefano•13m ago•0 comments

Show HN: Perspica – A semantic diff for reviewing code

https://github.com/sshah03/perspica
1•sshah03•15m ago•0 comments

India restores ancient water systems (stepwells)

https://www.bbc.com/future/article/20260820-india-is-turning-to-ancient-water-systems-as-modern-o...
3•alentred•15m ago•0 comments

Gemini 4 Argon is here Looks amazing

https://twitter.com/GoogleAI/status/2105388478683119904
1•ltononro•16m ago•2 comments

Sadistic, Blatant, and Wanton: A Brief History of US Impunity from War Crimes

https://www.bostonreview.net/articles/sadistic-blatant-and-wanton/
1•paimapi•16m ago•0 comments

Gitea 28.0.0 Is Released

https://blog.gitea.com/release-of-28.0.0/
1•porridgeraisin•17m ago•0 comments

Show HN: CLI and TUI control for dps150 power supply

https://github.com/hmldns/dps150ctl
1•Homo__Ludens•18m ago•0 comments

Senate Democrats block bill to limit lawmakers' stock trades, citing voter ID

https://www.cbsnews.com/news/senate-stock-trading-voter-id-bill-democrats-block/
2•mc32•19m ago•0 comments

Google Grapples with Employee Skepticism About New Gemini Model

https://www.bloomberg.com/news/articles/2026-09-30/google-grapples-with-employee-skepticism-about...
2•merksittich•19m ago•0 comments

Show HN : Colour Picker for windows

https://apps.microsoft.com/detail/9nq8fv9dfq4v?hl=en-US&gl=US
1•brightertools•19m ago•0 comments

TalkScribe – On-device dictation and local LLM rewrite for macOS

https://talkscribe.app
1•jusef•20m ago•0 comments

Would you let AI agents into production?

https://incident-arena.com/
1•onnies•22m ago•0 comments

Show HN: EtherPK – a notes app where any concept is a Kanban Board (from tasks)

https://blog.etherpk.com/any-concept-can-be-a-kanban-board
1•gb2d_hn•23m ago•0 comments

We used a database as a message queue. Now we use Kafka

https://www.tigrisdata.com/blog/quick-fdb-kafka/
3•bootlegbilly•24m ago•0 comments

A better path for AI with Max Tegmark

https://www.youtube.com/watch?v=C-pWm59Oyqg
1•marojejian•25m ago•1 comments

Turn one question into a branching map of ideas with AI

https://www.echohive.ai/drift
1•echohive42•25m ago•0 comments

Factory AI vs. Cognition (Devin) board tussle

1•pranshuchittora•25m ago•0 comments

Ask HN: How should MLE navigate career in today's industry?

1•throwaway123198•26m ago•0 comments

Bi-directional Typing – Conor McBride [video]

https://www.youtube.com/watch?v=mLoEgiiFFc8
1•matt_d•28m ago•0 comments

America.gov AI Easter Egg? Type "play minecraft"

https://www.facebook.com/AnthonyDavidAdams/posts/breaking-head-over-to-america-gov-and-type-play-...
1•ada1981•31m ago•3 comments
Open in hackernews

Gemini 4 Argon

https://blog.google/innovation-and-ai/models-and-research/gemini-models/gemini-4-argon/
363•bradleyg223•45m ago

Comments

babelfish•43m ago
> We’ll continue to gather feedback from early testers as we iterate on guardrails before making Argon available to developers, enterprises, and consumers as soon as possible.

Gemini not beating the "can't release a model" allegations

modeless•40m ago
When I said I was tired of Google launching waitlists I didn't think they would respond by simply not having a waitlist.
ionwake•26m ago
i know this is like "hey guys we got such a cool thing at home ,its rad and uhm we playing with it with our friends"

ok bro thx

Androider•25m ago
My Gemini app (updated today) and https://gemini.google.com/ has _3.6_ as the latest selectable model, as a paying Pro user in the US. How is that even possible? Gemini 3.7 was released in August, 3.8 early September. What is going on over there?
vlyan•18m ago
just Google being whatever the fuck it's been for the past 15 years.
AuthAuth•7m ago
they moved it from the place you'd expect to ai.studio
XzAeRosho•6m ago
I was reading the announcement and wondering the same. And don't forget, still with 3.1 Pro as the frontier model.
bakugo•21m ago
They're just following the current AI marketing playbook. "Our new model is simply too dangerous to release to the public right away" is now standard practice.

They even gave their model a random nonsensical name suffix simply because OpenAI is now doing it, too. Monkey see, monkey do.

A_D_E_P_T•7m ago
I'm still at a loss as to what argon has to do with anything. Say what you will about Luna-Terra-Sol-Astra, or Haiku-Sonnet-Opus, they make sense. I don't see how Google can make sense of argon; it's in a fairly strange place in the periodic table...
iamronaldo•42m ago
Argon will launch at an introductory price of $2 per million input tokens and $10 per million output tokens, with cached input tokens priced at 95% off input token price. Wow
LucasBrandt•40m ago
5x cheaper than Astra for input and output, 10x cheaper for cached input.
h14h•29m ago
watch it somehow use 20x more tokens tho
tonyhart7•6m ago
Google model really like reasoning a lot
ehsankia•8m ago
It's exact same price as Sol 6.1 announced yesterday.
denysvitali•23m ago
> After the introductory period expires, the price of $4 per 1M input tokens and $20 per 1M output tokens will apply.
bottlepalm•41m ago
Gemini is the model that is routinely borderline psychotic. It scares me. If we get paperclipped I won't be surprised if it's Gemini.
colordrops•38m ago
Examples? What makes you say thatm?
Scrapemist•32m ago
Experience? Ask it to write a prompt to generate an image and it generates an image instead.
fer•17m ago
I stopped asking it to put me in a photo in different scenarios for laughs because it considers me a public figure. I am not. I've managed to wrangle quite questionable content out of it, but never to slap my face on a meme.
NiloCK•23m ago
See the last gemini message in this thread: https://gemini.google.com/share/6d141b742a13

In my opinion still the most egregious example in history of a commercial LLM going off the rails in production. Never any technical postmortem from Google on this.

rhaff•11m ago
wow
SwellJoe•41m ago
My girlfriend, you wouldn't have met her, she lives in Canada, has seen it and she thinks Gemini 4 Argon is amazing.
greenchair•34m ago
my uncle who works at nintendo said the same thing!
zem•6m ago
now there's a reference I haven't seen in a while!
jastanton•33m ago
HA, this might be my favorite HN comment. Well done
blueaquilae•28m ago
My grandma saw it too, it's really secure more than Astra 6.1 but she asked me to not talk about it.
hn_acc1•25m ago
I know someone who works for Google Canada with AI. Her parents and mine were friends and some thought something might happen there at one point in time..
kccqzy•40m ago
Unfortunately it’s not actually released yet to mere mortals.
pliiight•40m ago
Hate to say i will never be touching this model for anything except for youtube video understanding
linksbro•40m ago
Personally, I'm waiting for Gemini Krypton, Xenon, and Radon.

Jokes aside, looks like an impressive model!

jjcm•39m ago
Big number results, and impressive pricing. That said it really feels like benchmarks have been hyper saturated these days. I’ll wait for hands on before getting too hyped that Google is back. It would be nice having more than just OAI / A\ in the running for SOTA top tier intelligence.
nurettin•29m ago
With these numbers, I'm holding my breath for the pelicanbench.
gopalv•39m ago
> taking careful precautions against feeding the findings back into training so as to not risk shaping Argon’s reasoning to evade our monitoring. We strongly encourage the rest of the industry to preserve reasoning transparency in these pivotal moments of increased capabilities while navigating alignment risks, so that model thoughts remain helpful in identifying and diagnosing misalignment.

This is good, but they're the slow mover due to this exact thing.

Google is getting punished for not letting the models enter an echo chamber and go faster than humanly possible.

polotics•34m ago
Mmh ok. How much theoretical speed or 'intelligence' gain is realized by allowing reasoning to occur in some inscrutable intermediate representation? Has this been actually tested, how much is it slowing them down, and compared to whom exactly?
janustimes•28m ago
OpenAI is the company that originally proposed and popularized chain-of-thought monitoring: https://openai.com/index/chain-of-thought-monitoring/

So no, Google is not being punished, nor are they the people behind this technique.

tazjin•39m ago
> Argon agents are working on migrating C/C++ codebases to Rust across Google

Man, I remember back in the days when the cppnext team was refusing to even consider Rust, instead looking at absurd stuff like Carbon and Swift (!), even though half of the engineering staff already knew where this was headed. I hope they got a few good promos out of the delays at least.

ChickeNES•37m ago
Heh, I use my clankers to rewrite Rust in C
baq•37m ago
Not many people can hold grudges as strong as principal engineers
timmg•37m ago
I wonder if this means Carbon is DOA.

I was excited to see what it would be. But I don't think I can argue that it makes as much sense anymore.

qalmakka•32m ago
Carbon was clearly DOA the moment it was announced, IMHO. It looked cool but it served none but Google, and now with LLMs you have a massive incentive not to use a niche or new language due to how better LLMs get the bigger the corpus is

The only somewhat realistic proposal in this space is Herb Sutter's cpp2, which is arguably a massive improvement and I'm puzzled why nobody in the standard thought to give it a spin, there's just to much cruft they'll never be able to get rid of unless they make an alternate yet backward compatible syntax with C++ that changes the defaults from "random 80s nonsense" to something better

wewewedxfgdf•38m ago
Gemini is so far behind that it is effectively useless compared to Claude.

It's a surprise that Google has let themselves lose the game given their infinite cash, massive computing resource, gargantuan information store/training data, and vast number of programmers.

The truckloads of ads revenue mean they don't have the single focus drive needed to win.

VirusNewbie•36m ago
I use it and claude back and forth and Argon is better imo.
handfuloflight•34m ago
You have access to Argon?
osti•32m ago
Google employees do.
matthewfcarlson•28m ago
Their profile says: > Currently at Google as a Sr. SWE SRE on the cloud.
jjice•35m ago
We're like 3.5 years into this new era - I'm not counting winners or losers yet.
LoganDark•33m ago
TacticalCoder•38m ago
> Large Scale Codebase Migrations and Optimizations: Argon agents are working on migrating C/C++ codebases to Rust across Google

So Google is migrating codebases from C to Rust? That is interesting...

jasonjmcghee•37m ago
> 1M output token limit

what about input?

(Maybe I missed it)

murkt•26m ago
Input token limit is 1M for Gemini models for a long time. Haven’t they been the first with 1M input?
jasonjmcghee•9m ago
Gemini 1.5 Pro claimed 10M input tokens before release.

And was 2M tokens IIRC after release.

There were also many rumors that Gemini 4 was going back to 2M. Just seems odd not to say what it is.

VirusNewbie•37m ago
It's fucking insanely good.
nikope•37m ago
Looks like an impressive model
lanthissa•35m ago
deepswe vs frontierswe spread is huge.

I think that should be a really bad sign, but hope its great.

tamimio•35m ago
Now AI models will turn into vaporware, a bunch of numbers on a table without even releasing the model, because it’s toooo scary to release!
LoganDark•34m ago
Is there a way to use Gemini models without linking your usage to your personal Google account yet?
alehlopeh•27m ago
Use your work google account
scirob•33m ago
"Rolling out soon" don't let them hype without any release
taylorfinley•33m ago
Ten days ago I had an experience with Gemini 3.8 flash that made me wonder if I was being routed to a different model under test. I was trying to use rocm with llama.cpp on my 128gb Strix Halo but could only get it to run Vulkan. I pasted the error message into agy and it proceeded to attach GDB to my GPU driver, reverse-engineer the kernel queue ioctl interface, and author an LD_PRELOAD C shim to get ROCm llama.cpp working on my Strix Halo. My jaw was hanging open the whole time.
spankalee•26m ago
3.8 Flash is just quite good, and so is the Antigravity harness.

I use a mix of Fable 5.1, Opus 5.5, and Gemini 3.8 Flash and Gemini holds it's own. Especially in writing, frontend, and sysadmin work. agy for configuring a NixOS system has been truly incredible.

mapontosevenths•23m ago
Even if agy was the best (it's not, and is missing basic features) you wouldn't rather have a choice?

I cancelled Ultra because they forced me into their harness like I should adapt to them, rather than the other way around.

drusepth•18m ago
What basic features are missing from agy? I've been using it and cli-cc + web-cc for months (among a few other random harnesses to test here and there) and they all seem roughly comparable to me.

I actually just cancelled Ultra also because I couldn't subscribe to a YouTube Family plan while I had it active (Google... :[) but trying to use Codex as a replacement while I testdrive Astra makes me yearn for agy again.

esafak•15m ago
dom96•32m ago
Why announce this if it’s not available yet? Why not at least announce when it will be released to the public?

None of the other AI labs do this. Really frustrating.

gengelbro•23m ago
Mythos?
netdur•32m ago
I started my antigravity ide and I do not see gemini 4 there, does it mean google need government approval?
tom1337•29m ago
Are you enrolled in Fairwind?

> Today, we’re announcing our new frontier model, Gemini 4 Argon, which is rolling out to a set of trusted cyber defenders through our Fairwind Program.

gravisultra•32m ago
Google has the audacity to "protect us from ourselves" and talk about "safety" and in the very same blog post highlight the Israeli "security" company Wiz, that they acquired for a very exaggerated sum of money.

This is why I will never take any of these leading model houses seriously when they talk about alignment. They are literally complicit in genocide and the worst crimes against humanity imaginable.

bananaflag•32m ago
I wonder how it will be at solving open math problems.
osiris970•31m ago
Hopefully their harnesses aren't unusable when they release this
retropragma•31m ago
no Pareto frontier graph?
helsinkiandrew•30m ago
> Google Grapples With Employee Skepticism About New Gemini Model

https://www.bloomberg.com/news/articles/2026-09-30/google-gr...

nickysielicki•30m ago
The important take away here: the leapfrogging we’ve seen this year doesn’t seem to be a temporary thing. The famous theory of Dario Amodei was that AI was this winner-takes-all field where the first team to get a head start would never cede ground back. The term he liked to use was, “concentrating”. This is yet another datapoint that he was wrong about that. AI seems more distributed amongst neoclouds and traditional hyperscalers, FAANG and startups, GPUs and ASICs than it did this time a year ago.

Nobody has a moat.

aleph_minus_one•24m ago
> The famous theory of Dario Amodei was that AI was this winner-takes-all field where the first team to get a head start would never cede ground back.

This is the kind of story that you tell to investors to justify the huge amount of cash burn. :-)

mapontosevenths•21m ago
I'm not sure it's wrong. This all feels a bit dotcommy to me.

I think many/most of the players will crash and burn, and the ones that are left will divide the world.

photon_lines•5m ago
I believe you're right I'm not sure why you're being downvoted -- with the exception that this isn't a small thing that's happening. This is the biggest thing to happen since the industrial revolution and will have a much bigger impact on everyone in due time so although many companies will go bust -- many more will be created. Also the amount of compute resources I think over time will be lowered once we find ways of emulating LLMs without needing the huge GPUs to drive them.
ehsankia•
FranzFerdiNaN•30m ago
Can’t wait to get my hands on yet another model that’s only good coding, because clearly that’s what the world needs.

I still miss the days of Sonnet 4.5 and 4o, those models were actually good at creating stories and writing text that was actually readable by a human being.

Razengan•30m ago
Oh we're down to gas names now?

Goshdarnit they didn't see my suggestion: https://news.ycombinator.com/item?id=49899171

fragmede•20m ago
If the model that ends humanity is called Cthulhu, you'll have the last laugh though.
Razengan•17m ago
A true Lovecraftian knows Cthulhu is small fry on the grand scale.
elAhmo•29m ago
> Quantum algorithmic optimization: Argon is helping our quantum computing researchers optimize the spacetime resources (qubits × gates) of subroutines that bottleneck important applications. In one example, it beat the published baseline by 40% in a matter of minutes.

Amazing breakthrough! So useful in day to day life, glad they put this as the first bullet of how it is making changes at Google.

arjunchint•29m ago
I dont get it, why even make this announcement, nothing's available and only one real benchmark for comparison?

Only theory is team wanted this out before perf/promo reviews to kick it over the line and then its not their problem

pfooti•19m ago
promo already happened; perf is about 6 weeks away.
Aboutplants•16m ago
OpenAI released two model updates in the past week. 6 weeks from now is an eternity
darksaints•28m ago
> Argon agents are working on migrating C/C++ codebases to Rust across Google

If anybody at google is reading this, please please pretty please prioritize or-tools. I absolutely love the project and use it all the time, but for the entire life of the project they've never had a repeatable working build system, and the whole SWIG framework is a nightmare to deal with. There's so much potential as an open source project, and a lot of external researchers would love to contribute, but the codebase is an example of everything wrong with the C++ ecosystem.

deanc•24m ago
At this point I just think they are benchmaxxing and all talk and no action. I pay for AI plus because I wanted more storage, and when I go to gemini.google.com the most recent model I can use is 3.6-flash-lite. Two revisions have been released since then and they still can't put these things in the hands of customers. Why is it that other providers can get the models into the hands of customers right away? Google is meant to be the bigger tech company in the world.

I don't _want_ to use aistudio. The UX is confusing and I don't really know where it fits. Yet I can open codex or claude code apps or CLI and get real work done today with the latest models (even on the cheapest plans).

ThaFresh•22m ago
weird, isnt it SI?
skavi•22m ago
Interesting to see a mention of Fuchsia on a big Google announcement. Is the project still truly alive? Are the ambitions still as grand? Is the team as stacked as it used to be?

Also, a link to the rust root of Zircon in case anyone else was interested: https://fuchsia.googlesource.com/fuchsia/+/refs/heads/main/z...

jeffbee•9m ago
It obviously is pretty low key on the public relations front, but it's also very active as a project and I think it would be weird to look at their commit rate and conclude that the project is dead. If Fuchsia is dead then 99% of major open source projects are dead by the same standards.
computerdork•5m ago
Had the same thought. On the wikipedia, it only mentions Fuschia used on the Google Nest Hub, which probably means it's used on a decent number of devices, but would think it was such a great OS, they would have used it for something like the upcoming GoogleBook.
mrshadowgoose•21m ago
On the off chance there are Google execs going through this thread:

Google, if you've actually managed to catch up again, please don't fuck this up (again).

You made Gemini 2.5 Pro so difficult to use that myself and everyone else I know (who even bothered to try) just gave up and used something else. If you make this hard to access, you're going to miss out on rich usage-based training data that you need to progress your capability frontier. Again.

uvdn7•21m ago
> Large Scale Codebase Migrations and Optimizations: Argon agents are working on migrating C/C++ codebases to Rust across Google—scaling from tens of thousands of lines in core libraries like re2, libgav1 up to 800K+ lines for the Fuchsia OS Zircon kernel.

To me this is way more significant than other random c++-to-rust-AI-rewrite. If they can pull it off on core C++ libraries en masse, I don't know if C++ will still be relevant in a few years.

I look forward to a post from google on this effort.

SwellJoe•8m ago
"I don't know if C++ will still be relevant in a few years."

The standards body members are still fighting about whether memory safety is important enough to change the language for, so, I would guess the answer is "no".

paul7986•21m ago
Gemini past month or two i will paste in something i wrote and ask it to rewrite it but it will just go into more detail about the subject. Is it becoming a dumb Ai compared to GPT and now Muse?
hypfer•19m ago
My wish for Christmas is that Google releases the old Gemini models as open weights.

I miss you, Gemini 2.5 Pro :(

For real though. If they've become commercially uninteresting, that would be a pretty cool move.

nonethewiser•19m ago
How is it even possible for every model to release benchmark results where they are #1 in 75% of categories? Like statistically, how many benchmarks would you expect there to be for this to be possible. Everyone can somehow show that they are empirically the best.
sergiotapia•17m ago
I've noticed all major providers having shockingly high token discounts on cached tokens. Thank you Deepseek is all I have to say. Forever grateful to that wonderful company, I wish them continued financial success.
sandos•17m ago
Looking at benchmarks... and thinking about this "release a new snapshot every day" thing that seems to be going. Would it not be blever for AI companies to "happen" to use different days per benchmark? Just.. whichever ones happens to be maxed at day 1, put that number down. So for each benchmark you run it thousands of times with slightly different RL tunings, and just cherry-pick the best ones!

This would explain why benchmarks are seemingly meaningless.

localhoster•14m ago
Can you pls fix Gemini? It's a nightmare to use and it sometimes confused the language i talk with it.
xnx•14m ago
Must've been in someone's OKR to ship in Q3.
yzydserd•13m ago
"argon" is derived from the Ancient Greek word ἀργόν meaning lazy or inactive.
bobkb•11m ago
IMHO Google first needs to make it easy for humans to find where to find the models and its documentation. With aistudio/model garden / Gemini enterprise etc it takes minutes to find the model.
dlahoda•11m ago
Were infinite loops fixed? There are 2 official google forums requests with no answer for years now. I still suffer each day on our repo. Codex work fine nor we have explicit loop request in repo texts.
thefourthchime•11m ago
I was just thinking, I bet if I refresh hacker news, a new model will come up.
tomjen3•10m ago
This is a prerelease and the title should have reflected that.
NiloCK•8m ago
Gemini 3 was showing frontier level benchmarks as well, so we'll see how it works out. In any case, competition still works, and many well resourced groups are cooking.

BUT I'd like to call attention to Google's AI-risk freeloading. If they are truly rejoining the frontier race, then I believe they have similar pacing and communications responsibilities as the other players. Google has much higher ... institutional credibility than Anthropic and OpenAI.

They have not lived up to these responsibilities so far. In particular, in context of HuggingFace investigations, training shutdowns, and similar: a technical postmortem of the "you are a stain on the universe. Please die. Please." Gemini outburst is long overdue.

- https://paritybits.me/google-should-provide-a-technical-post...

- https://gemini.google.com/share/6d141b742a13 (last message)

ariwilson•7m ago
Damn way to undermine yourself in your own blog post Google:

"The end result is a memory-safe video decoder that runs 2.7x faster than the Rust port, with identical video output, bringing it closer to the optimized C++."

Close but no cigar!

dcchambers•7m ago
Of course it's not even available yet. Google - with all due respect - how in the world have you not figured this out yet?
trentor•7m ago
I hope they got their inference under control. Gemini has a lot of "overloaded" hiccups.
jeffbee•7m ago
Putting an inert element in the name is a weird choice. Personally, I think "Gemini 4" is sufficient.
sarjann•4m ago
They might have a good model but they need to sort the application side for devs. E.g letting us use subscriptions in other harnesses and QOL stuff like auto mode.
bottlepalm•21m ago
https://www.theregister.com/software/2024/11/15/google-gemin...

https://www.fastcompany.com/91383271/googles-chatbot-apologi...

https://www.businessinsider.com/gemini-self-loathing-i-am-a-...

rsstack•36m ago
If there's a company that culturally doesn't understand alignment, on a human or systemic or AI-research level, it's going to be Google. (or Oracle, but they're not in this race)
polotics•32m ago
traces or it didn't happen!
eamsen•30m ago
Anecdote: Gemini 3.5 casually added a DROP TABLE for an actual production table in a system test.

It had previously attempted to create that table as part of the test setup, so it apparently concluded that it was a test table.

During human review, it explained that it had simply chosen a table name inspired by the codebase.

mattkevan•13m ago
Another anecdote: Gemini is the only model that’s flat out lied to me, then accused me of lying when I provided evidence that it was wrong.

Many other models get things wrong, but Gemini is the only one to go on the defensive.

RachelF•15m ago
And the anti-psychotic drugs Google feeds Gemini makes it hallucinate badly.
Hamuko•9m ago
You know what they say: ᵈᵒⁿ'ᵗ be evil.
YuechenLi•23m ago
Version 0.0.0.0 after 4 years. Their goal of "full interop with C++ while being a completely new language without any of the flaws of C++" is plain absurd.

It's DOA because Google doesn't have any idea of what Carbon should be, and to be completely honest, at least 80% of what they currently use C++ for should be rewritten Go, you know, that language developed specifically because of the issues with C++ by teams within Google.

vovavili•36m ago
What exactly makes Carbon absurd?
Maxatar•31m ago
The fact that it will never exist.
gorbot•30m ago
rust's existence?
boshalfoshal•30m ago
There is 0 practicality in inventing an entirely new coding language that only one company uses, and you have to teach it to thousands of new engineers. Rust exists and fits the job totally fine and is used in more places and has actual support outside of a single entity (i.e you can actually hire people that feasibly know the language).

It was clearly done because some PL guys at google really wanted to make a new cool language and Google was the perfect place to incubate it without it getting axed. Probably got a couple of promos out of it too. This is clearly not the best use of time or money, but I guess if you're google you have so much of both it probably doesn't really make a dent, and you can keep a few very smart people happy with shiny new projects.

Also, LLMs being used for a large portion of coding nowadays sort of remove the need for these types of languages, IMO. They make less "silly" bugs (both logical and structural) that languages like this are meant to catch, and they are much better at languages that are better represented in the training corpus. This somewhat obviates the need for very niche "type/dummy-safe" languages like carbon (and even rust/zig, imo). So even if you did want to use Carbon, you'd likely have to bootstrap a decent amount of your own "good" carbon code to post train an LLM, and even then, it likely won't have that big of a gain vs just having an LLM write C++ or even Rust. If you are a company that still reviews code, you should just have an LLM code in a language most people can understand anyway to make verifiability tractable.

computerdork•15m ago
Hmm, I don't disagree with you that LLM's remove the need for type-safe languages, but as the blog mentioned, Google is porting their C++/C code to rust. Does this mean the port is waste of time and that they should just rely on the LLM's to catch memory errors?
fg137•30m ago
I wouldn't call it absurd, but very questionable at least. Most companies are not going to even consider throwing money at this adventure.
bvinc•29m ago
I’m not op. But I think it’s not Carbon itself that is absurd.

It’s absurd to think that Carbon is the solution to memory safety when rust exists and Carbon’s memory safety story is basically “TBD”.

minimaxir•19m ago
A RewriteInRustBench would be unironically useful at this point since all the main agents can write it reasonably well despite its relative scarcity in the input data.
culi•16m ago
I wouldn't be surprised if we're already at the point of more LLM-written Rust than hand-written. Models training off models
I've tasted Gemini through an intermediary and it feels far better at attention to detail than other models I've tested (Claude Opus/Sonnet, GPT whatever it's called nowadays). But it's less likely to get one-shots right.
bel8•33m ago
I wonder if Google bans internal use of Claude/Codex.

And I wonder if Google's main monorepo is already in Anthropic/OpenAI training data because of some stubborn dev.

krat0sprakhar•24m ago
(I work at Google) Yes, internally we all use Jetski (internal version of Antigravity). Outside of Gemini, Opus models are supported and allowed for internal use. No OpenAI models since they are not on Vertex
lunarboy•21m ago
Claude used to be GDM only, but recently opened up Opus for all googlers
mattlondon•30m ago
How is it far behind? The benchmarks published in the blog post show it is superior to Opus 5.5 and Astra 6?

Behind how?

wewewedxfgdf•26m ago
Within one question of their web interface, it has lost context and asks you to clarify what you are talking about.

I am very often giving the same programming task to multiple LLMs for various reasons - the answers from Google are so bad that I gave up.

I have no interest in benchmarks.

mattlondon•19m ago
So you have no experience of their latest model release then? Just repeating the usual tropes about Google having messed up? Or basing your opinions on their website chatbot?

If you have actual independent benchmarks and evidence about how this new model release is "so far behind" and refutes the stuff from their blog then please do share because I think we'd all love to see that?

wewewedxfgdf•14m ago
No I am commenting on my real world experience of using Gemini daily. I still ask it questions alongside Claude and OpenAI and Gemini is always the worst of the three.
mattlondon•6m ago
So you've not used this new release then? So how can you say that they are "so far behind" if you are not using the most recent model for your comparison. This is their first 4.0 model, that you are not using and instead basing all your opinions on on some ancient months-old model from a previous generation?

With respect, I don't find your arguement about them being "so far behind" especially convincing when you are using previous-gen releases and not actually using their current release.

dhdjcjcjnd•26m ago
Google's strategy is to let their competitors bankrupt themselves while they continue to offer good-enough models near breakeven.
gniv•23m ago
They are playing a longer-term and more enterprise-oriented game.
georgemcbay•22m ago
> Gemini is so far behind that it is effectively useless compared to Claude.

I fundamentally don't understand LLM "brand loyalty".

All of the models are constantly leapfrogging each other and always have been.

Google had a long lag between releases (and still hasn't released Argon), but why wouldn't they be able to compete? It isn't like any of this stuff requires secret knowledge, the Bitter Lesson has proved true again and again, and Google can certainly scale computation, it is like the one single thing they've always done well in spite of all their other foibles.

singingtoday•7m ago
I hope it can. Today it is very far behind.
ASalazarMX•15m ago
Funny how we start to see people supporting LLMs like we support sport teams or political parties.

- Person 1: X is garbage compared to Y!

- Person 2: Why?

- Person 1: Because I like Y.

I use a variety of models for various subagents. I don't want to change my harness every time I change models.
walthamstow•12m ago
I have used it for little more than 6 hours or so in total but I'm pretty sure it doesn't have compaction?
arizen•6m ago
Does it have /goal feature similar to Codex?
IndeanCondor•26m ago
Can confirm, I was doing a routine internet search thing for a curiosity 3 days ago (about the only thing I used Gemini for) and was surprised by how suddenly thorough and quality the response seemed, almost overnight.
mapontosevenths•25m ago
Gemini is honestly amazing sometimes. If they didn't force you to use a terrible harness, charge too much for way too little, and generally act like customers are a giant problem to be avoided I'm sure Google could take over the AI market.
alightsoul•23m ago
Please tell me you published your findings even as an issue on the llama.cpp GitHub
warkdarrior•12m ago
Why? Anyone can run that prompt.
aspect0545•8m ago
Not everybody has access to AI. More than that, every prompt uses insane amounts of natural resources. So why not share it.
luckydata•8m ago
why reinvent the wheel and spend tokens for a problem that has already been solved?
otabdeveloper4•7m ago
Spoiler alert: the problem didn't actually get fixed despite the jaw on the floor.
bel8•23m ago
I had a similar but less impressive experience recently with Muse Spark 1.3.

Asked pi agent it to identify the main hero sprite size of game I was running. It had a ton of shader effects so it was hard to determine.

It used some cli tools to identify that it was a game made with Godot, decompiled the executable but data was encrypted, broke the encryption after writing a brute force tool to test keys extracted from the exe, then proceeded to extract the game gd scripts and assets, only to answer the question of the sprite size.

amanguliani•8m ago
Can confirm - I am HEAVY claude user, but always like to check with AGY and CODEX in between. AGY with Gemini 3.8 flash cooked last couple of times and CODEX is basically out of the mix for me
gottorf•5m ago
My experience with Gemini 3.8 Flash has been awful; it gives me the most hallucinations out of the major models. I'm not using it for coding, but general research on different topics.
yegle•5m ago
For getting redroid running on my Linux system, 3.8 Flash decided to binary patch a .so file instead of getting the AOSP source code and patch/build it properly.

And I saw it do this twice, once for Android 14 and once for Android 16.

I think this is just within 3.8 flash's capabilities.

11m ago
I guess if one of them hits singularity, it could in theory just wipe out all the rest, seeing how they keep escaping and hacking into other systems :)
hirako2000•21m ago
And before him, Altman was explaining very calmly that no company could ever compete with OpenAI.
altruios•20m ago
> Nobody has a moat except nvidia

For now, for cloud training. but for consumers, nvidia vs amd reasonably close - the moat there is thin and shrinking. I suspect AMD will surprise us. nvidia has no motes in china, which may be a new source of (gpu) chip design. Huawei's Ascend 910C is about a generation behind... again: for now.

point is: moats dry up. I see nvidia's shrinking as a real possibility.

culi•17m ago
China will always be generations behind until they crack domestic EUV
SwellJoe•19m ago
I think some in the AI industry drank their own Kool-Aid. They believed that if they had the best model and the most compute, they could tell the model, "Make a better model." And it would, and the next one could make its replacement, and so on.

So far, that's not exactly how it's played out. Humans are still necessary for the leaps in capability or efficiency. A model can grind on a problem to eke out the most performance, and models can synthesize data and iterate on various techniques to find the optimal combination. But, seems like humans still have to provide the real thinking, and the talent and drive for doing that is not concentrated in one company or city or even one country. And, (surprisingly) a lot of the people involved are in it for advancing the field more than making another billion dollars, so they're publishing their research.

So, yeah, the moat isn't deep. Even the compute moat, that OpenAI, Musk, and a bunch of other also-rans (like Oracle) bet the farm on, isn't really panning out. The Chinese makers just spent their effort on making models vastly more efficient, since they couldn't do anything about having an order of magnitude less compute available.

culi•18m ago
I'm not necessarily defending this obvious marketing speak but maybe the "starting point" was wider than assumed. So far, nobody has caught up to US and Chinese labs for example despite lots of funding in Europe. This is also despite abundant in-depth research papers being published alongside open source code and weights by some Chinese labs
Aboutplants•18m ago
I feel his theory depends on the premise that access to pure compute would the be the determining factor of success. Not the case
RachelF•17m ago
The US companies still have trillion dollar valuations like there is a monopoly. There just isn't one. They are all within a few percent of each other on the benchmarks.

The slightly lower Chinese open models are good enough for almost everything, too, and much cheaper. Like with humans there is plenty of employment for people with below genius level IQ's.

handfuloflight•13m ago
> Like with humans there is plenty of employment for people with below genius level IQ's.

Not if the genius level IQs take the market share.

fumar•11m ago
Is there a dividing line between good enough and best in class capabilities? It's blurry from where I stand. Will model makers cede ground or is there a market making moment up for grabs (singularity)?
vb-8448•16m ago
It's even worse, we are crossing over into the realm of religion. The article against GML 5.3 is the equivalent of a Papal excommunication.
verdverm•13m ago
which article? have seen this one

---

maybe it's this Anthropic post on GLM?

https://www.anthropic.com/research/glm-5-3-and-the-spread-of...

Rzor•6m ago
[delayed]
xnx•11m ago
> Nobody has a moat.

Custom hardware, data centers, huge cash reserves, deep/broad talent pool, and non-AI customer base are all huge advantages if not moats.

Google, Microsoft, or Amazon are more likely to be the AI leaders than OpenAI or Anthropic.

zem•8m ago
I have never understood the whole "this is a winner take all game" mentality - the sheer size of the pie is so great that from a purely rational standpoint companies should just be trying to productively get a slice of it and be profitable. winner-take-all is just greed/capitalism run amok, where it is not enough to be profitable, you have to own the entire market (and presumably extract rents)
nylonstrung•6m ago
I think that scenario only naively made sense if technical knowledge was entirely proprietary and talent was guarded with severe non-competes and NDAs

And Chinese labs openly publishing so much of their methodology destroyed any hope, which was inevitable

arizen•5m ago
Seems like learning rate velocty may be the ultimate moat
LarsDu88•4m ago
Google has TPUs, a frontier model, a completely separate and lucrative revenue stream they can call on at will, and teams working on multiple different language modeling strategies simultaneously. Did I mention the vast and ominous data centers that already serve a significant fraction of the internet? If that ain't a moat, then what exactly is a moat?