frontpage.
newsnewestaskshowjobs

Open Source @Github

fp.

Open in hackernews

Grok 4.7

https://x.ai/news/grok-4-7
58•meetpateltech•47m ago

Comments

ls1911•35m ago
after using cursor grok & trae.ai for several months , grok curor is highly superior results to trae.ai
simianwords•27m ago
I guess it’s only my opinion but having used grok for personal chat: it’s by far the worst one amongst Claude, ChatGPT and even Deepseek, Gemini etc.

The personality is bland and it doesn’t work nearly as hard or even tries to help.

artemonster•14m ago
I used openrouter to send same prompt to qwen, derpseek, gemini and grok and found that grok does good research and produces less bullshit, especially when prompted to be critical of an idea
slowin•13m ago
This has been my experience as well. Grok will end tasks almost immediately and claim "Done!". It's definitely the laziest and most "dishonest" of all the models. The others aren't perfect, but I can't use Grok for any serious coding task.
Capricorn2481•9m ago
> The personality is bland

I don't use Grok, but do you want your LLM to have a personality? "Personality" is exactly what people don't like about Claude.

moojacob•23m ago
Apparently Grok 4.7 has 40% more weights than Grok 4.6, but the price ($6 output token, $2 input) is the same.

Given that the decrease in their margin and the fact they delayed the release of Grok 4.7 almost two weeks past the original date, XAI must not have been happy with the results for 4.7. And XAI also waited the day before Opus 5.5 is rumored to launch. I imagine Opus 5.5 will blow Grok 4.7 out of the water benchmark wise.

However, I have become skeptical of benchmarks. Grok 4.5 solved some issues setting up a buildroot system that Fable 5 couldn't do. I find the post cursor groks are phenomenal at frontend web development, though Claude is much better at backend ruby.

My favorite part of the new Groks has been how they speak in plain english. I simply cannot stand Claudish. Or even GPT, which doesn't have Claude's ticks but definitely likes to handwave explaining technical concepts. Still, nothing beats Claude 3.5 and 4 with explaining since it seems all models have regressed. I wonder if Grok 4.7 will also regress with English because of all the RL.

jasonjmcghee•12m ago
For what it's worth - over the last few years or whatever, it seems like Anthropic benchmaxxes the least.

That being said, I currently prefer Sol / Astra to Opus / Fable as I find both to be a better cost payoff to me.

dumberquestions•5m ago
Token price doesn't tell you much without knowing token efficiency.
kristofferR•19m ago
What's with the deceptive graph on top? Not including Astra can't have been an oversight, did the model compare poorly to it?
Iolaum•15m ago
I wonder if that means that SpaceX evals show that they consider astra better than fable or that they hate Sam&co so much they don't want to show their stuff.
babelfish•14m ago
this is exactly it.
Jcampuzano2•12m ago
https://openai.com/index/our-decision-on-cursor-following-it...

Its because of this. You can't use Astra in Cursor, and cursorbench uses cursor as the harness. They can't actually benchmark it using their harness hence why its not included.

babelfish•11m ago
They have Astra in other benchmarks lower on the page. They just don't want to show it winning
Jcampuzano2•10m ago
The chart is cursorbench though and they asked about the "deceptive graph"
toader•14m ago
How anyone that values democracy in the United States could support any of Elon's ventures is difficult to understand.
AtlanticThird•9m ago
Weird, that's the main reason I purchase all of Elon's products https://time.com/5936036/secret-2020-election-campaign/
thoman23•4m ago
Привет, fellow American!
ctrlkctrls•6m ago
Judging by Elon's staggering success in all of his ventures I'd say you're out of touch.
jackfischer•3m ago
The public very much voted for massive administrative reform. Are you refering to DOGE, Elon Musk's influence on elections, something else?
Jcampuzano2•11m ago
https://openai.com/index/our-decision-on-cursor-following-it...

This explains why. Mentioned in another comment, but cursorbench explicitly tests with Cursor as the harness, and OpenAI doesn't allow them to use Astra in Cursor.

scottyah•5m ago
Deceptive? An extremely quick google search would answer your question. OpenAI pulled out of Cursor before they released Astra so it never got that benchmark.

Googlebook – Meet the Lineup and pre-order

https://googlebook.google/shop/
1•bretpiatt•1m ago•0 comments

Grok 4.7 Intelligence, Performance and Price Analysis

https://artificialanalysis.ai/models/grok-4-7
1•theanonymousone•1m ago•0 comments

Show HN: Agent Chaperone – Screen AI agent tool calls and results with Jev

https://github.com/agent-chaperone/agent-chaperone
1•sepehrsafari•3m ago•0 comments

Show HN: flyOS – A fruit fly connectome simulated in real time on an iPhone

https://www.becomethefly.com
1•coryetzkorn•3m ago•0 comments

Saudi Arabia's Ceer launches flagship electric vehicles

https://www.agbi.com/manufacturing/2026/09/saudi-arabias-ceer-launches-flagship-electric-vehicles/
3•abdullahalharir•5m ago•0 comments

Modulate ML Team Announces New Public Entity Transcription Benchmark

https://www.modulate.ai/blog/modulate-ml-team-announces-new-public-entity-transcription-benchmark
1•voicevector5•6m ago•0 comments

WTF Is Up with Napster's AI Pivot?

https://tedium.co/2026/09/21/napster-ai-pivot/
2•speckx•6m ago•0 comments

Alcor: Simulate cpuid and sgdt/sidt results per-process

https://github.com/er-azh/alcor
1•Tiberium•7m ago•0 comments

Google hit with €403M fine by Irish data watchdog over GDPR violations

https://www.bbc.co.uk/news/articles/ck1e52v16ngxo
2•philbo•7m ago•0 comments

Raspberry Pi founder Eben Upton: 'I'm an Omni-geek

https://www.ft.com/content/2285c111-b103-4b04-96c7-fe26a3c04c3e
1•imichael•8m ago•0 comments

Grok 4.7 is here with Electrical engineering benchmark which beats fable 5.1 max

https://twitter.com/hive_echo/status/2102072241420898708
1•echohive42•9m ago•0 comments

Delta A21N at Kahului on Sep 19th 2026, fuel fumes on board

https://avherald.com/h?article=541628f4
2•r2sk5t•12m ago•0 comments

WWLD #1: Domains and gas station hotdogs

https://chaosguru.substack.com/p/wwld-1-domains-and-gas-station-hotdogs
3•taubek•14m ago•0 comments

The agents, they just want to talk

https://snats.xyz/pages/articles/political_ecology/the_agents_they_just_want_to_talk.html
2•snats•15m ago•1 comments

Amazon Blocks Meta's Muse AI Agent from Its Retail Site

https://www.bloomberg.com/news/articles/2026-09-21/amazon-blocks-meta-s-muse-ai-agent-from-its-re...
3•k2m•15m ago•1 comments

Show HN: Foremerge – Catch Intent Conflicts Between Parallel Coding Agents

https://github.com/naw103/foremerge
4•foremerge•15m ago•0 comments

Measuring nothing (with great accuracy) (2014)

https://seths.blog/2014/01/measuring-nothing-with-great-accuracy/
2•compiler-guy•16m ago•0 comments

Thinking, Fast and Slow

https://en.wikipedia.org/wiki/Thinking,_Fast_and_Slow
3•vinhnx•16m ago•0 comments

Building a Browser Without V8: What Broke, and What Worked

https://medium.com/@Koukyosyumei/building-a-browser-without-v8-what-broke-and-what-worked-bb3bd78...
2•syumei•17m ago•0 comments

Show HN: Fusion-runtime – self-hosted voice agents, STT+LLM+TTS in one process

https://github.com/SamarthUrs18/fusion-runtime
2•samarthurs18•17m ago•0 comments

Show HN: Code Graph View – What if you could navigate a codebase as a graph?

https://github.com/otobongfp/code-graph-view
2•otobong•18m ago•0 comments

How Secrets Work in Docker

https://infisical.com/blog/docker-secrets-management
3•FinnLobsien•18m ago•0 comments

Steerable Cultural Preference Optimization of Reward Models

https://arxiv.org/abs/2606.18606
2•measurablefunc•21m ago•0 comments

pwn.college - Learn to Hack

https://pwn.college/
2•mihau•22m ago•1 comments

First Autonomous Flight Across the United States

https://www.jobyaviation.com/news/joby-completes-first-ever-fully-autonomous-flight-across-the-un...
2•geox•23m ago•1 comments

This Digital Radio Gets Messages to the World’s Remotest Locations

https://spectrum.ieee.org/hermes-shortwave-radio-digital-data
3•SamuraiLion•23m ago•0 comments

Fable 5 – Median thinking declined in August

https://twitter.com/Lon/status/2101793422487204027
9•espeed•23m ago•0 comments

Nginx Control API: View In-Memory Configuration and Reload via HTTP Requests

https://blog.nginx.org/blog/nginx-control-api-view-in-memory-configuration-and-reload-via-http-re...
2•aaaaffine•24m ago•1 comments

The reason children no longer enjoy reading

https://spectator.com/article/the-real-reason-children-no-longer-enjoy-reading/
6•netfortius•25m ago•0 comments

Barry Blitt's "Pulling the Plug"

https://www.newyorker.com/culture/cover-story/cover-story-2026-09-28
1•rbanffy•27m ago•0 comments