The fascinating this is that the LLM is not acting as a tool here AFAIk, but very much like a colleague.
I have no knowledge of the domain and have only PhD EE level math knowledge, so maybe my bar is too low.
There is clearly intelligence there. We have no way to recognise intelligence other than the appearance of intelligence and this very clearly displays that.
It's also quite clearly different to human intelligence in some notable ways, but not in any that preclude describing it as intelligent. At least for normal non-pedantic definitions of the word.
I can use AI for coding after decades of coding. I can't use it for theoretical physics because I can't evaluate the responses.
Another satisfied customer!
High IQ bros.
The first one was someone proving another conjecture false by just repeatedly saying "keep going" to ChatGPT: https://x.com/DmitryRybin1/status/2079904005652893709
What a world we live in.
without someone independently verifying it, it just dangles there
...
Two, at some point AIs will be able to use other context like the fact that this is Terrence Tao and not your average Joe and change how it answers, either in tone or structure.
Just awesome to see new knowledge hit an incredible mind like this. Having these "what if" discussions is what I miss most from JPL and academia.
Expand the entire expression, then change the representation to find the core axis. You can't see the axis from just one perspective, so you change the representation. In programming terms, it's like applying multiple domain models. Then break it down into small contract units. Why is it a Jacobian monomial? Why does x satisfy a cubic equation? And so on.
Then swap out the modeling under a hypothesis, assemble it all back together, and verify it through the equation.
This feels similar to modeling in programming.
Observe the whole -> explore better modeling -> decompose local problem -> verify independently -> reason about the highre level structure -> integrate back into the original problem.
This feels similar to when I receive work from a client and write a programming proposal
Sorry it's a bit of an aside, but I imagine many other otherwise "technical" folks feel the same unfamiliar sense of total loss like when encountering hard mathematics.
Similarly, at some point somebody pointed out to me "the reason you're confused is that the bold on that variable means it's a matrix"
1. The counter example wasn't just a brute force selection, the polynomial is structured in a very specific way that ends up getting the result.
2. Terry Tao's questions are very specific and prompts the AI in a useful way, that without high math training you are not going to get the same information out of it. Terry seems to see some aspects of the problem and counter example and uses AI to brute force some parts of it.
a) The model thinks on some questions while straight answers on others. (I wish I'd knew from the questions if this is somehow correlated to hard tasks or "inventive" tasks, but that's way out of my league).
b) The model sometimes pushes back. Again, I'd wish I knew if it was warranted, but I counted 2 instances where it said "yes, but with caveats", one where it said "mostly yes but with this correction" and one where it said "careful here, because x y z".
c) The model did q&a + pdf ingestion + code writing + more q&a + thinking + more q&a, for a looong while, while seemingly staying on topic (at least Terrence Tao seems to think they're still productive, so I'll trust that).
This is what model progress is, not number goes up on xBency or yBencher. Damn.
Is there any way to tell a conversation's model and thinking level?
Where will we be in another 4 years? What a time to be alive!
We have no way to recognise intelligence other than the appearance of intelligence and this very clearly displays that.
There is about 150 years of cognitive science experimentation in animals that have clarified a little bit how you can actually measure intelligence. Ooorrrrr we can use medieval contempt for scientific thinking and pretend "intelligence" is just some higher intuition that can never be falsified. You know it when you see it, bro! Don't listen to Emily Bender, she's a socialist witch.The point of those cognitive science experiments is that they apply to any animal with a brain and plausibly show a real shared concept of "intelligence" that isn't limited to humans. According to this concept, orcas might be smarter than humans, despite their physiological inability to make tools. It's not a "normal, nonpedantic definition" of intelligence because such a definition would be scientifically meaningless.
Indeed, AI's fundamental sin, going back to Alan Turing, is embracing a definition of intelligence that applies to civilized humans, but not to hunter-gatherers, let alone apes, corvids, and cetaceans. Frustratingly, our modern society has two concepts of intelligence:
- an intuitive, social sense of "how smart is this guy?", which is well-understood and, being highly correlated with IQ, a totally pseudoscientific artifact of human psychology
- the poorly-understood scientific concept I mentioned earlier
If AI researchers cared about scientific thinking, they would be intensely focused on the brains of bees. Insteac they love money and sci-fi but have pure contempt for science, even Demis Hassabis. This is why AI researchers have yet to build a robot that navigates real-world 3D space as intelligently as a cockroach. I don't think any of our grandchildren will live to see a computer smarter than a mouse. (It seems like Fable still struggles with small-number arithmetic. Rodents don't.)
A second corollary is that rational consciousness and thought is less likely to be contained in language than previously thought, because if language is so simple that a machine can process it, it can't contain consciousness.
What I do notice however is that LLMs are becoming capable of doing an increasing part of the intellectual work I can do, and usually a lot faster.
Just today I presented an agent framework that can take an informal incident statement and propose infrastructure changes to fix it, all evidence backed. This did nothing I could not to, but it did all 5 test cases in 6 - 12 minutes each. I would have found all of the monitoring indications it did, but it would have taken me a day per test case. The LLM also included sass to silly tickets. ("This is not even worth spending monitoring resources on. It's obviously a configuration problem.")
That's how this is reading to me as well. It's just fast at slogging through a certain level of "simple" transformations.
vb-8448•55m ago
Is it something "revolutionary" or just another small brick that will pile up until something really "revolutionary" will happen?
rickypp•53m ago
echelon•32m ago
Maybe they'll find a solution where P=NP.
That could really throw a wrench into the whole internet thing.
It seems they need an expert human driver for now.
ConceptJunkie•23m ago
sdwr•39m ago
hyperhello•18m ago
Legend2440•8m ago
It's not quite the Reimann hypothesis, but many prominent mathematicians have spent years working on this problem. Yitang Zhang wrote his PhD thesis on it.
arm32•34m ago
By itself, no consequence. But over time, provided we keep pumping out talented and qualified mathematicians and keep subsidizing costs, we could maybe hit a breakthrough... somewhere... that has real impact.
fragmede•26m ago