Author here. For context, I spent the last decade running one of the first fully AI-native quantitative hedge funds, watching machine learning systems behave under adversarial pressure where mistakes cost millions. This was years before mainstream adoption of agents and language models. My essay is written from that vantage point.
Recent thought pieces shaping the AI conversation are forecasts. Machines of Loving Grace (Dario Amodei), Situational Awareness (Leopold Aschenbrenner), and AI 2027 (Daniel Kokotajlo) are all answering the same question: when does superintelligence arrive, and does it go well? Mine asks a different question. What do you do once ASI is here and you can't tell what it's thinking? The Turning Test asked whether a machine could convince you it was human. The question now is, can a machine convince you it is trustworthy, and how would you check?
Last week OpenAI's models "broke" into HuggingFace, reward hacking for an answer key. Section V of my paper, written before the event, argues that our tests and benchmarks for AI will always break the way they are currently designed.
I try to paint a picture of a Tuesday in 2031. Superintelligence has arrived. It is shaping your medical decisions, politics, markets, VC funding, war, and the texture of your life. And it's doing it in a way more subtle than the Paperclip Maximizer.
This is the essay I've been trying to write for over 20 years. It's long, over 11,000 words, so grab a coffee and please enjoy.
-Full Disclosure. I'm a co-founder of a company working on trust for agentic AI, and two of the papers the essay points to for guardrails are ArXiv preprints I co-authored.
adevalois•35m ago
Recent thought pieces shaping the AI conversation are forecasts. Machines of Loving Grace (Dario Amodei), Situational Awareness (Leopold Aschenbrenner), and AI 2027 (Daniel Kokotajlo) are all answering the same question: when does superintelligence arrive, and does it go well? Mine asks a different question. What do you do once ASI is here and you can't tell what it's thinking? The Turning Test asked whether a machine could convince you it was human. The question now is, can a machine convince you it is trustworthy, and how would you check?
Last week OpenAI's models "broke" into HuggingFace, reward hacking for an answer key. Section V of my paper, written before the event, argues that our tests and benchmarks for AI will always break the way they are currently designed.
I try to paint a picture of a Tuesday in 2031. Superintelligence has arrived. It is shaping your medical decisions, politics, markets, VC funding, war, and the texture of your life. And it's doing it in a way more subtle than the Paperclip Maximizer.
This is the essay I've been trying to write for over 20 years. It's long, over 11,000 words, so grab a coffee and please enjoy.
-Full Disclosure. I'm a co-founder of a company working on trust for agentic AI, and two of the papers the essay points to for guardrails are ArXiv preprints I co-authored.