I started to build Is AI Dumber Today because I kept seeing people complaining like "GPT/Claude/Gemini get dumber, is it just me?", usually after an update, or before new models launch.
We saw benchmarks on lots of platforms, but they don't keep tracking models qualities, and the model channels, harness, configurations make such tracking more complex.
So I built this tool to track how people describe their real experiences hourly, from the platforms like Reddit, Hacker News, Zhihu, Rednote etc., which are the major platforms in English and Chinese communities. It's even more interesting to see how people feel and react on the same models in different regions.
You can use it to:
- See which models perform better or worse than usual and the reasons from user opinions, with hourly latency (Because I am a data engineer)
- View model performance from user opinions across the lifecycle of model
- Compare models based on user’s feedback and recommend top models in categories like coding, reasoning, speed, etc.
- Report your own experience with one tap
For the limitation, I would like to say it is an observational experience index, but not a real benchmark. Although I don't think benchmarks can help us to make decisions to choose models, I also don't think Is AI Dumber Today is more accurate. It just provides us with information from another perspective. I believe as the AI developing, and the iteration of this product, we can extract more useful opinions, like which models do users prefer exactly, and how they are using them. We can learn from others easily.
I would appreciate feedback on the methodology, and what would make the tool more useful or trustworthy.
schafberg•49m ago
I started to build Is AI Dumber Today because I kept seeing people complaining like "GPT/Claude/Gemini get dumber, is it just me?", usually after an update, or before new models launch.
We saw benchmarks on lots of platforms, but they don't keep tracking models qualities, and the model channels, harness, configurations make such tracking more complex.
So I built this tool to track how people describe their real experiences hourly, from the platforms like Reddit, Hacker News, Zhihu, Rednote etc., which are the major platforms in English and Chinese communities. It's even more interesting to see how people feel and react on the same models in different regions.
You can use it to: - See which models perform better or worse than usual and the reasons from user opinions, with hourly latency (Because I am a data engineer) - View model performance from user opinions across the lifecycle of model - Compare models based on user’s feedback and recommend top models in categories like coding, reasoning, speed, etc. - Report your own experience with one tap
For the limitation, I would like to say it is an observational experience index, but not a real benchmark. Although I don't think benchmarks can help us to make decisions to choose models, I also don't think Is AI Dumber Today is more accurate. It just provides us with information from another perspective. I believe as the AI developing, and the iteration of this product, we can extract more useful opinions, like which models do users prefer exactly, and how they are using them. We can learn from others easily.
I would appreciate feedback on the methodology, and what would make the tool more useful or trustworthy.
I also write weekly newsletters based on the data I collected and my observations here: https://isaidumbertoday.substack.com/
You can also follow me on X: https://x.com/isaidumber