There are no sites like artificialanalysis.ai or arena.ai for healthcare benchmarks so I decided to build it. I ran the leading models from openai, anthropic, and google on MedXpertQA: https://github.com/TsinghuaC3I/MedXpertQA#leaderboard
The results show scores across organ systems and across diagnosis, treatment, and basic science. Gemini 3.8 flash and gemini 3.1 pro seem to lead across the other LLMs. The last time eval results were run and published online for this dataset was 2025 so huge progress has been made: https://medxpertqa.github.io/
Overall I spent about ~$150 on LLM tokens for the evals. I also discovered that the text only dataset for medxpertqa is pretty dirty. The field would benefit greatly from cleaner and better maintained datasets, since this dataset is currently being publicly used by Google and Meta.
toshvelaga•39m ago
The results show scores across organ systems and across diagnosis, treatment, and basic science. Gemini 3.8 flash and gemini 3.1 pro seem to lead across the other LLMs. The last time eval results were run and published online for this dataset was 2025 so huge progress has been made: https://medxpertqa.github.io/
Overall I spent about ~$150 on LLM tokens for the evals. I also discovered that the text only dataset for medxpertqa is pretty dirty. The field would benefit greatly from cleaner and better maintained datasets, since this dataset is currently being publicly used by Google and Meta.