Nice! Is is missing Codex in the agent harnesses comparison IMO.
grim_io•49m ago
I'd expect google to do well here, since they were historically strong at multimodal and physics.
gizmodo59•46m ago
Yet another "benchmark to promote their own harness"
giwook•41m ago
Please forgive my naivety, but are world models (once they are in a consumer-ready form) expected to outperform any currently existing LLM on these sorts of tasks (i.e. of the physical world)?
jespinel•49m ago