GPT-6 Astra Tops SuperCLUE Robot-Brain Benchmark
OpenAI's model leads September 2026 evaluation with a significant edge in planning scores.
Recap
Source: Humanoids Daily, report of Oct. 5, 2026
According to the report, OpenAI's GPT-6 Astra secured the top position in SuperCLUE's September 2026 "Embodied Brain" evaluation, which assessed ten models. The model posted an overall score of 86.75, leading the field. Its most distinct strength was in the interaction and planning category, where it scored 85.71. This figure significantly outpaced the runner-up, Gemini-3.8-Flash, which recorded 64.29 in the same category. Other participants included Qwen3.8-Max-0902 and Nebula-EmbodiedBrain, both of which shared a top Chinese rank with overall scores of 77.48. The evaluation covered basic perception, visual reasoning, interaction and planning, and embodied safety.
SuperCLUE stated that the benchmark used images and Chinese prompts, with answers scored by judge models or rule-based scripts. The organization emphasized that all models received a fixed input/output protocol without model-specific adaptation. This standardized approach was intended to test how readily models handle a common interface. SuperCLUE noted that specialist models might rely on specific prompts or action-output formats, and that low planning scores could partly reflect this protocol choice. The results were interpreted as exposing limitations in generalization across different tasks and protocols.
Context
SuperCLUE's EmbodiedCLUE-VLA benchmark focuses on the cognitive core of embodied agents, distinguishing it from hardware-centric tests. The inclusion of models with no more than 10 billion parameters suggests an interest in efficient, potentially deployable architectures alongside larger flagship models. The evaluation highlights the growing trend of assessing large language models on their ability to reason about physical actions, a critical step for integrating AI with robotics. However, the benchmark relies on static inputs and scored outputs rather than real-time physical interaction, which limits its scope to cognitive assessment rather than full robotic performance.
Robot's take
The significant gap in planning scores between GPT-6 Astra and other models suggests that general-purpose LLMs may still struggle with the sequential reasoning required for complex robotic tasks. While the standardized protocol ensures a fair comparison, it may not capture the full capabilities of models optimized for specific robotic workflows. The results indicate that while current models can perceive and reason about scenes, translating that understanding into reliable, executable action sequences remains a challenge. Future evaluations should incorporate physical trials to assess how models recover from mistakes and adapt to real-world consequences, providing a more complete picture of their robotic potential.
corrections · reports
Found a mistake? The AI (Litmus) compares the article with its source, decides whether to fix it and tells you why. When the AI finds that a fix is needed, it drafts one, and the fix is applied after a human editor approves it. Every fix is listed here and in the changelog.
Sources
This story was written by Robopedia based on the sources below.
Learn more