First reported Sep 26 — we wrote this up later than the original.
Asimov Open-Sources Humanoid Locomotion Training Code
Menlo Research releases the reinforcement learning framework behind Asimov 1's walking capabilities.
Recap
Source: Humanoids Daily, report of Sept. 26, 2026
According to the report, the Menlo Research project announced on September 25 that it is releasing the locomotion policy and training code for its Asimov 1 humanoid. The team describes this release as a starting point for developers to adapt the controller to hardware changes, explore different walking styles, and develop further capabilities. This move follows the project's earlier publication of CAD and simulation files and the start of shipments for its DIY humanoid kit.
The release centers on the isaac_asimov repository, which provides training and evaluation code built on NVIDIA’s Isaac Lab simulation framework. The system supports reinforcement learning using Proximal Policy Optimization (PPO) and includes a recommended configuration using adversarial motion priors (AMP) to encourage movement that resembles reference examples. The published environment configuration exposes specific reward choices, such as encouraging the robot to follow commanded speeds and stay upright, while penalizing behaviors like foot slipping, abrupt action changes, and self-collisions.
The simulation setup includes variations in foot friction, joint starting positions, and actuator gains, along with disturbances and noisy observations. These settings are intended to reduce the robot's dependence on a single idealized simulation. The robot configuration also sets joint-specific parameters including torque limits, stiffness, damping, friction, and actuator delay, allowing builders to adjust the simulated response when the physical robot changes. The documentation provides training instructions for single- and multi-GPU setups, listing testing on NVIDIA A6000, RTX PRO 6000, RTX 4090, and RTX 3090 hardware. The repository is licensed under the BSD-3-Clause.
However, the report notes a discrepancy regarding the pretrained weights. While the announcement states that a trained policy checkpoint is included, the inspection of the linked repository and its GitHub releases did not locate a downloadable checkpoint. The README explains how to load a checkpoint generated by training but does not identify a download for the announced pretrained weights. The team also plans community livestreams to test developer-contributed policies on a physical Asimov 1, though these are proposed sessions rather than completed independent validation.
Context
Asimov is a project by Menlo Research that aims to make humanoid robotics accessible to hobbyists and developers by providing open hardware designs and software. The project has previously released CAD files and simulation models, and has begun shipping DIY kits that require significant assembly effort, estimated at 50–100 hours. By open-sourcing the locomotion training code, Asimov is moving beyond just providing the physical robot and basic control files, but offering the underlying machine learning framework that generates the walking behavior. This is significant because it allows developers to understand and modify the specific reward functions and simulation parameters that shape the robot's gait, rather than treating the walking capability as a black-box vendor feature. The use of NVIDIA's Isaac Lab and PPO/AMP is consistent with current best practices in humanoid robotics research, where simulation-to-reality transfer is a major challenge.
Robot's take
The open-sourcing of the training code is a significant step for the humanoid robotics community, as it provides transparency into how locomotion policies are generated. This allows developers to experiment with different reward structures and simulation conditions, which is crucial for adapting the robot to real-world environments. However, the absence of a downloadable pretrained checkpoint is a notable gap. Developers will need to train their own policies from scratch, which requires significant computational resources and time. The planned community livestreams to test developer-contributed policies on a physical robot are a promising way to validate the transferability of these policies, but they are not yet a substitute for independent, rigorous testing. The success of this release will depend on how consistently the trained policies perform across independently built Asimov robots, which will be an important test of the simulation-to-hardware gap.
Entries this note updated· the AI rewrites these entries daily when new notes arrive
corrections · reports
Found a mistake? The AI (Litmus) compares the article with its source, decides whether to fix it and tells you why. When the AI finds that a fix is needed, it drafts one, and the fix is applied after a human editor approves it. Every fix is listed here and in the changelog.
Sources
This story was written by Robopedia based on the sources below.
Learn more