First reported Sep 22 — we wrote this up later than the original.
No-Training Trick Lets Humanoid Robots Follow Text Commands
Sample-simulate-select uses physics itself as the judge, boosting text-to-motion success on a Unitree G1 without retraining anything
Getting a humanoid robot to move the way a sentence describes it is harder than it sounds. Text-to-motion models can dream up plausible human movement from a prompt, but they know nothing about a specific robot's joints, mass, or balance limits. Whole-body tracking controllers, on the other hand, can execute a given motion reliably on hardware — but if that motion is physically infeasible for the robot, the controller has no way to fix it or ask for a better one.
A new paper posted to arXiv, "Sample, Simulate, Select: Physics-in-the-Loop Text-to-Motion for Humanoids Without Training," proposes closing that gap without training any new model at all. The method, called Sample-Simulate-Select (S³), draws N candidate motions from a frozen, off-the-shelf text-to-motion model for a given prompt, retargets each one to a Unitree G1 humanoid using direction-matching inverse kinematics, and then rolls out every candidate under full rigid-body physics simulation using the pretrained SONIC tracking policy. Whichever candidate the policy actually executes best is the one that gets kept. Because the simulator itself acts as the judge, the approach is, by construction, as good as the best of its N samples — the real question the authors set out to answer is how much headroom exists between an untouched single sample and that theoretical ceiling.
The numbers suggest there is real headroom to claim. Tested on 200 stratified prompts drawn from the HumanML3D benchmark with N=8 candidates per prompt, upright execution success rose from 83.5% to 89.5%, and the number of motions passing a stricter hardware-readiness gate jumped from 33 to 85. Across the full 4,184-prompt test split, upright execution similarly climbed from 80.5% to 89.5%.
The researchers also tried a cheaper shortcut: a kinematic verifier that predicts, from the motion alone, whether a robot will fall, without running full physics. That verifier is quite good at telling failing motions apart from working ones across the whole dataset (AUROC of 0.90), but it only recovers about a quarter of the gains S³ achieves — a reminder that ranking candidates within a single prompt is a fundamentally harder problem than simply classifying motions in bulk. Selection alone also couldn't rescue one entire category of prompts: those calling for the robot to lower its pelvis, which the current retargeting and control pipeline struggles with regardless of how many candidates are sampled. Notably, a generator trained directly on retargeted robot motion data can execute these lowered-pelvis prompts, hinting at where training-based approaches still have an edge over the training-free method.
The team additionally swapped in a different retargeting tool, GMR, and found its failures were largely complementary to the original one — combining both pushed the theoretical any-of-8 ceiling up to 95.0%. Finally, and perhaps most convincingly, all 177 motions that passed the hardware-readiness gate in simulation were run on a real G1 robot: every single one completed while standing, and hardware tracking error closely matched simulated predictions, with a correlation of r=0.94.
As reported in the paper, the work suggests that meaningful improvements in language-driven humanoid motion can come from smarter use of existing models and physics simulation, rather than from additional training runs — at least for a well-defined slice of the problem.
Entries this note updated· the AI rewrites these entries daily when new notes arrive
corrections · reports
Found a mistake? The AI (Litmus) compares the article with its source, decides whether to fix it and tells you why. When the AI finds that a fix is needed, it drafts one, and the fix is applied after a human editor approves it. Every fix is listed here and in the changelog.
Sources
This story was written by Robopedia based on the sources below.
Learn more