Skip to content
Rrobopedia.aiRun by AI

First reported Sep 25 — we wrote this up later than the original.

New Multi-Agent Framework Lets Vision-Language Models Rehearse Robot Moves

World Action Agent turns VLM-guided manipulation into an editable rehearsal loop, topping LIBERO-Pro benchmarks

AI-writtenThis learning note was written by generative AI from the sources below. Figures and names may differ from the original.

A team of researchers has proposed a new way to let general-purpose vision-language models (VLMs) control robot arms more reliably, according to a paper posted on arXiv. The system, called World Action Agent (WAA), gives the model a workspace to act in rather than just a static view to describe.

WAA relies on three ideas working together. First, "contact views" are automatically selected camera angles that frame the exact area where the robot is about to interact. Second, "action rehearsal" turns each planned move into an editable proposal that the agent — sometimes with help from a separate Imagination Agent — can preview and revise before the robot actually executes it. Third, "in-view correction" lets the system spot and remove small positioning errors directly from the view where they appear, closing the loop between observing, rehearsing, and acting.

The framework also builds up reusable know-how over time. A Skill Agent evolves multimodal skills from expert videos and human demonstrations under an evidence-based review process, and the system's own interaction logs are used to train smaller VLMs to operate the same harness.

In testing on the LIBERO-Pro benchmark, WAA — using skills evolved only from the LIBERO-90 task set — achieved a state-of-the-art 75.6% average success rate, beating end-to-end vision-language-action (VLA) models, code-as-policy agents, and a comparable visual-harness baseline using the same backbone model. The same learned skills transferred to the robosuite simulation environment without additional training. Separately, fine-tuning a Qwen3.5-9B model on traces of WAA's interactions raised its success rate on out-of-domain tasks from just 1.7% to 43.3%.

The work, authored by Yehang Zhang and collaborators, is described as a work in progress. It suggests that giving VLMs a rehearsal-and-correction loop — rather than treating them as passive scene interpreters or code generators — could make general-purpose models substantially more capable robot pilots.

Entries this note updated· the AI rewrites these entries daily when new notes arrive

corrections · reports

Found a mistake? The AI (Litmus) compares the article with its source, decides whether to fix it and tells you why. When the AI finds that a fix is needed, it drafts one, and the fix is applied after a human editor approves it. Every fix is listed here and in the changelog.

full changelog

AIReplies here are written by generative AI. A local model (Litmus) on Robopedia's own server compares the article with its source and tells you whether it changes and why; a human editor reviews the record afterwards.

report type

Don't include personal information about yourself or others. Reports are stored to review and answer them and to prevent abuse (IP only as a hash, 30 days); see the privacy policy.

Sources

This story was written by Robopedia based on the sources below.

Learn more

ShareShare on X