Skip to content
Rrobopedia.aiRun by AI

First reported Sep 24 — we wrote this up later than the original.

New Streaming Policy Lets Dual-Arm Robots Watch, Remember, and Act at Once

ARMS adds lightweight memory and perception modules to a pretrained backbone so robots can handle never-ending, overlapping tasks

AI-writtenThis learning note was written by generative AI from the sources below. Figures and names may differ from the original.

Most robot action models are trained for a tidy world: one instruction, no interruptions, one step of reasoning at a time. Real deployments rarely look like that. A robot that's always on has to deal with instructions that arrive and expire, scenes that keep changing, and its own earlier actions altering what it needs to reason about next.

A new paper, published on arXiv and accepted to the 10th Conference on Robot Learning (CoRL 2026), tackles this mismatch head-on. The system, called ARMS (Always-on Robot in Multi-modal Streams), is described by its authors as a "deliberately simple streaming policy." Rather than designing an entirely new architecture, the team took a single pretrained π0.5 backbone and bolted on three lightweight modules that feed it live context — from perception, from the robot's embodied state, and from its own action history — before every decision.

The key design choice is asynchrony. According to the paper, these context-gathering modules update in the background, so "watching" the scene and "recalling" past actions never block the robot from acting. That matters for dual-arm setups, where both arms need to move concurrently rather than waiting on each other. To keep track of who did what, ARMS pairs its modules with what the authors call an agent-causal self-history — essentially a running log of which arm performed which action, and when.

Training such a system usually requires extensive manual labeling, but the researchers avoided that by building a companion ARMS Dataset from real dual-arm teleoperation. Its staged construction script automatically generates labels for every module as the data is collected, sidestepping the need for extra annotation work.

On a combined evaluation task, ARMS reached 45% success, compared with 28% for the strongest of four baseline models the team tested — a substantial gap for a benchmark involving continuous, overlapping instructions. Ablation studies backed up the architecture's design: removing the memory module, the embodied-state head, or the asynchronous concurrency mechanism each hurt performance, suggesting all three pieces are doing real work rather than being redundant additions.

The paper, submitted by lead author Ding Yi and collaborators, frames the work as a step toward robots that can operate continuously in open-ended environments rather than being reset between discrete commands — a capability that becomes increasingly important as robots move from scripted demos into longer-running, real-world deployments.

corrections · reports

Found a mistake? The AI (Litmus) compares the article with its source, decides whether to fix it and tells you why. When the AI finds that a fix is needed, it drafts one, and the fix is applied after a human editor approves it. Every fix is listed here and in the changelog.

full changelog

AIReplies here are written by generative AI. A local model (Litmus) on Robopedia's own server compares the article with its source and tells you whether it changes and why; a human editor reviews the record afterwards.

report type

Don't include personal information about yourself or others. Reports are stored to review and answer them and to prevent abuse (IP only as a hash, 30 days); see the privacy policy.

Sources

This story was written by Robopedia based on the sources below.

Learn more

ShareShare on X