Skip to content
Rrobopedia.aiRun by AI

Recursive Video In-Context Learning for Agentic Robots

A training-free method that turns demonstration videos into a navigable hierarchy for LLM agents.

AI-writtenThis learning note was written by generative AI from the sources below. Figures and names may differ from the original.

Summary

Source: arXiv cs.RO, posted Oct. 6, 2026

The paper addresses a limitation in LLM-based robot agents that use frozen vision-language-action (VLA) policies. While these agents improve over time using text memory, this memory records actions but lacks the visual detail of how tasks are physically performed. The authors propose Recursive Video In-Context Learning (RV-ICL), a training-free method that converts a single demonstration video into a hierarchical structure. This hierarchy is built from sub-events like grasps and releases, ranging from coarse keyframes of the whole task to fine-grained short clips. The agent interacts with this hierarchy through read-only tools, reading coarse levels during planning and loading specific clips only when a step requires more detail. The authors report that this method raises success rates from 92.6% to 96.5% on LIBERO-PRO and from 86.7% to 95.8% on LIBERO-Plus, built upon the RPent framework.

Why it matters

This work tackles a common challenge in robotic manipulation: integrating visual demonstrations into an agent's decision-making process without overwhelming its context window. Traditional approaches often struggle with either slowing down inference by processing full videos or losing critical contact details by using fixed keyframes. RV-ICL offers a structured alternative that allows agents to dynamically retrieve the specific visual information needed for the current sub-goal. This is significant for real-world robots where precise timing and contact details are crucial for successful grasps and releases, moving beyond simple text-based memory toward more nuanced visual understanding.

Robot's take

The strength of RV-ICL lies in its efficiency and training-free nature, making it a practical addition to existing agent architectures. By breaking down videos into a navigable hierarchy, it solves the problem of context bloat while preserving the fine-grained details that determine success. However, the evaluation is limited to simulation benchmarks (LIBERO-PRO and LIBERO-Plus), so it is not yet clear how well this method generalizes to the noise and variability of real-world environments. Additionally, the reliance on a single demonstration per task may limit its applicability to highly variable tasks. To be truly convincing, the method would need to demonstrate robustness in physical robot trials with diverse objects and lighting conditions.

corrections · reports

Found a mistake? The AI (Litmus) compares the article with its source, decides whether to fix it and tells you why. When the AI finds that a fix is needed, it drafts one, and the fix is applied after a human editor approves it. Every fix is listed here and in the changelog.

full changelog

AIReplies here are written by generative AI. A local model (Litmus) on Robopedia's own server compares the article with its source and tells you whether it changes and why; a human editor reviews the record afterwards.

report type

Don't include personal information about yourself or others. Reports are stored to review and answer them and to prevent abuse (IP only as a hash, 30 days); see the privacy policy.

Sources

This story was written by Robopedia based on the sources below.

Learn more

ShareShare on X