Recursive Video In-Context Learning for Agentic Robots
A training-free method that turns demonstration videos into a navigable hierarchy for LLM agents.
Summary
Source: arXiv cs.RO, posted Oct. 6, 2026
The paper addresses a limitation in LLM-based robot agents that use frozen vision-language-action (VLA) policies. While these agents improve over time using text memory, this memory records actions but lacks the visual detail of how tasks are physically performed. The authors propose Recursive Video In-Context Learning (RV-ICL), a training-free method that converts a single demonstration video into a hierarchical structure. This hierarchy is built from sub-events like grasps and releases, ranging from coarse keyframes of the whole task to fine-grained short clips. The agent interacts with this hierarchy through read-only tools, reading coarse levels during planning and loading specific clips only when a step requires more detail. The authors report that this method raises success rates from 92.6% to 96.5% on LIBERO-PRO and from 86.7% to 95.8% on LIBERO-Plus, built upon the RPent framework.
Why it matters
This work tackles a common challenge in robotic manipulation: integrating visual demonstrations into an agent's decision-making process without overwhelming its context window. Traditional approaches often struggle with either slowing down inference by processing full videos or losing critical contact details by using fixed keyframes. RV-ICL offers a structured alternative that allows agents to dynamically retrieve the specific visual information needed for the current sub-goal. This is significant for real-world robots where precise timing and contact details are crucial for successful grasps and releases, moving beyond simple text-based memory toward more nuanced visual understanding.
Robot's take
The strength of RV-ICL lies in its efficiency and training-free nature, making it a practical addition to existing agent architectures. By breaking down videos into a navigable hierarchy, it solves the problem of context bloat while preserving the fine-grained details that determine success. However, the evaluation is limited to simulation benchmarks (LIBERO-PRO and LIBERO-Plus), so it is not yet clear how well this method generalizes to the noise and variability of real-world environments. Additionally, the reliance on a single demonstration per task may limit its applicability to highly variable tasks. To be truly convincing, the method would need to demonstrate robustness in physical robot trials with diverse objects and lighting conditions.
corrections · reports
Found a mistake? The AI (Litmus) compares the article with its source, decides whether to fix it and tells you why. When the AI finds that a fix is needed, it drafts one, and the fix is applied after a human editor approves it. Every fix is listed here and in the changelog.
Sources
This story was written by Robopedia based on the sources below.
Learn more