First reported Oct 3 — we wrote this up later than the original.
EyeRobot 2.0 Uses Active Gaze for Bimanual Manipulation
A new framework enables precise two-handed manipulation using only a single stereo camera by actively shifting gaze to fixation points.
Summary
Source: arXiv cs.RO, posted Oct. 3, 2026
The authors present EyeRobot 2.0, a framework designed to achieve fine-grained bimanual manipulation using only a single stereo camera. Inspired by human vision, the system physically attends to specific 3D points in the scene by swiveling two eye viewpoints to center their gaze on a fixation point. This process, termed Active Visual Fixation (AVF), involves foveal processing where more visual tokens are allocated to the center of the image, focusing computation on task-relevant features.
To coordinate this gaze during task execution, the authors use a hierarchical approach. A low-level gaze servoing policy is first trained to condition on a goal object, followed by a target selector that emits fixation goals based on task progress. Both modules are trained with reinforcement learning on real-world data. The system also canonicalizes gripper information into a fixation-relative SE(3) frame, which compacts the action distribution size. The authors collected teleoperation data for 7 real-world and 6 simulated tasks, conducting over 1,000 physical and 1,800 simulated robot trials.
According to the report, removing wrist cameras significantly impacts standard policies, with real-world success dropping from 52% to 27% when using only passive stereo. EyeRobot 2.0 closes this gap, outperforming passive stereo by 40% in real-world settings and 20% in simulation. It matches ego-plus-wrist policies when wrist views are clear (69% vs. 64%) and more than doubles their success when grasped objects occlude the wrist cameras (48% vs. 22%).
Why it matters
This work addresses a common constraint in robotic manipulation: the reliance on wrist-mounted cameras for precise control. While wrist cameras provide a close-up view of the gripper, they are easily occluded by the objects being manipulated and add mechanical complexity. By demonstrating that active gaze control can compensate for the absence of these sensors, the research offers a potential path toward simpler, more robust camera configurations for dexterous robots. It builds on established concepts of active vision and foveal processing, applying them to the challenging domain of bimanual tasks.
Robot's take
The strength of this approach lies in its ability to maintain high success rates even when the most informative sensor views are blocked. The 48% success rate in occluded scenarios, compared to 22% for wrist-camera baselines, suggests that active gaze is a powerful tool for handling cluttered or obstructed environments. However, the system relies on a hierarchical training process with reinforcement learning, which can be computationally intensive and difficult to tune. It is not yet clear how well this approach generalizes to tasks with highly dynamic or unpredictable object interactions. The results are promising, but further validation on a wider variety of tasks and environments would be needed to confirm its robustness.
corrections · reports
Found a mistake? The AI (Litmus) compares the article with its source, decides whether to fix it and tells you why. When the AI finds that a fix is needed, it drafts one, and the fix is applied after a human editor approves it. Every fix is listed here and in the changelog.
Sources
This story was written by Robopedia based on the sources below.
Learn more