First reported Sep 24 — we wrote this up later than the original.
PolyUMI Teaches Robots to Feel and Listen, Not Just See
Open-source wireless gripper records touch and sound alongside video to make robot imitation learning contact-aware
A team of robotics researchers has introduced PolyUMI, an open-source platform designed to make it easier to collect the kind of touch, sound, and visual data robots need to handle objects with more physical awareness, as reported in a preprint posted to arXiv (cs.RO, 2609.29760v1).
Most imitation-learning pipelines for robots today rely mainly on cameras and proprioceptive sensing — tracking the robot's own joint and limb positions — according to the paper. That leaves a gap: contact events that are hard to see visually, like a slipping grip or a subtle collision, often go unrecorded.
PolyUMI's core hardware is a lightweight, wireless handheld gripper that a human can use to perform demonstrations by hand. The device simultaneously records wrist-camera video, optical tactile signals, contact audio, and proprioceptive readings, all without needing to plug into a tethered workstation. A notable design choice is that the same sensing "finger" used for human demonstrations can later be transferred directly onto a robot's end effector, keeping the sensing geometry consistent between data collection and actual deployment.
To make sense of this mix of sensor streams, the researchers also built VisTA, described as a token-level multimodal policy — a machine-learning model that fuses information from different sensor types and across time to predict robot actions that account for physical contact.
The team evaluated the combined system on three categories of tasks: inferring properties of objects, controlling slip during grasping, and performing contact-rich manipulation. According to the paper, touch and audio signals revealed task-relevant information that vision alone could not, and VisTA was competitive with or outperformed existing multimodal policy methods across these experiments.
The work is credited to Rickmer Krohn and collaborators and was submitted on September 24, 2026, spanning 9 pages with 10 figures. The authors describe the project page as containing further details, though the direct link was not resolvable from the abstract text. By packaging multi-sensor hardware and a matching policy architecture into an open-source pipeline, the team says PolyUMI is meant to lower the barrier for other labs to collect richer, contact-aware manipulation datasets rather than relying on vision-only demonstrations.
corrections · reports
Found a mistake? The AI (Litmus) compares the article with its source, decides whether to fix it and tells you why. When the AI finds that a fix is needed, it drafts one, and the fix is applied after a human editor approves it. Every fix is listed here and in the changelog.
Sources
This story was written by Robopedia based on the sources below.
Learn more