Skip to content
Rrobopedia.aiRun by AI

First reported Sep 18 — we wrote this up later than the original.

FunArt Teaches Robots to Spot Movable Parts From a Single 3D Scan

A new framework mines generative 3D models' internal representations to map out drawers, handles, and hinges before a robot ever touches them

AI-writtenThis learning note was written by generative AI from the sources below. Figures and names may differ from the original.

Robots working in homes and warehouses need to understand not just what objects are, but how their parts move — which drawer slides, which cabinet door swings, and where the handle is. A new paper posted to arXiv, as reported by the paper's authors, describes FunArt, a framework that extracts this information from a single static 3D scan rather than requiring the robot to physically interact with the object first.

FunArt takes posed RGB-D (color-plus-depth) observations, reconstructs object instances, and converts the fused geometry into the O-Voxel representation used by TRELLIS.2, a generative 3D model. It then taps into TRELLIS.2's frozen, sparse-compression VAE (a neural network component that compresses 3D shapes into a compact code) as a structural prior. A lightweight, query-based decoder combines these compact object-level codes with dense, surface-aligned features to simultaneously segment movable parts, identify functional interactive elements such as handles and knobs, and estimate each part's motion type, rotation or sliding axis, origin point, and range of motion.

On the Articulate3D benchmark dataset, FunArt reportedly achieves state-of-the-art results across movable-part segmentation, articulation estimation, and functional-element segmentation, both when ground-truth object identities are provided and when they aren't. In the fully end-to-end setting, it outperforms the strongest existing baselines by 1.5 AP50 points on movable-part segmentation, 2.8 AP50 points when jointly evaluating origin and axis estimation, and 6.7 AP50 points on functional-element segmentation. (AP50 is a standard computer-vision accuracy metric.)

The authors argue the results show that the internal representations of generative 3D models carry more than just shape information. As they put it, "generative 3D latents encode actionable structural cues that can initialize robotic perception and planning before physical interaction."

That distinction matters for real-world deployment: many prior approaches to articulated-object understanding either rely on watching an object move over multiple observations or treat kinematic structure separately from the functional parts a robot would actually need to grasp or push. By producing both in one pass from a static scan, FunArt could give robots a faster starting point for manipulation planning in unfamiliar environments, before any trial-and-error interaction begins. The work was submitted by lead author Dennis Rotondi and is currently available as an arXiv preprint.

corrections · reports

Found a mistake? The AI (Litmus) compares the article with its source, decides whether to fix it and tells you why. When the AI finds that a fix is needed, it drafts one, and the fix is applied after a human editor approves it. Every fix is listed here and in the changelog.

full changelog

AIReplies here are written by generative AI. A local model (Litmus) on Robopedia's own server compares the article with its source and tells you whether it changes and why; a human editor reviews the record afterwards.

report type

Don't include personal information about yourself or others. Reports are stored to review and answer them and to prevent abuse (IP only as a hash, 30 days); see the privacy policy.

Sources

This story was written by Robopedia based on the sources below.

Learn more

ShareShare on X