First reported Sep 17 — we wrote this up later than the original.
GIST AI Teaches Robots to Grasp What Handles and Buttons Actually Do
A new open-vocabulary 3D scene graph links objects, their small parts, and the functions those parts serve.
Knowing where a microwave sits in a room doesn't help a robot heat up food. It also needs to find the door handle to open it, and locate the button that starts the cooking cycle. Humans grasp this instantly; robots historically have not, according to a report from South Korea's Robot Newspaper (로봇신문).
A research team led by Professor Kim Ui-hwan in GIST's AI Department, with integrated master's-PhD student Kim Ee-rum as first author, has developed an AI framework designed to close that gap. The team introduced a concept called the Unified 3D Scene Graph (Unified 3DSG), which links objects and their parts to the functions and possible actions associated with them, all within a single structure. To build this graph automatically from ordinary camera footage, they created OP3DSG (Open-Vocabulary Part-Aware 3D Scene Graph Generation).
A 3D scene graph is essentially a spatial "knowledge map" that helps a robot understand its surroundings and plan movement or task sequences by connecting objects and their relationships. Existing 3D scene graphs, however, have mostly represented whole objects—a chair, a microwave—without capturing the smaller components that make those objects usable, or what those components actually do. Small parts are difficult to track because they're tiny and look very different depending on the camera angle, so prior systems either detected only the whole object or failed to match the same part seen from different viewpoints.
OP3DSG works in stages. First, it detects a main object in the scene, then actively searches for the sub-parts it expects to find—for an oven, that means looking for buttons, a door, and a handle. Next, it merges information from multiple camera views into a single 3D space, comparing not just each part's location and semantic meaning but also its color, to correctly link the same small part across different shots. Once basic spatial relationships are established from position and distance data, a large language model (LLM) reviews and adds functional relationships and likely actions—for instance, connecting a microwave-handle spatial link to the functional fact that the handle "is used to open the door."
The team evaluated OP3DSG using a new benchmark they created, called UniGraph3D. The correct part appeared among the AI's top three candidates 83.6% of the time, a 31.2 percentage-point jump over the best-performing existing method, which scored 52.4%. OP3DSG also outperformed prior best methods by 7.9 points on spatial-relationship accuracy and 9.2 points on functional-relationship accuracy.
To test real-world feasibility, the researchers deployed the system on a Stretch 3 mobile robot. The robot successfully counted nearby chairs, located a remote control to turn on a TV, and—given the instruction "tidy up the laundry"—planned a sequence to find towels and place them in a basket. Notably, when told "I'm thirsty, bring me something I can drink from" without naming a specific object, the robot used the functional relationships in its scene graph to identify a suitable item and set its location as a navigation goal.
"The core of this research is connecting not just information about what exists and where, but what parts can be used and how they can be used," Kim Ui-hwan said, adding that the technology could serve as a foundation for service robots to translate natural-language instructions into concrete task plans in complex real-world settings.
The work was supported by South Korea's Ministry of Science and ICT and the Institute of Information & Communications Technology Planning & Evaluation (IITP) through the AI Bots Collaboration Platform and Self-Organizing AI programs, the Ministry of Science and ICT and National Research Foundation of Korea's Young Researcher support program, and the Ministry of Science and ICT and National Research Council of Science & Technology's Global TOP Strategic Research program. The paper, titled "OP3DSG: Open-Vocabulary Part-Aware 3D Scene Graph Generation for Real-World Environments," was presented at the European Conference on Computer Vision (ECCV) 2026 in Malmö, Sweden, held September 8–12 and organized by the European Computer Vision Association (ECVA).
corrections · reports
Found a mistake? The AI (Litmus) compares the article with its source, decides whether to fix it and tells you why. When the AI finds that a fix is needed, it drafts one, and the fix is applied after a human editor approves it. Every fix is listed here and in the changelog.
Sources
This story was written by Robopedia based on the sources below.
Learn more