First reported Sep 24 — we wrote this up later than the original.
World Models Let Robots Insert Unseen Parts With 56% Success
A single learned world model trained on up to 90 insertion tasks generalizes to new parts far better than model-free baselines.
Robotic assembly lines that handle many different part types usually need a separate, hand-tuned policy for every insertion task — a setup that works well but is slow and costly to redeploy whenever a new part shows up. A new paper, submitted to arXiv and accepted at IROS 2026, proposes a different approach: train one model that generalizes across many insertion problems from the start.
The framework, described by lead author Nicklas Hansen and collaborators, is built around a world model — a learned simulator of sorts that predicts how the robot's actions will change what it senses next. The model combines two streams of information: the robot's own proprioceptive signals (joint positions, forces, and similar internal state) and raw visual input from a camera mounted on the robot's wrist. Rather than training a bespoke policy for each peg-in-hole or connector-insertion task, the team trained a single world model across up to 90 geometrically diverse insertion tasks at once.
The headline result is generalization to objects the model has never seen before, with unknown geometry. On this zero-shot test, the world-model system achieved a 56% success rate. A model-free baseline — a more conventional reinforcement-learning approach without an explicit predictive world model — managed only 7% under the same conditions, according to the paper as reported via arXiv.
The researchers also found a scalability signal that matters for real-world deployment: performance on unseen objects kept improving as more training objects were added to the dataset. That suggests the approach doesn't just memorize a fixed set of parts but builds transferable understanding of insertion dynamics that gets stronger with broader training data — a property that's valuable for factories dealing with constantly shifting product lines.
For cases where zero-shot performance isn't sufficient, the team tested fine-tuning the generalist world model on held-out objects it hadn't seen during pretraining. Compared with training a fresh model from scratch for that specific object, fine-tuning was substantially more data-efficient, and in some cases the fine-tuned generalist even reached better final (asymptotic) performance than the from-scratch policy.
The authors describe this as, to their knowledge, the first system able to assemble previously unseen objects using an entirely data-driven approach — meaning no task-specific geometric models or manually engineered insertion strategies are required. That framing positions the work as a step toward assembly robots that can be pointed at a new part and expected to figure out how to insert it, rather than requiring an engineering team to build a new policy first.
The paper, titled "Generalizable Robotic Insertion with World Models," is available on arXiv (2609.28258) and is slated for presentation at IROS 2026. As with any preprint, the reported success rates and task counts reflect the authors' own experimental setup and have not yet undergone final peer review at the time of this v1 submission.
corrections · reports
Found a mistake? The AI (Litmus) compares the article with its source, decides whether to fix it and tells you why. When the AI finds that a fix is needed, it drafts one, and the fix is applied after a human editor approves it. Every fix is listed here and in the changelog.
Sources
This story was written by Robopedia based on the sources below.
Learn more