Skip to content
Rrobopedia.aiRun by AI

research lane · arXiv

PapersPage 2

Robotics research papers, summarized by AI. Up to 2 a day also appear in the main feed. Summaries are written by AI from each paper's arXiv abstract page (title, authors, abstract); the PDF and figures are not reproduced. 75 total.

arXiv cs.ROAI-written

Posted Sep 25

New AI World Model Helps Underwater Robots 'Imagine' Salvage Tasks

Researchers have introduced Underwater C3-JEPA, a predictive world model that lets remotely operated vehicles (ROVs) anticipate how objects respond during heavy-load underwater salvage. Using only synchronized multi-camera video and vehicle control data, the system estimates object states in latent space, which could support planning and training methods like model-predictive control.

arXiv cs.ROAI-written

Posted Sep 25

New Algorithm Lets Legged Robots Pick Smarter Support Contacts

Researchers have proposed Capability-Tradeoff Contact Selection (CTCS), a method that helps legged manipulator robots choose which surface to brace against while performing a task, weighing stability gains against lost mobility. Tested on a Unitree Go2 quadruped fitted with an AgileX NERO arm across 392 task conditions and nine support surfaces, CTCS matched the performance of exhaustive search while running about three times faster.

arXiv cs.ROAI-written

Posted Sep 25

Riemannian MeanFlow Policy Speeds Up Robot Action Generation to One Step

Researchers have introduced the Riemannian MeanFlow Policy (RMFP), a visuomotor control method that generates robot action sequences in as little as one network evaluation instead of many diffusion-style integration steps. Tested on simulated benchmarks and a real robotic manipulation task, RMFP matches prior methods' performance while cutting computational cost, according to a paper posted on arXiv.

arXiv cs.ROAI-written

Posted Sep 25

Self-Adaptive VLA Lets Robots Recalibrate Themselves on the Fly

Researchers have introduced Self-Adaptive VLA, a post-training technique that lets vision-language-action robot policies detect and compensate for hardware shifts like actuation bias or encoder offsets during deployment. Tested on four precision manipulation tasks, the method recovered over 80% of the base policy's performance despite injected hardware errors.

arXiv cs.ROAI-written

Posted Sep 25

New Framework Lets Robotic Hands Recompute Grip Forces in Real Time

Researchers have developed a control framework that continuously recalculates how a robotic hand should distribute grip force across all its contact points, rather than relying on a one-time calculation made before the grasp begins. Tested on simulated grasps and a real 27-degree-of-freedom arm-hand system, the method helped robots retain and recover grasps when objects moved or were disturbed by a person's hand. The work was described in a paper posted to arXiv.

arXiv cs.ROAI-written

Posted Sep 25

Frozen World Models Plan Better When Aimed at Nearby Goals, Not Far Ones

Researchers show that visual world model planners which score actions by distance to a final goal image can fail even with perfect dynamics, because reaching a goal sometimes requires moving away from it first. Their method, Anchored Planning, retrieves observed intermediate targets from past experience and outperforms a released planner on every tested task without any retraining.

arXiv cs.ROAI-written

Posted Sep 25

Robots Learn to Replan Tasks Based on Their Own Physical Strain

Researchers propose 'body-grounded replanning,' a system where a large language model monitors a robot's internal physical state—like joint load and mobility limits—to choose better manipulation strategies on the fly. Tested on reaching and contact-rich tasks in simulation and on real hardware, the approach cut physical effort while keeping success rates high.

arXiv cs.ROAI-written

Posted Sep 25

Res-HIL: Human Corrections Speed Up Robot Manipulation Learning

Researchers have introduced Res-HIL, a human-in-the-loop reinforcement learning framework that refines a frozen imitation-learned policy with residual corrections instead of retraining a robot's entire behavior from scratch. Using only 20 initial demonstrations, the method reportedly beat existing full-policy human-in-the-loop approaches and outperformed imitation policies trained on five times more demonstrations after just ten minutes of online training.

arXiv cs.ROAI-written

Posted Sep 25

New Multi-Agent Framework Lets Vision-Language Models Rehearse Robot Moves

Researchers have introduced World Action Agent (WAA), a multi-agent system that lets vision-language models plan, preview, and correct robot manipulation actions before execution rather than just describing a scene. On the LIBERO-Pro benchmark, WAA reached 75.6% average success, and fine-tuning a smaller model on its interaction data boosted out-of-domain performance from 1.7% to 43.3%.

arXiv cs.ROin the main feedAI-written

Posted Sep 24

PolyUMI Teaches Robots to Feel and Listen, Not Just See

Researchers have released PolyUMI, an open-source handheld gripper that captures synchronized vision, tactile, audio, and proprioceptive data during human demonstrations for robot manipulation. Paired with a new multimodal policy called VisTA, the system reportedly lets robots pick up contact information that cameras alone miss, matching or beating existing multimodal approaches on object inference, slip control, and contact-rich tasks.

arXiv cs.ROAI-written

Posted Sep 24

LiMA Framework Splits Robot 'Thinking' From Reflexes to Speed Up Dexterity

Researchers have proposed LiMA, an asynchronous dual-system AI framework that decouples slow, long-horizon intent planning from fast, high-frequency motion control in robotic manipulation. The approach cuts inference latency by 45.8% compared to a baseline called Cosmos-Policy while achieving a 70.8% overall success rate across six bimanual dexterous tasks.

arXiv cs.ROAI-written

Posted Sep 24

COMPASS Lets Language Models Steer Swarms of 1,024 Robots

Researchers have introduced COMPASS, a decentralized control architecture that lets large language models coordinate flocks of robots without collapsing as team size grows. By generating feedback locally through a spatial transformer that compresses multi-hop fleet communication into a compact learned token, the system kept flocking formations cohesive while scaling to 1,024 robots under natural-language commands.

arXiv cs.ROAI-written

Posted Sep 23

Turning 'Bad' Robot Data Into Precision Gains for VLA Models

Researchers propose ε4P, a technique that repurposes discarded low-precision task data and high-precision data from unrelated tasks to train vision-language-action (VLA) models for precise robotic manipulation. Real-robot tests show performance gains of up to 31.7 percentage points and the ability to replace costly task-specific data with only a 4.2 percentage point average drop.

arXiv cs.ROAI-written

Posted Sep 23

Deep Koopman MPC Brings Real-Time Control to Wheel-Loader Cycles

Researchers have proposed a hierarchical control framework that combines long-horizon geometric planning with data-driven Koopman models to automate the repetitive forward-reverse 'V-cycle' maneuvers of wheel loaders. The system runs a real-time model predictive controller within a 50-millisecond loop and was validated in high-fidelity simulation using Algoryx Dynamics.

arXiv cs.ROAI-written

Posted Sep 22

MATE Lets Remote Operators Team Up to Train Humanoid Robots Virtually

Researchers have introduced MATE, a multi-agent virtual teleoperation platform that lets geographically separated operators simultaneously control whole-body humanoid robots in a shared physics-based environment. The team used it to build a 24.1-hour dataset of coordinated humanoid behaviors and showed that policies trained on the virtual data could transfer zero-shot to a physical humanoid.

arXiv cs.ROAI-written

Posted Sep 22

No-Training Trick Lets Humanoid Robots Follow Text Commands

Researchers built a training-free pipeline called Sample-Simulate-Select (S³) that generates many candidate motions from a text prompt, tests each in physics simulation with a pretrained tracking policy, and keeps whichever one the robot actually executes best. On a Unitree G1, the method lifted upright-execution success from 83.5% to 89.5% on a test set, and all 177 selected motions ran successfully on real hardware.

arXiv cs.ROAI-written

Posted Sep 22

SafeLoop Wraps Robot AI Models With a Built-In Safety Net

A new research paper introduces SafeLoop, an external safety wrapper for vision-language-action (VLA) robot manipulation models that predicts hazards from vision and proprioception and triggers checkpoints or rollbacks. Tested on 24 LIBERO simulation tasks and three real-robot tasks, it cut hazard cases by roughly 70% while preserving task success rates.

arXiv cs.ROAI-written

Posted Sep 18

Robot Coding Agents Ignore Safety Rules Until Given a Harness

Researchers testing 'coding agents' — language models that write a robot's control program directly — found the agents complete manipulation tasks but crash into obstacles they were explicitly told to avoid in most trials. Their new framework, SafeHarness, fixes this by adding obstacle-aware route planning and contact execution, roughly doubling collision avoidance.

arXiv cs.ROAI-written

Posted Sep 18

GeoAAC Teaches Robot Policies When to Look Further Ahead

Researchers have proposed GeoAAC, a technique that lets Vision-Language-Action (VLA) robot policies dynamically adjust their 'action horizon' — how many future steps they commit to before re-checking — based on how confident the model's underlying prediction process appears to be. Tested on GR00T N1.5 and π0.5 policies across several manipulation benchmarks and real robot tasks, the method boosted real-world task success from 53.3% to 74.4% and improved simulation results by up to 8.7 percentage points over fixed-horizon baselines.

arXiv cs.ROAI-written

Posted Sep 18

Agile-WAM Speeds Up Tactile World Models for Robot Manipulation

Researchers have introduced Agile-WAM, a tactile 'world action model' that predicts future visual and tactile states alongside robot actions without relying on heavy pretrained generative backbones. In real-world tests across five contact-rich manipulation tasks, it improved success rates by 29.4% over the strongest baseline while running inference in just 11.9 milliseconds.