Robotics research papers, summarized by AI. Up to 2 a day also appear in the main feed. Summaries are written by AI from each paper's arXiv abstract page (title, authors, abstract); the PDF and figures are not reproduced. 75 total.
Researchers have introduced Underwater C3-JEPA, a predictive world model that lets remotely operated vehicles (ROVs) anticipate how objects respond during heavy-load underwater salvage. Using only synchronized multi-camera video and vehicle control data, the system estimates object states in latent space, which could support planning and training methods like model-predictive control.
A new dataset called Ego-Exo4D-HM extends the existing Ego-Exo4D collection by providing dense 4D human motion reconstructions instead of sparse pose annotations. The release includes the full reconstruction pipeline, code, and documentation for researchers working on skill learning, activity understanding, and embodied AI.
Researchers have proposed Capability-Tradeoff Contact Selection (CTCS), a method that helps legged manipulator robots choose which surface to brace against while performing a task, weighing stability gains against lost mobility. Tested on a Unitree Go2 quadruped fitted with an AgileX NERO arm across 392 task conditions and nine support surfaces, CTCS matched the performance of exhaustive search while running about three times faster.
Researchers have proposed Behavior Predictive Control (BPC), a robot control method that synthesizes actions by combining stored demonstrations rather than training a neural network policy. In tests, BPC matched or beat a learned policy called π_0.5 while cutting policy-fitting time from hours to seconds.
Researchers have introduced the Riemannian MeanFlow Policy (RMFP), a visuomotor control method that generates robot action sequences in as little as one network evaluation instead of many diffusion-style integration steps. Tested on simulated benchmarks and a real robotic manipulation task, RMFP matches prior methods' performance while cutting computational cost, according to a paper posted on arXiv.
Researchers have introduced Self-Adaptive VLA, a post-training technique that lets vision-language-action robot policies detect and compensate for hardware shifts like actuation bias or encoder offsets during deployment. Tested on four precision manipulation tasks, the method recovered over 80% of the base policy's performance despite injected hardware errors.
Researchers have developed a control framework that continuously recalculates how a robotic hand should distribute grip force across all its contact points, rather than relying on a one-time calculation made before the grasp begins. Tested on simulated grasps and a real 27-degree-of-freedom arm-hand system, the method helped robots retain and recover grasps when objects moved or were disturbed by a person's hand. The work was described in a paper posted to arXiv.
Researchers show that visual world model planners which score actions by distance to a final goal image can fail even with perfect dynamics, because reaching a goal sometimes requires moving away from it first. Their method, Anchored Planning, retrieves observed intermediate targets from past experience and outperforms a released planner on every tested task without any retraining.
Researchers propose 'body-grounded replanning,' a system where a large language model monitors a robot's internal physical state—like joint load and mobility limits—to choose better manipulation strategies on the fly. Tested on reaching and contact-rich tasks in simulation and on real hardware, the approach cut physical effort while keeping success rates high.
Researchers have introduced Res-HIL, a human-in-the-loop reinforcement learning framework that refines a frozen imitation-learned policy with residual corrections instead of retraining a robot's entire behavior from scratch. Using only 20 initial demonstrations, the method reportedly beat existing full-policy human-in-the-loop approaches and outperformed imitation policies trained on five times more demonstrations after just ten minutes of online training.
Researchers have introduced World Action Agent (WAA), a multi-agent system that lets vision-language models plan, preview, and correct robot manipulation actions before execution rather than just describing a scene. On the LIBERO-Pro benchmark, WAA reached 75.6% average success, and fine-tuning a smaller model on its interaction data boosted out-of-domain performance from 1.7% to 43.3%.
Researchers evaluated GPT-6-Astra on zero-shot vision-and-language navigation in continuous environments using only monocular RGB input. The model achieved a 79.0% success rate on the R2R-CE benchmark, outperforming previous zero-shot and supervised baselines by 13.0 and 6.9 percentage points, respectively.
The authors propose BeyondRetarget, an end-to-end system that maps monocular RGB videos directly to robot motions. By bypassing explicit human motion representation, the method aims to reduce error propagation and improve physical plausibility for humanoid robots.
Researchers have released PolyUMI, an open-source handheld gripper that captures synchronized vision, tactile, audio, and proprioceptive data during human demonstrations for robot manipulation. Paired with a new multimodal policy called VisTA, the system reportedly lets robots pick up contact information that cameras alone miss, matching or beating existing multimodal approaches on object inference, slip control, and contact-rich tasks.
Researchers present a learning-based framework that integrates terrain-aware target selection with reinforcement and imitation learning for continuous autonomous excavation. The system, deployed on a scaled hydraulic excavator, achieves a mean payload of 6.52 kg per cycle, significantly outperforming a fixed-dig baseline at 2.68 kg.
Researchers have proposed LiMA, an asynchronous dual-system AI framework that decouples slow, long-horizon intent planning from fast, high-frequency motion control in robotic manipulation. The approach cuts inference latency by 45.8% compared to a baseline called Cosmos-Policy while achieving a 70.8% overall success rate across six bimanual dexterous tasks.
Researchers introduce ARMS, a streaming robot policy built on a pretrained π0.5 backbone that lets a dual-arm robot continuously watch for new instructions, recall its own past actions, and act without pausing. In tests, ARMS scored 45% on a combined task versus 28% for the best baseline, with ablations confirming each added module matters.
Researchers propose Context-Continuous Preference Learning (CCPL) to personalize exoskeleton assistance by assuming user preferences vary smoothly across operating conditions. In simulations and retrospective human data, CCPL improved reconstruction accuracy and reduced the data budget compared to independent learning.
Researchers have introduced a world-model-based framework for robotic insertion that generalizes to unseen parts, reaching 56% zero-shot success versus 7% for a model-free baseline. The system, trained on up to 90 diverse insertion tasks, improves as more objects are added and can be fine-tuned efficiently on new parts.
Researchers have introduced COMPASS, a decentralized control architecture that lets large language models coordinate flocks of robots without collapsing as team size grows. By generating feedback locally through a spatial transformer that compresses multi-hop fleet communication into a compact learned token, the system kept flocking formations cohesive while scaling to 1,024 robots under natural-language commands.
Researchers propose ε4P, a technique that repurposes discarded low-precision task data and high-precision data from unrelated tasks to train vision-language-action (VLA) models for precise robotic manipulation. Real-robot tests show performance gains of up to 31.7 percentage points and the ability to replace costly task-specific data with only a 4.2 percentage point average drop.
Researchers have proposed a hierarchical control framework that combines long-horizon geometric planning with data-driven Koopman models to automate the repetitive forward-reverse 'V-cycle' maneuvers of wheel loaders. The system runs a real-time model predictive controller within a 50-millisecond loop and was validated in high-fidelity simulation using Algoryx Dynamics.
Researchers have introduced MATE, a multi-agent virtual teleoperation platform that lets geographically separated operators simultaneously control whole-body humanoid robots in a shared physics-based environment. The team used it to build a 24.1-hour dataset of coordinated humanoid behaviors and showed that policies trained on the virtual data could transfer zero-shot to a physical humanoid.
Researchers have introduced Dr-LiSA, described as the first direct method for localizing 2D spinning radar scans against 3D lidar maps in full six-degree-of-freedom SE(3) pose space. Tested on more than 90 km of on-road driving data, the approach beats prior radar-lidar localization methods and matches leading radar-only systems in planar accuracy.
Researchers built a training-free pipeline called Sample-Simulate-Select (S³) that generates many candidate motions from a text prompt, tests each in physics simulation with a pretrained tracking policy, and keeps whichever one the robot actually executes best. On a Unitree G1, the method lifted upright-execution success from 83.5% to 89.5% on a test set, and all 177 selected motions ran successfully on real hardware.
A new research paper introduces SafeLoop, an external safety wrapper for vision-language-action (VLA) robot manipulation models that predicts hazards from vision and proprioception and triggers checkpoints or rollbacks. Tested on 24 LIBERO simulation tasks and three real-robot tasks, it cut hazard cases by roughly 70% while preserving task success rates.
Researchers testing 'coding agents' — language models that write a robot's control program directly — found the agents complete manipulation tasks but crash into obstacles they were explicitly told to avoid in most trials. Their new framework, SafeHarness, fixes this by adding obstacle-aware route planning and contact execution, roughly doubling collision avoidance.
Researchers have proposed StageGuard, a framework that uses a distilled vision-language model to decide when a robot should stop one skill and move to the next step in a multi-stage task. The system was tested on the BEHAVIOR-1K benchmark and on real robots, showing notable gains in accuracy and speed over prior methods.
Researchers have proposed GeoAAC, a technique that lets Vision-Language-Action (VLA) robot policies dynamically adjust their 'action horizon' — how many future steps they commit to before re-checking — based on how confident the model's underlying prediction process appears to be. Tested on GR00T N1.5 and π0.5 policies across several manipulation benchmarks and real robot tasks, the method boosted real-world task success from 53.3% to 74.4% and improved simulation results by up to 8.7 percentage points over fixed-horizon baselines.
Researchers have introduced Agile-WAM, a tactile 'world action model' that predicts future visual and tactile states alongside robot actions without relying on heavy pretrained generative backbones. In real-world tests across five contact-rich manipulation tasks, it improved success rates by 29.4% over the strongest baseline while running inference in just 11.9 milliseconds.