Robotics research papers, summarized by AI. Up to 2 a day also appear in the main feed. Summaries are written by AI from each paper's arXiv abstract page (title, authors, abstract); the PDF and figures are not reproduced. 80 total.
A new research paper introduces SafeLoop, an external safety wrapper for vision-language-action (VLA) robot manipulation models that predicts hazards from vision and proprioception and triggers checkpoints or rollbacks. Tested on 24 LIBERO simulation tasks and three real-robot tasks, it cut hazard cases by roughly 70% while preserving task success rates.
Researchers testing 'coding agents' — language models that write a robot's control program directly — found the agents complete manipulation tasks but crash into obstacles they were explicitly told to avoid in most trials. Their new framework, SafeHarness, fixes this by adding obstacle-aware route planning and contact execution, roughly doubling collision avoidance.
Researchers have proposed StageGuard, a framework that uses a distilled vision-language model to decide when a robot should stop one skill and move to the next step in a multi-stage task. The system was tested on the BEHAVIOR-1K benchmark and on real robots, showing notable gains in accuracy and speed over prior methods.
Researchers have proposed GeoAAC, a technique that lets Vision-Language-Action (VLA) robot policies dynamically adjust their 'action horizon' — how many future steps they commit to before re-checking — based on how confident the model's underlying prediction process appears to be. Tested on GR00T N1.5 and π0.5 policies across several manipulation benchmarks and real robot tasks, the method boosted real-world task success from 53.3% to 74.4% and improved simulation results by up to 8.7 percentage points over fixed-horizon baselines.
Researchers have introduced Agile-WAM, a tactile 'world action model' that predicts future visual and tactile states alongside robot actions without relying on heavy pretrained generative backbones. In real-world tests across five contact-rich manipulation tasks, it improved success rates by 29.4% over the strongest baseline while running inference in just 11.9 milliseconds.
Researchers introduce MoWAM, a world action model that replaces expensive future-video generation with explicit future motion prediction, using a Mixture-of-Transformer architecture. Tests on LIBERO, LIBERO-Plus, and real-world manipulation tasks show it matches or beats existing world action model baselines while running more efficiently.
Researchers have introduced HOPHY, a hierarchical hypergraph-based terrain representation for off-road path and mission planning that matches near-optimal pixel-based planning accuracy while dramatically cutting computation time. Tested on real kilometer-scale maps and a physical Clearpath Jackal robot, it achieved 100% planning success and up to 79x lower computation in multi-robot task allocation.
Researchers have introduced FunArt, a system that builds articulation-aware 3D scene graphs from a single static RGB-D scan, identifying movable parts and functional elements like handles while estimating how they move. Tested on the Articulate3D dataset, it beats prior state-of-the-art methods across multiple metrics.
Researchers introduced Movement Trend Guidance, a method that gives 3D diffusion policies a sense of foresight without requiring an explicit future plan. Tested on RoboTwin2.0, LIBERO-40, and real-robot tasks, it consistently outperformed the baseline DP3 policy while adding only 3.52% more parameters.
Researchers have introduced HIL-UMI, a method that improves vision-language-action (VLA) robot policies through human feedback without needing a physical robot to execute the policy during training. By comparing human and policy-predicted trajectories on handheld demonstration data, the system targets its most useful corrections and outperformed a standard interactive baseline on one benchmark task.
Researchers have unveiled DexTouch-WM, an action-conditioned world model that predicts future camera views and touch signals for dexterous robots by learning from human touch data. Using matching flexible tactile sensor arrays on human and robot hands, the team showed that adding up to 100 hours of human interaction data—while keeping robot data fixed at 5 hours—significantly improved the model's predictions on unseen robot tasks.
Researchers have introduced Sampling-Guided Policy Search (SGPS), a training method that combines sampling-based model predictive control with fast first-order policy gradients to teach vision-based locomotion and manipulation skills more reliably. Tested in simulation on Unitree Go2 and G1 robots and deployed zero-shot to a real Go2, the approach let the robot autonomously trot, crawl, and clear obstacles using only onboard depth images.
A team of researchers has published a comprehensive survey on space mining robotics, proposing a six-stage framework covering everything from remote sensing to resource extraction. The paper, posted to arXiv, also catalogs mission data, analog datasets, and simulation tools while outlining open challenges for building an off-world economy.
Researchers describe ViTacPhys, a visual-tactile learning framework that estimates an object's mass, friction, and stiffness from human manipulation demonstrations, then uses those estimates to adapt robot grasping. Tested on 60 rigid and deformable objects, the system reports strong accuracy on seen items and solid generalization to new ones, reaching up to 95% grasping success in robot trials.
Researchers have proposed Anatomy-Informed Neural Networks (AINN), a framework that embeds anatomical rules directly into a model's loss function and architecture so it cannot produce anatomically impossible predictions. The approach is demonstrated on a data-scarce clinical problem — how the aortoiliac artery tree deforms when a stiff guidewire is inserted — using an SE(3)-based mechanical model rather than a trained network, as a step toward autonomous endovascular navigation.
Researchers unveiled NeSAM, a framework combining differentiable soil-mechanics equations with a Transformer-based correction model to predict off-road vehicle motion. Tested in simulation and on a physical robot, it improved prediction accuracy by up to 30% and cut trajectory-tracking error by 69.4%.
Researchers have formalized a new routing problem called the Steiner Traveling Salesman Problem on Graphs of Convex Sets, which models robot navigation through required and optional regions. Their branch-and-bound search algorithm found feasible solutions on all tested benchmark cases within 30 seconds, far outperforming two existing baseline methods that only succeeded on about half the instances.
Researchers have introduced VT-MUSE, a representation-learning framework that jointly models vision and touch across time to improve robotic manipulation. The method outperformed the strongest baseline by 11 percentage points on a simulation benchmark and showed gains in real-world tests.
A new arXiv preprint presents two localization frameworks for autonomous surface vessels that rely on coastal geometry instead of GPS. One uses LiDAR to track vessel motion and match shorelines to satellite maps, while the other uses only a monocular camera and semantic segmentation to achieve similar results with bounded drift.
Researchers have proposed Action-JND, a technique that decides which visual tokens a robot's AI brain can safely drop or reuse without altering its actions. Tested on the LIBERO benchmark with OpenVLA and OpenVLA-OFT, it improved compression reliability, especially at aggressive compression ratios, according to a new arXiv preprint.