Robotics research papers, summarized by AI. Up to 2 a day also appear in the main feed. Summaries are written by AI from each paper's arXiv abstract page (title, authors, abstract); the PDF and figures are not reproduced. 74 total.
The authors present SNAP, a self-supervised model that uses Novel View Synthesis to learn robust 3D geometric representations. By restricting the decoder's expressivity and using latent-space targets, SNAP outperforms standard methods across five robotics and vision tasks.
The authors introduce EyeRobot 2.0, a system that performs fine-grained bimanual manipulation using a single stereo camera. By actively swiveling its viewpoint to center on 3D fixation points and allocating more computational resources to the image center, the system compensates for the lack of wrist cameras. In real-world trials, this approach outperformed passive stereo baselines by 40% and matched or exceeded ego-plus-wrist camera policies, particularly when objects occluded the wrist views.
Researchers introduce RPG, a framework that improves robot manipulation capabilities by generating practice tasks in simulation and refining symbolic skills without updating model weights. The system achieves 95.0% success on held-out tasks after 15 practice rounds and succeeds in all 30 physical trials.
Researchers introduce DuoMind, a distributed framework that coordinates multiple robots through semantic communication. The system pairs VLM-based orchestrators with VLA-based action models, and is evaluated on a new benchmark called RoboPoly.
Researchers introduce HumanoidToolBench, an 18-task benchmark for evaluating humanoid tool use, alongside ToolBook, a dataset of 3.1k demonstrations. The study reveals significant gaps between tool selection and task completion across seven simulation and three real-robot policies.
Researchers introduce PrefPI, an iterative method that steers pretrained robot policies using only relative preferences over self-generated trajectories. The approach aims to push behaviors beyond the initial policy's effective support, where desired actions are rarely observed. In real-world tests, the method increased object transport height from 10.7 cm to 19.8 cm using just 150 preference-labeled trajectories.
The authors introduce Dream4ACT, a world model that standardizes action representations for multi-embodiment video-action modeling by rendering joint configurations as images from virtual cameras. The system achieves an 88.98% success rate on RoboTwin 2.0 and a 65.66 score on TriWorldBench.
Researchers introduce TacEx, a method that directs robot exploration toward tactile feedback to learn manipulation skills more efficiently. The approach helps robots discover contact dynamics without task rewards or expert demonstrations.
Researchers introduced ASENA, a system where general-purpose coding agents program robot behavior without retraining model weights. The approach integrates a 4B-parameter navigation policy, achieving top-tier results on standard benchmarks and improving performance through iterative self-correction.
Researchers propose DSDyn-VLA, a dual-stream system that combines motion-aware planning with real-time correction to improve robot manipulation of moving objects. The authors report significant performance gains over existing methods in both simulation and real-world dynamic settings.
Researchers introduce LocoWM, a control framework that combines a base locomotion policy with a world model to predict future physical states. A residual adapter then generates corrective actions based on these predictions to improve precision and robustness.
Researchers propose a 'Universal Adversarial Object'—a sphere with optimized surface texture—that degrades the performance of Vision-Language-Action (VLA) models. The authors report that this object reduces average task success rates by 31.2% to 39.9% for two representative models, Pi0 and RDT, in both simulated and real-world settings.
Researchers analyzed how vision-language-action models behave under camera faults like blackouts and freezing. They found that while task success rates are similarly low, the physical failure modes differ significantly, with freezing causing extreme joint behavior and blackouts leading to object drops.
Researchers introduce SteerQuant, a 4-bit quantization method for world-action models that prioritizes numerical accuracy where it matters most for final robot actions. The approach achieves up to 2.23x speedup in simulation and 1.35x on a real dual-arm robot while maintaining task success rates close to full precision.
The Rho Team has released Rho, a family of open-weights Vision-Language-Action (VLA) models designed for bimanual manipulation. The authors report that Rho variants for YAM Box, UR AI Trainer, and FR3 Duo match or outperform existing open-weights baselines in both simulation and physical robot experiments. The models also demonstrate an online adaptation capability, requiring as few as 15 corrected episodes to handle tasks at the fringe of their training distribution.
Researchers introduce ReWAM, a world-action model that leverages pre-trained DINO features to enhance robot policy learning. The model achieves a 93.6% success rate on RoboTwin 2.0 and an 8.28% success rate on RoboDojo, demonstrating strong performance without relying on generative video pre-training.
The authors introduce CrossBFM, a framework that creates a shared latent behavior space for multiple humanoid embodiments. By using a unified encoder, the method reduces training costs from hundreds of GPU-hours to less than one, while enabling cross-robot transfer of motion tracking, goal reaching, and reward optimization.
The authors introduce WorldLine, a visual simulator that predicts robot manipulation outcomes by learning dynamics from over 10,000 hours of action-free videos and grounding them with 2,000 hours of action trajectories across ten embodiments. It achieves 74% mean accuracy in predicting trajectory success and improves task success by up to 21.4 percentage points over direct policy execution in out-of-domain settings.
The authors introduce EVO-WAM, a method that improves robot policies on new tasks without collecting additional expert demonstrations. By using a vision-language model and an inverse dynamics model to verify the consistency of generated video-action pairs, the framework iteratively refines the underlying world action model. Reported results show significant success rate improvements on both simulated RoboTwin 2.0 tasks and real-world composite tasks.
Researchers introduced EgoAlign, a data-construction framework that transforms egocentric human demonstrations into supervision signals for humanoid robots. The method uses simulator feedback to align human motion with robot capabilities, enabling zero-shot deployment on physical hardware for tasks like object relocation.
Researchers present DexRoam, a system that learns mobile bimanual dexterous manipulation from egocentric human demonstrations. By using a tracker-free capture setup and three alignment stages, the method preserves fine-grained whole-body motion, improving average success rates from 29% to 56% on GR00T N1.7 and from 32% to 57% on pi0.5.
Researchers present RoboCompiler, a graph-native framework that compiles robot mechanism graphs into a shared mechanical interface for consistent modeling, control, and simulation. The authors report significant speedups in evaluation time for complex robots like the Kangaroo.
Researchers introduce Holo-M, a discrete Vision-Language-Action (VLA) model designed for humanoid loco-manipulation. It uses a unified action tokenizer and grouped discrete diffusion decoding to handle the high-dimensional action space of humanoids, achieving top success rates on the SIMPLE benchmark.
Researchers present MM-ABC, a foundation model designed to improve mobile manipulation by coordinating arm and base actions. The system achieves high success rates across multiple simulation benchmarks and real-world tasks.
The authors introduce CollisionSplatting, a GPU-accelerated distance metric that operates directly on 3D Gaussian Splatting scenes to enable collision-aware motion planning. By combining this metric with learned image-conditioned rewards, the system achieves joint geometric and visual planning with high throughput and low memory usage.
Researchers propose F4R, a closed-loop system that diagnoses real robot failures, reconstructs them as simulation environments, and uses them to refine vision-language-action policies. The authors report 90.0% out-of-distribution success on four manipulation tasks, outperforming a baseline by 18.75 percentage points without collecting new real-world demonstrations.
Researchers introduce AD-WM, a world model designed for model predictive control that explicitly preserves differences between candidate actions rather than only minimizing prediction error. The method sharply improves simulated manipulation success rates and, paired with a frozen V-JEPA 2 encoder, boosts zero-shot pick-and-place performance on a real Franka robot from 42.2% to 71.1%.
Researchers have introduced RAPID, a system that generates, verifies, and refines robot programs from just one visual human demonstration. Tested on contact-rich manipulation tasks in simulation and on a real Franka arm, it generalized across object pose, shape, material, and environment.
Researchers have introduced Rolling-WAM, a World Action Model that denoises robot action and video predictions in a rolling, staggered fashion rather than all at once. Tested on LIBERO, RoboTwin, and a real Unitree G1 humanoid, it matches standard approaches on manipulation performance while cutting steady-state replanning time by 4.5x.
Researchers have introduced Underwater C3-JEPA, a predictive world model that lets remotely operated vehicles (ROVs) anticipate how objects respond during heavy-load underwater salvage. Using only synchronized multi-camera video and vehicle control data, the system estimates object states in latent space, which could support planning and training methods like model-predictive control.