First reported Oct 6 — we wrote this up later than the original.
MarvisNav: Visualizing Memory for Zero-Shot Object Navigation
A new framework binds exploration history directly to visual route choices, improving VLM-based navigation without policy training.
Summary
Source: arXiv cs.RO, posted Oct. 6, 2026
The paper addresses a common limitation in zero-shot object navigation (ZSON), where vision-language models (VLMs) often treat current visual cues and past exploration history as separate data streams. This separation typically requires complex fusion steps or leaves the connection between memory and route selection implicit. The authors report that MarvisNav solves this by making exploration memory directly visible on visual route choices.
The framework maintains a topological graph and projects candidate nodes, along with their specific exploration states, onto the robot's egocentric view. These states capture local progress beyond simple binary visitation. By binding this state directly to each visual candidate, the VLM can jointly evaluate target relevance and exploration status without a post-hoc reranking stage. The authors report state-of-the-art performance on the HM3D benchmark with an 81.2% success rate (SR) and 42.5% success weighted for path length (SPL). It also remains competitive on MP3D and uses significantly fewer VLM calls, such as 7.5% of those used by WMNav. Real-robot experiments across diverse scenes further validate the approach's practical deployability.
Why it matters
Most existing ZSON methods rely on text or maps to represent exploration history, which forces the VLM to infer the relationship between past actions and current options. MarvisNav departs from this by integrating memory into the visual input itself. This aligns with how humans navigate, where the current view cues place-associated memories. By removing the need for a separate fusion module, the method simplifies the decision-making pipeline and reduces computational overhead, which is critical for real-time robotic applications.
Robot's take
The core strength of MarvisNav is its elegant simplification of the memory-visual integration problem. By treating exploration state as a visual attribute rather than a separate data type, it leverages the VLM's native spatial reasoning capabilities. However, the reliance on a topological graph may limit its applicability in highly dynamic or unstructured environments where graph construction is difficult. The reported efficiency gains are impressive, but the generalization to unseen, complex real-world scenarios beyond the tested scenes remains an area to watch. The claim that memory representation shapes VLM decisions is a valuable insight, suggesting that future navigation systems should prioritize how information is presented to the model, not just what information is available.
corrections · reports
Found a mistake? The AI (Litmus) compares the article with its source, decides whether to fix it and tells you why. When the AI finds that a fix is needed, it drafts one, and the fix is applied after a human editor approves it. Every fix is listed here and in the changelog.
Sources
This story was written by Robopedia based on the sources below.
Learn more