- 📚 Preliminaries
- 🧠 Context Modeling
- 🔮 Imaginative Prediction
- 🎯 Decision Planning
⚠️ Reflection from Error- 🌐 Broader Navigation Literature
- 🚀 Future Directions
| Year | Paper | Venue | Resources |
|---|---|---|---|
| 2026 | ActiveVLN: Towards Active Exploration via Multi-Turn RL in Vision-and-Language Navigation | 2026 ICRA | |
| 2026 | StreamVLN: Streaming Vision-and-Language Navigation via SlowFast Context Modeling | 2026 ICRA | |
| 2025 | NaVILA: Legged Robot Vision-Language-Action Model for Navigation | 2025 RSS | |
| 2021 | VLN-BERT: A Recurrent Vision-and-Language BERT for Navigation | 2021 CVPR |
| Year | Paper | Venue | Resources |
|---|---|---|---|
| 2026 | Structured Observation Language for Efficient and Generalizable Vision-Language Navigation | 2026 arXiv | — |
| 2025 | Embodied Navigation Foundation Model | 2025 arXiv | |
| 2025 | Uni-NaVid: A Video-based Vision-Language-Action Model for Unifying Embodied Navigation Tasks | 2025 RSS |
| Year | Paper | Venue | Resources |
|---|---|---|---|
| 2026 | MapDream: Task-Driven Map Learning for Vision-Language Navigation | 2026 arXiv | — |
| 2026 | MonoDream: Monocular Vision-Language Navigation with Panoramic Dreaming | 2026 AAAI | |
| 2025 | PanoGen++: Domain-Adapted Text-Guided Panoramic Environment Generation for Vision-and-Language Navigation | 2025 Neural Networks | — |
| 2024 | Imagine Before Go: Self-Supervised Generative Map for Object Goal Navigation | 2024 CVPR | |
| 2023 | PanoGen: Text-Conditioned Panoramic Environment Generation for Vision-and-Language Navigation | 2023 NeurIPS |
| Year | Paper | Venue | Resources |
|---|---|---|---|
| 2026 | WNM-3D: A World Navigation Model with 3D Scene Conditioning for Closed-Loop VLN | 2026 arXiv | — |
| 2026 | NavDreamer: Video Models as Zero-Shot 3D Navigators | 2026 arXiv | — |
| 2026 | Sparse Video Generation Propels Real-World Beyond-the-View Vision-Language Navigation | 2026 arXiv | |
| 2025 | AstraNav-World: World Model for Foresight Control and Consistency | 2025 arXiv | — |
| 2025 | VISTAv2: World Imagination for Indoor Vision-and-Language Navigation | 2025 arXiv | — |
| 2025 | DreamNav: A Trajectory-Based Imaginative Framework for Zero-Shot Vision-and-Language Navigation | 2025 arXiv | — |
| 2025 | VISTA: Generative Visual Imagination for Vision-and-Language Navigation | 2025 arXiv | — |
| 2025 | Navigation World Models | 2025 CVPR | — |
| 2025 | Do Visual Imaginations Improve Vision-and-Language Navigation Agents? | 2025 CVPR | — |
| 2021 | PathDreamer: A World Model for Indoor Navigation | 2021 ICCV | — |
| Year | Paper | Venue | Resources |
|---|---|---|---|
| 2025 | GC-VLN: Instruction as graph constraints for training-free vision-and-language navigation | 2025 arXiv | — |
| 2025 | Constraint-aware zero-shot vision-language navigation in continuous environments | 2025 TPAMI | — |
| 2024 | Boosting efficient reinforcement learning for vision-and-language navigation with open-sourced llm | 2024 RA-L | — |
| Year | Paper | Venue | Resources |
|---|---|---|---|
| 2026 | VLN-NF: Feasibility-Aware Vision-and-Language Navigation with False-Premise Instructions | 2026 arXiv | |
| 2024 | Mind the error! detection and localization of instruction errors in vision-and-language navigation | 2024 IROS |
- Covered by papers already listed above.
| Year | Paper | Venue | Resources |
|---|---|---|---|
| 2026 | Where Did It Go Wrong? Capability-Oriented Failure Attribution for Vision-and-Language Navigation Agents | 2026 arXiv | — |
- Covered by papers already listed above.
| Year | Paper | Venue | Resources |
|---|---|---|---|
| 2026 | SeqWalker: Sequential-Horizon Vision-and-Language Navigation with Hierarchical Planning | 2026 arXiv | — |
| 2026 | EvolveNav: Empowering LLM-Based Vision-Language Navigation via Self-Improving Embodied Reasoning | 2026 TPAMI |
| Year | Paper | Venue | Resources |
|---|---|---|---|
| 2024 | Vision-Language Navigation with Energy-Based Policy | 2024 NeurIPS | — |
| Year | Paper | Venue | Resources |
|---|---|---|---|
| 2025 | MemoryVLA: Perceptual-Cognitive Memory in Vision-Language-Action Models for Robotic Manipulation | 2025 arXiv | — |
| 2024 | π0: A Vision-Language-Action Flow Model for General Robot Control | 2024 arXiv | — |
| 2024 | OpenVLA: An Open-Source Vision-Language-Action Model | 2024 arXiv |
Representative outlook papers explicitly discussed in the manuscript are listed below.
- Action grounding beyond static image-text reasoning.
- Reasoning grounded in predicted visual futures.
- Tighter coupling between geometry and policy learning.