Skip to content

Latest commit

History

27 Commits

Folders and files

NameName
Last commit message
Last commit date

Repository files navigation

🧭 Awesome Visual Language Navigation

Awesome list badgeBibTeXPaper

🌟 Overview

📚 Preliminaries

Simulation, Data and Control

Toward More Realistic Simulation

YearPaperVenueResources
2026Image2Sim: Scaling Embodied Navigation via Generative Neural Simulator2026 arXivCodeData
2026Habitat-GS: A High-Fidelity Navigation Simulator with Dynamic Gaussian Splatting2026 arXivProjectCode
2026NavGSim: High-Fidelity Gaussian Splatting Simulator for Large-Scale Navigation2026 ICRACode
2025VLNVerse: A Benchmark for Vision-Language Navigation with Versatile, Embodied, Realistic Simulation and Evaluation2025 arXivProjectCode
2025Towards Physically Executable 3D Gaussian for Embodied Navigation2025 arXiv
2025Rethinking the embodied gap in vision-and-language navigation: A holistic study of physical and visual disparities2025 ICCVProjectCode
2023Habitat 3.0: A co-habitat for humans, avatars and robots2023 arXiv
2021Habitat-Matterport 3D Dataset (HM3D): 1000 Large-scale 3D Environments for Embodied AI2021 NeurIPS D&B
2020Beyond the nav-graph: Vision-and-language navigation in continuous environments2020 ECCV
2019PyRobot: An Open-source Robotics Framework for Research and Benchmarking2019 arXivCode
2019Habitat: A Platform for Embodied AI Research2019 ICCVCode
2018Gibson env: Real-world perception for embodied agents2018 CVPR
2017AI2-THOR: An Interactive 3D Environment for Visual AI2017 arXivProject
2017Matterport3D: Learning from RGB-D Data in Indoor Environments2017 3DVCode

Towards More Diverse Datasets&Benchmarks

YearPaperVenueResources
2026HumanoidVLN: A Physics-Grounded Simulator and Benchmark for Vision-Language Navigation Across Diverse Humanoid Embodiments2026 arXivProject
2026DaViNCi: A Dataset Towards Outdoor Vision-and-Language Navigation with Continuous Actions and Dynamic Elements2026 arXivProject
2026Beyond Isolation: A Unified Benchmark for General-Purpose Navigation2026 RSSProjectCode
2026MiniVLA-Nav v1: A Multi-Scene Simulation Dataset for Language-Conditioned Robot Navigation2026 arXiv
2026Explore with Long-term Memory: A Benchmark and Multimodal LLM-based Reinforcement Learning Framework for Embodied Exploration2026 arXiv
2026CoNavBench: Collaborative Long-Horizon Vision-Language Navigation Benchmark2026 ICLRProject
2026Rule-VLN: Bridging Perception and Compliance via Semantic Reasoning and Geometric Rectification2026 arXiv
2026AirNav: A Large-Scale Real-World UAV Vision-and-Language Navigation Dataset with Natural and Diverse Instructions2026 arXiv
2025IndoorUAV: Benchmarking Vision-Language UAV Navigation in Continuous Indoor Environments2025 arXiv
2025VL-LN Bench: Towards Long-horizon Goal-oriented Navigation with Active Dialogs2025 arXivCodeData
2025OpenFly: A Comprehensive Platform for Aerial Vision-Language Navigation2025 arXivCodeData
2025Collaborative Instance Object Navigation: Leveraging Uncertainty-Awareness to Minimize Human-Agent Dialogues2025 ICCVProjectCodeData
2025RoomTour3D: Geometry-Aware Video-Instruction Tuning for Embodied Navigation2025 CVPRCodeData
2025Towards long-horizon vision-language navigation: Platform, benchmark and method2025 CVPRCode
2025CityWalker: Learning Embodied Urban Navigation from Web-Scale Videos2025 CVPRCodeData
2025UrbanNav: Learning Language-Guided Urban Navigation from Web-Scale Human Trajectories2025 arXivCode
2025UAV-Flow Colosseo: A Real-World Benchmark for Flying-on-a-Word UAV Imitation Learning2025 arXiv
2024Embodiedcity: A benchmark platform for embodied agent in real-world city environment2024 arXivCode
2024InstruGen: Automatic Instruction Generation for Vision-and-Language Navigation Via Large Multimodal Models2024 arXivCode
2024Towards realistic uav vision-language navigation: Platform, benchmark, and methodology2024 arXivProjectCodeData
2024CityNav: Language-goal aerial navigation dataset with geographic information2024 arXivCode
2024Mind the error! detection and localization of instruction errors in vision-and-language navigation2024 IROSProjectCode
2024Hazard challenge: Embodied decision making in dynamically changing environments2024 arXivProject
2023AerialVLN: Vision-and-Language Navigation for UAVs2023 ICCVCode
2023Iterative vision-and-language navigation2023 CVPRCode
2023Scaling Data Generation in Vision-and-Language Navigation2023 ICCV
2022REVE-CE: Remote Embodied Visual Referring Expression in Continuous Environment2022 RA-L
2021Talk2Nav: Long-Range Vision-and-Language Navigation with Dual Attention and Spatial Memory2021 IJCV
2020Room-across-room: Multilingual vision-and-language navigation with dense spatiotemporal grounding2020 arXiv
2018Vision-and-language navigation: Interpreting visually-grounded navigation instructions in real environments2018 CVPR

Towards More Continuous Control

YearPaperVenueResources
2026FreqCache: Accelerating Embodied VLN Models with Adaptive Frequency-Guided Token Caching2026 arXiv
2026LiveVLN: Breaking the Stop-and-Go Loop in Vision-Language Navigation2026 arXivCode
2026VLN-Cache: Enabling Token Caching for VLN Models with Visual/Semantic Dynamics Awareness2026 arXiv
2026AutoFly: Vision-Language-Action Model for UAV Autonomous Navigation in the Wild2026 arXiv
2025Embodied Navigation Foundation Model2025 arXivProject
2022Bridging the gap between learning in discrete and continuous environments for vision-and-language navigation2022 CVPR
2021Hierarchical cross-modal agent for robotics vision-and-language navigation2021 ICRACode

🧠 Context Modeling

Context Perception

Spatial Perception

YearPaperVenueResources
2026GA-VLN: Geometry-Aware BEV Representation for Efficient Vision-Language Navigation2026 arXivCode
2026What Limits Vision-and-Language Navigation ?2026 arXivProjectCode
2026SpaAct: Spatially-Activated Transition Learning with Curriculum Adaptation for Vision-Language Navigation2026 arXiv
2026GIST: Multimodal Knowledge Extraction and Spatial Grounding via Intelligent Semantic Topology2026 arXiv
2026Instruction-as-State: Environment-Guided and State-Conditioned Semantic Understanding for Embodied Navigation2026 arXiv
2026DyGeoVLN: Infusing Dynamic Geometry Foundation Model into Vision-Language Navigation2026 arXiv
2026SPAN-Nav: Generalized Spatial Awareness for Versatile Vision-Language Navigation2026 arXiv
2026Spatial-VLN: Zero-Shot Vision-and-Language Navigation With Explicit Spatial Perception and Exploration2026 arXiv
2026JanusVLN: Decoupling Semantics and Spatiality with Dual Implicit Memory for Vision-Language Navigation2026 ICLRProjectCode
2025Efficient-VLN: A Training-Efficient Vision-Language Navigation Model2025 arXivProject
2025MG-Nav: Dual-Scale Visual Navigation via Sparse Spatial Memory2025 arXiv
2024NaVid: Video-based VLM Plans the Next Step for Vision-and-Language Navigation2024 RSSCode

Temporal Perception

YearPaperVenueResources
2026ActiveVLN: Towards Active Exploration via Multi-Turn RL in Vision-and-Language Navigation2026 ICRACode
2026StreamVLN: Streaming Vision-and-Language Navigation via SlowFast Context Modeling2026 ICRACodeData
2025NaVILA: Legged Robot Vision-Language-Action Model for Navigation2025 RSSCodeBench
2021VLN-BERT: A Recurrent Vision-and-Language BERT for Navigation2021 CVPRCode

Context Memory

Memory Representation

YearPaperVenueResources
2026MemVLN: Episodic and Procedural Memory for Vision-and-Language Navigation2026 arXiv
2026LASAR: Towards Spatio-temporal Reasoning with Latent Cognitive Map2026 CVPRCode
2026Explore Like Humans: Autonomous Exploration with Online SG-Memo Construction for Embodied Agents2026 arXiv
2026Hypothesis Graph Refinement: Hypothesis-Driven Exploration with Cascade Error Correction for Embodied Navigation2026 arXiv
2026HiMemVLN: Enhancing Reliability of Open-Source Zero-Shot Vision-and-Language Navigation with Hierarchical Memory System2026 arXiv
2026HaltNav: Reactive Visual Halting over Lightweight Topological Priors for Robust Vision-Language Navigation2026 arXiv
2026GSMem: 3D Gaussian Splatting as Persistent Spatial Memory for Zero-Shot Embodied Exploration and Reasoning2026 arXiv
2026Structured Observation Language for Efficient and Generalizable Vision-Language Navigation2026 arXiv
2025Open-Nav: Exploring Zero-Shot Vision-and-Language Navigation in Continuous Environment with Open-Source LLMs2025 ICRACode
2025MapNav: A Novel Memory Representation via Annotated Semantic Maps for VLM-based Vision-and-Language Navigation2025 ACL
2024MapGPT: Map-Guided Prompting with Adaptive Path Planning for Vision-and-Language Navigation2024 ACLCode
2022BEVBert: Multimodal Map Pre-training for Language-guided Navigation2022 arXiv
2021Structured Scene Memory for Vision-Language Navigation2021 CVPRCode

Memory Retrieval and Utilization

YearPaperVenueResources
2026LookStep: Efficient Vision-Language Navigation with Linguistic Foresight and Event Driven Memory2026 EMNLPCode
2026Dual-Anchoring: Addressing State Drift in Vision-Language Navigation2026 arXiv
2026Beyond Textual Knowledge-Leveraging Multimodal Knowledge Bases for Enhancing Vision-and-Language Navigation2026 arXiv
2026CMMR-VLN: Vision-and-Language Navigation via Continual Multimodal Memory Retrieval2026 arXiv
2026One Agent to Guide Them All: Empowering MLLMs for Vision-and-Language Navigation via Explicit World Representation2026 arXiv
2026SpatialNav: Leveraging Spatial Scene Graphs for Zero-Shot Vision-and-Language Navigation2026 arXiv
2025Dream to Recall: Imagination-Guided Experience Retrieval for Memory-Persistent Vision-and-Language Navigation2025 arXiv
2025FSR-VLN: Fast and Slow Reasoning for Vision-Language Navigation with Hierarchical Multi-modal Scene Graph2025 arXiv
2025OpenIN: Open-Vocabulary Instance-Oriented Navigation in Dynamic Domestic Environments2025 RA-L
2024OrionNav: Online Planning for Robot Autonomy with Context-Aware LLM and Open-Vocabulary Semantic Scene Graphs2024 arXiv

Context Efficiency

Reducing Temporal Redundancy

YearPaperVenueResources
2026DecoVLN: Decoupling Observation, Reasoning, and Correction for Vision-and-Language Navigation2026 arXiv
2025RANGER: A Monocular Zero-Shot Semantic Navigation Framework through Contextual Adaptation2025 arXiv
2025C-NAV: Towards self-evolving continual object navigation in open world2025 arXiv
2025AstraNav-Memory: Contexts Compression for Long Memory2025 arXiv
2025COSMO: Combination of Selective Memorization for Low-cost Vision-and-Language Navigation2025 ICCV

Reducing Spatial Redundancy

YearPaperVenueResources
2026Structured Observation Language for Efficient and Generalizable Vision-Language Navigation2026 arXiv
2025Embodied Navigation Foundation Model2025 arXivProject
2025Uni-NaVid: A Video-based Vision-Language-Action Model for Unifying Embodied Navigation Tasks2025 RSSCode

Computational Reuse and Acceleration

YearPaperVenueResources
2026FreqCache: Accelerating Embodied VLN Models with Adaptive Frequency-Guided Token Caching2026 arXiv
2026VLN-Cache: Enabling Token Caching for VLN Models with Visual/Semantic Dynamics Awareness2026 arXiv
2026History-Conditioned Spatio-Temporal Visual Token Pruning for Efficient Vision-Language Navigation2026 arXiv

🔮 Imaginative Prediction

Spatial Imagination

Occupancy Prediction

YearPaperVenueResources
2026SPAN-Nav: Generalized Spatial Awareness for Versatile Vision-Language Navigation2026 arXivCode
2026SparseOccVLA: Bridging Occupancy and Vision-Language Models via Sparse Queries for Unified 4D Scene Understanding and Planning2026 arXiv
2026SparseWorld: A Flexible, Adaptive, and Efficient 4D Occupancy World Model Powered by Sparse and Dynamic Queries2026 AAAIProjectCode
2025Occ-llm: Enhancing autonomous driving with occupancy-based large language models2025 ICRA
2024Occllama: An Occupancy-Language-Action Generative World Model for Autonomous Driving2024 arXiv
2023PEANUT: Predicting and Navigating to Unseen Targets2023 ICCVCode

View Completion

YearPaperVenueResources
2026MapDream: Task-Driven Map Learning for Vision-Language Navigation2026 arXiv
2026MonoDream: Monocular Vision-Language Navigation with Panoramic Dreaming2026 AAAIProject
2025PanoGen++: Domain-Adapted Text-Guided Panoramic Environment Generation for Vision-and-Language Navigation2025 Neural Networks
2024Imagine Before Go: Self-Supervised Generative Map for Object Goal Navigation2024 CVPRCode
2023PanoGen: Text-Conditioned Panoramic Environment Generation for Vision-and-Language Navigation2023 NeurIPSCode

3D Representation

YearPaperVenueResources
2026GSMem: 3D Gaussian Splatting as Persistent Spatial Memory for Zero-Shot Embodied Exploration and Reasoning2026 arXiv
2026ThinkMatter: Panoramic-Aware Instructional Semantics for Monocular Vision-and-Language Navigation2026 TIP
2025monoVLN: Bridging the Observation Gap between Monocular and Panoramic Vision and Language Navigation2025 ICCV
2024UnitedVLN: Generalizable Gaussian Splatting for Continuous Vision-Language Navigation2024 arXiv

Temporal Imagination

Pixel Level

YearPaperVenueResources
2026WNM-3D: A World Navigation Model with 3D Scene Conditioning for Closed-Loop VLN2026 arXiv
2026NavDreamer: Video Models as Zero-Shot 3D Navigators2026 arXiv
2026Sparse Video Generation Propels Real-World Beyond-the-View Vision-Language Navigation2026 arXivCode
2025AstraNav-World: World Model for Foresight Control and Consistency2025 arXiv
2025VISTAv2: World Imagination for Indoor Vision-and-Language Navigation2025 arXiv
2025DreamNav: A Trajectory-Based Imaginative Framework for Zero-Shot Vision-and-Language Navigation2025 arXiv
2025VISTA: Generative Visual Imagination for Vision-and-Language Navigation2025 arXiv
2025Navigation World Models2025 CVPR
2025Do Visual Imaginations Improve Vision-and-Language Navigation Agents?2025 CVPR
2021PathDreamer: A World Model for Indoor Navigation2021 ICCV

Feature Level

YearPaperVenueResources
2026Towards Dual-Brain Minimal Sufficient Representation for Vision-Language Navigation2026 arXiv
2026NavWM: A Unified Navigation World Model for Foresight-Driven Planning2026 ECCV
2026WAM-Nav: Asymmetric Latent World-Action Modeling for Unified Visual Navigation2026 arXiv
2026WorldVLN: Autoregressive World Action Model for Aerial Vision-Language Navigation2026 arXivProjectCode
2026LatentPilot: Scene-Aware Vision-and-Language Navigation by Dreaming Ahead with Latent Visual Reasoning2026 arXivProjectCode
2026PROSPECT: Unified Streaming Vision-Language Navigation via Semantic-Spatial Fusion and Latent Predictive Representation2026 arXiv
2026FantasyVLN: Unified Multimodal Chain-of-Thought Reasoning for Vision-Language Navigation2026 arXivProjectCode
2026MonoDream: Monocular Vision-Language Navigation with Panoramic Dreaming2026 AAAIProject
2025NavForesee: A Unified Vision-Language World Model for Hierarchical Planning and Dual-Horizon Navigation Prediction2025 arXiv
2025NavMorph: A Self-Evolving World Model for Vision-and-Language Navigation in Continuous Environments2025 NeurlPSCode

Inverse Dynamics Modeling and Causality

Inverse Dynamics

YearPaperVenueResources
2026FutureNav: Unified World-Action Modeling for Vision-and-Language Navigation2026 arXivCodeProject
2026SpaAct: Spatially-Activated Transition Learning with Curriculum Adaptation for Vision-Language Navigation2026 arXiv
2026Sparse Video Generation Propels Real-World Beyond-the-View Vision-Language Navigation2026 arXivCode
2026LatentPilot: Scene-Aware Vision-and-Language Navigation by Dreaming Ahead with Latent Visual Reasoning2026 arXivCodeProject
2026ImagiNav: Scalable Embodied Navigation via Generative Visual Prediction and Inverse Dynamics2026 arXiv
2026NaVIDA: Vision-Language Navigation with Inverse Dynamics Augmentation2026 arXivCode

Causal Learning

YearPaperVenueResources
2025CAST: Counterfactual Labels Improve Instruction Following in Vision-Language-Action Models2025 arXivCodeData
2025CVLN-Think: Causal Inference with Counterfactual Style Adaptation for Continuous Vision-and-Language Navigation2025 IROS
2024Causality-based cross-modal representation learning for vision-and-language navigation2024 arXiv
2024Vision-and-Language Navigation via Causal Learning2024 CVPRCode
2022Counterfactual Cycle-Consistent Learning for Instruction Following and Generation in Vision-Language Navigation2022 CVPRCode

🎯 Decision Planning

Structured Decision

Instruction Decomposition

YearPaperVenueResources
2025GC-VLN: Instruction as graph constraints for training-free vision-and-language navigation2025 arXiv
2025Constraint-aware zero-shot vision-language navigation in continuous environments2025 TPAMI
2024Boosting efficient reinforcement learning for vision-and-language navigation with open-sourced llm2024 RA-L

Sub-Goal Formulation

YearPaperVenueResources
2026Traj-VLN: Learning Pixel-Space Interaction via Autoregressive Trajectory Generation2026 arXiv
2026Goal2Pixel: Grounding Goals to Pixels for Vision-Language Navigation2026 arXivProject
2026Ground Slow, Move Fast: A Dual-System Foundation Model for Generalizable Vision-and-Language Navigation2026 ICLRProjectCode
2025Landmark-Guided Knowledge for Vision-and-Language Navigation2025 ICIC
2025SmartWay: Enhanced Waypoint Prediction and Backtracking for Zero-Shot Vision-and-Language Navigation2025 IROS
2025Affordances-oriented planning using foundation models for continuous vision-language navigation2025 AAAIProject
2025Hierarchical semantic-augmented navigation: Optimal transport and graph-driven reasoning for vision-language navigation2025 NeurIPS
2023GridMM: Grid Memory Map for Vision-and-Language Navigation2023 ICCVCode

Navigation Progress Estimation

YearPaperVenueResources
2026From Routes to Steps: Separating Semantic Progress from Local Execution in Vision-and-Language Navigation2026 arXivProject
2025Progress-Think: Semantic Progress Reasoning for Vision-Language Navigation2025 arXiv
2024PASTS: Progress-Aware Spatio-Temporal Transformer Speaker for Vision-and-Language Navigation2024 EAAI
2022One step at a time: Long-horizon vision-and-language navigation with milestones2022 CVPR
2020Vision-Language Navigation with Self-Supervised Auxiliary Reasoning Tasks2020 CVPR
2019The regretful agent: Heuristic-aided navigation through progress estimation2019 CVPRCode
2019Self-monitoring navigation agent via auxiliary progress estimation2019 ICLRCode

Reasoning-based Decision

LLM as External Reasoner

YearPaperVenueResources
2026Stop Wandering: Efficient Vision-Language Navigation via Metacognitive Reasoning2026 arXiv
2026TagaVLM: Topology-Aware Global Action Reasoning for Vision-Language Navigation2026 arXiv
2026ReasonNavi: Human-Inspired Global Map Reasoning for Zero-Shot Embodied Navigation2026 arXivProject
2026DV-VLN: Dual Verification for Reliable LLM-Based Vision-and-Language Navigation2026 arXiv
2026Towards Open Environments and Instructions: General Vision-Language Navigation via Fast-Slow Interactive Reasoning2026 arXiv
2025GeoNav: Empowering MLLMs with Explicit Geospatial Reasoning Abilities for Language-Goal Aerial Navigation2025 arXiv
2025SpatialGPT: Zero-Shot Vision-and-Language Navigation via Spatial CoT over Structured Spatial Memory2025 SIGSPATIAL
2024End-to-end navigation with vision language models: Transforming spatial reasoning into question-answering2024 arXiv
2024NavGPT-2: Unleashing Navigational Reasoning Capability for Large Vision-Language Models2024 ECCVCodeDataInstruct
2024InstructNav: Zero-Shot System for Generic Instruction Navigation in Unexplored Environment2024 arXivCode
2024NavGPT: Explicit Reasoning in Vision-and-Language Navigation with Large Language Models2024 AAAICode

Integrated Perception-Reasoning-Action Model

YearPaperVenueResources
2026ReflectVLN: Training Vision-Language Navigation Agents with Reflective Reasoning2026 arXivCode
2026AwareVLN: Reasoning with Self-awareness for Vision-Language Navigation2026 CVPRProjectCode
2026A Deployable Embodied Vision-Language Navigation System with Hierarchical Cognition and Context-Aware Exploration2026 arXiv
2026AsyncShield: A Plug-and-Play Edge Adapter for Asynchronous Cloud-based VLA Navigation2026 arXiv
2026AURA: Multimodal Shared Autonomy for Real-World Urban Navigation2026 arXiv
2026HiRO-Nav: Hybrid Reasoning Enables Efficient Embodied Navigation2026 arXiv
2026AgentVLN: Towards Agentic Vision-and-Language Navigation2026 arXivCode
2026EmergeNav: Structured Embodied Inference for Zero-Shot Vision-and-Language Navigation in Continuous Environments2026 arXiv
2026ABot-N0: Technical Report on the VLA Foundation Model for Versatile Embodied Navigation2026 arXiv
2026VLingNav: Embodied Navigation with Adaptive Reasoning and Visual-Assisted Linguistic Memory2026 arXiv
2025D3D-VLP: Dynamic 3D Vision-Language-Planning Model for Embodied Grounding and Navigation2025 arXiv
2025MobileVLA-R1: Reinforcing Vision-Language-Action for Mobile Robots2025 arXivProjectCode
2025CompassNav: Steering from Path Imitation to Decision Understanding in Navigation2025 arXiv
2025AdaNav: Adaptive Reasoning with Uncertainty for Vision-Language Navigation2025 arXivCode
2025Nav-r1: Reasoning and navigation in embodied scenes2025 arXivProjectCode
2025OctoNav: Towards Generalist Embodied Navigation2025 arXiv
2025Aux-Think: Exploring Reasoning Strategies for Data-Efficient Vision-Language Navigation2025 NeurlPSProjectData
2025NavCoT: Boosting LLM-based Vision-and-Language Navigation via Learning Disentangled Reasoning2025 TPAMI

Predictive Reasoning Model

YearPaperVenueResources
2026FantasyVLN: Unified Multimodal Chain-of-Thought Reasoning for Vision-Language Navigation2026 arXivProjectCode
2025NavForesee: A Unified Vision-Language World Model for Hierarchical Planning and Dual-Horizon Navigation Prediction2025 arXiv
2026WorldMAP: Bootstrapping Vision-Language Navigation Trajectory Prediction with Generative World Models2026 arXiv
2025ImagineNav++: Prompting Vision-Language Models as Embodied Navigator through Scene Imagination2025 arXiv

Hierarchical Decision

High-Level System

YearPaperVenueResources
2026LookasideVLN: Direction-Aware Aerial Vision-and-Language Navigation2026 arXiv
2026HTNav: A Hybrid Navigation Framework with Tiered Structure for Urban Aerial Vision-and-Language Navigation2026 arXiv
2026FloorPlan-VLN: A New Paradigm for Floor Plan Guided Vision-Language Navigation2026 arXiv
2026OpenFrontier: General Navigation with Visual-Language Grounded Frontiers2026 arXiv
2024VLFM: Vision-Language Frontier Maps for Zero-Shot Semantic Navigation2024 ICRA
2024ETPNav: Evolving Topological Planning for Vision-Language Navigation in Continuous Environments2024 TPAMI
2022Think Global, Act Local: Dual-Scale Graph Transformer for Vision-and-Language Navigation2022 CVPR

Low-Level System

YearPaperVenueResources
2026X-NavDP: Generalizing Navigation Diffusion Policy to Novel Behavior and Embodiments with Group Q-score Reweighted Matching2026 arXivProjectCode
2026WAM-Nav: Asymmetric Latent World-Action Modeling for Unified Visual Navigation2026 arXiv
2026NavDP: Learning sim-to-real navigation diffusion policy with privileged information guidance2026 ICRACode
2025LoGoPlanner: Localization Grounded Navigation Policy with Metric-aware Visual Geometry2025 arXivProject
2024NoMaD: Goal Masked Diffusion Policies for Navigation and Exploration2024 ICRACode
2024ViPlanner: Visual Semantic Imperative Learning for Local Navigation2024 ICRACode
2023iPlanner: Imperative Path Planning2023 arXivCode
2022Learning Forward Dynamics Model and Informed Trajectory Sampler for Safe Quadruped Navigation2022 RSS
2020DD-PPO: Learning Near-Perfect PointGoal Navigators from 2.5 Billion Frames2020 ICLR

Hybrid Architecture

YearPaperVenueResources
2026Token-Wise Latent Streaming from Slow Reasoners to Fast Planners for Dynamic Vision Language Navigation2026 arXivProject
2026SEDualVLN: A Spatially-Enhanced Dual-System for Vision-Language Navigation2026 arXivProjectCode
2026FineCog-Nav: Integrating Fine-grained Cognitive Modules for Zero-shot Multimodal UAV Navigation2026 arXivProject
2026TIC-VLA: A Think-in-Control Vision-Language-Action Model for Robot Navigation in Dynamic Environments2026 arXiv
2026Ground Slow, Move Fast: A Dual-System Foundation Model for Generalizable Vision-and-Language Navigation2026 ICLRProjectCode
2026OmniNav: A Unified Framework for Prospective Exploration and Visual-Language Navigation2026 ICLRProjectCode
2025LaViRA: Language-Vision-Robot Actions Translation for Zero-Shot Vision Language Navigation in Continuous Environments2025 arXiv
2025Breaking down and building up: Mixture of skill-based vision-and-language navigation agents2025 arXiv
2025FlexVLN: Flexible Adaptation for Diverse Vision-and-Language Navigation Tasks2025 arXiv
2025InternVLA-N1: An Open Dual-System Navigation Foundation Model with Learned Latent Plans2025 arXiv

⚠️ Reflection from Error

Sources and Taxonomy of Errors in VLN

Instruction Errors

YearPaperVenueResources
2026VLN-NF: Feasibility-Aware Vision-and-Language Navigation with False-Premise Instructions2026 arXivProject
2024Mind the error! detection and localization of instruction errors in vision-and-language navigation2024 IROSProjectCode

Perceptual Errors

  • Covered by papers already listed above.

Planning and Decision Errors

YearPaperVenueResources
2026Where Did It Go Wrong? Capability-Oriented Failure Attribution for Vision-and-Language Navigation Agents2026 arXiv

Error Prevention

Prevent Instruction Errors

YearPaperVenueResources
2025Seeing with Partial Certainty: Conformal Prediction for Robotic Scene Recognition in Built Environments2025 arXiv
2025Collaborative Instance Object Navigation: Leveraging Uncertainty-Awareness to Minimize Human-Agent Dialogues2025 ICCVProjectCodeData
2024I2EDL: Interactive Instruction Error Detection and Localization2024 RO-MAN

Prevent Perception Errors

  • Covered by papers already listed above.

Prevent Planning and Decision Errors

YearPaperVenueResources
2026SeqWalker: Sequential-Horizon Vision-and-Language Navigation with Hierarchical Planning2026 arXiv
2026EvolveNav: Empowering LLM-Based Vision-Language Navigation via Self-Improving Embodied Reasoning2026 TPAMICode

Learn From Error

Distribution Alignment

YearPaperVenueResources
2024Vision-Language Navigation with Energy-Based Policy2024 NeurIPS

DAgger-based

YearPaperVenueResources
2026Joint On-and-Off Policy Learning for Vision-and-Language Navigation2026 arXivProject
2026Path-level Hindsight Instructions for Semantic Exploration in Vision-Language Navigation2026 ECCV
2026GN0: Toward a Unified Paradigm for Generation, Evaluation, and Policy Learning in Visual-Language Navigation2026 arXivProjectCodeModel
2026LatentPilot: Scene-Aware Vision-and-Language Navigation by Dreaming Ahead with Latent Visual Reasoning2026 arXivProjectCode
2025Efficient-VLN: A Training-Efficient Vision-Language Navigation Model2025 arXivProject
2025CorrectNav: Self-Correction Flywheel Empowers Vision-Language-Action Navigation Model2025 arXivProject
2025DAgger Diffusion Navigation: DAgger Boosted Diffusion Policy for Vision-Language Navigation2025 arXiv

Exploration-based

YearPaperVenueResources
2026The Essence of Balance for Self-Improving Agents in Vision-and-Language Navigation2026 arXiv
2026Let's Reward Step-by-Step: Step-Aware Contrastive Alignment for Vision-Language Navigation in Continuous Environments2026 arXivCode
2026Trajectory-Diversity-Driven Robust Vision-and-Language Navigation2026 arXiv
2026Nipping the Drift in the Bud: Retrospective Rectification for Robust Vision-Language Navigation2026 arXiv
2026ActiveVLN: Towards Active Exploration via Multi-Turn RL in Vision-and-Language Navigation2026 ICRACode
2025SeeNav-Agent: Enhancing Vision-Language Navigation with Visual Prompt and Step-Level Policy Optimization2025 arXivCodeModel

Test-time Adaptation & Self Correstion

YearPaperVenueResources
2026SC2-WM: A Self-Correcting World Model with Closed-Loop Feedback for Vision-and-Language Navigation in Continuous Environments2026 ICMLCode
2026No Training, Better Flights: Test-Time Scaled VLMs for UAV Navigation2026 arXiv
2026Turning Adaptation into Assets: Cross-Domain Bridging for Online Vision-Language Navigation2026 ICML
2025Active Test-time Vision-Language Navigation2025 arXiv
2025Test-Time Adaptation for Online Vision-Language Navigation with Feedback-based Reinforcement Learning2025 ICML
2023Fast-slow test-time adaptation for online vision-and-language navigation2023 arXivCode

🌐 Broader Navigation Literature

Surveys and Evaluation

YearPaperVenueResources
2026A Comprehensive Survey and Systematic Real-World Evaluation of Embodied Vision-and-Language Navigation2026 IEEE TASE
2026A comprehensive review of recent advancements in vision-and-language navigation2026 Discover Computing
2026Robot Navigation via Foundation Language Models: A Review2026 ACM Computing Surveys
2026Vision-Language Navigation for Aerial Robots: Towards the Era of Large Language Models2026 arXiv
2026Vision-and-Language Navigation for UAVs: Progress, Challenges, and a Research Roadmap2026 arXiv
2026Mini-BEHAVIOR-Gran: Revealing U-Shaped Effects of Instruction Granularity on Language-Guided Embodied Agents2026 arXiv
2026NavTrust: Benchmarking Trustworthiness for Embodied Navigation2026 arXivProject
2024Vision-and-language navigation today and tomorrow: A survey in the era of foundation models2024 TMLRCode
2024Vision-language navigation: a survey and taxonomy2024 Neural Computing and Applications
2023Advances in embodied navigation using large language models: A survey2023 arXiv
2023A2Nav: Action-Aware Zero-Shot Robot Navigation by Exploiting Vision-and-Language Ability of Foundation Models2023 arXiv
2023Visual language navigation: A survey and open challenges2023 Artificial Intelligence Review
2022Vision-and-language navigation: A survey of tasks, methods, and future directions2022 ACLCode
2019General evaluation for instruction conditioned navigation using dynamic time warping2019 arXiv
2018On evaluation of embodied navigation agents2018 arXiv

General Navigation and Object Navigation

YearPaperVenueResources
2026Scaffolding Foundation Models into Physical-World Agents Pushes the Frontier of Long-Horizon Navigation2026 arXiv
2026LightNav-0: Eliciting VLM Spatial Intelligence for Generalist Embodied Navigation2026 arXivProjectCode
2026VoLN: Vision-Only Long-Horizon Navigation---Paradigm, Benchmark, and Method2026 arXivProjectCodeData
2026ABot-N1: Toward a General Visual Language Navigation Foundation Model2026 arXivProject
2026Qwen-RobotNav: A Scalable Navigation Model Designed for an Agentic Navigation System2026 arXivProject
2026GN0: Toward a Unified Paradigm for Generation, Evaluation, and Policy Learning in Visual-Language Navigation2026 arXivProjectCodeModel
2026R2F: Repurposing Ray Frontiers for LLM-free Object Navigation2026 arXiv
2026Hydra-Nav: Object Navigation via Adaptive Dual-Process Reasoning2026 arXiv
2026MerNav: A Highly Generalizable Memory-Execute-Review Framework for Zero-Shot Object Goal Navigation2026 arXiv
2025RANGER: A Monocular Zero-Shot Semantic Navigation Framework through Contextual Adaptation2025 arXiv
2025SocialNav-Map: Dynamic Mapping with Human Trajectory Prediction for Zero-Shot Social Navigation2025 arXivCode
2025NavSpace: How Navigation Agents Follow Spatial Intelligence Instructions2025 arXivCode
2025C-NAV: Towards self-evolving continual object navigation in open world2025 arXiv
20253D-Mem: 3D Scene Memory for Embodied Exploration and Reasoning2025 arXiv
2024TopV-Nav: Unlocking the top-view spatial reasoning potential of MLLM for zero-shot object navigation2024 arXiv
2023ViNT: A foundation model for visual navigation2023 arXivProjectCode
2022GNM: A general navigation model to drive any robot2022 arXivCode

Embodied / VLA Context

YearPaperVenueResources
2025MemoryVLA: Perceptual-Cognitive Memory in Vision-Language-Action Models for Robotic Manipulation2025 arXiv
2024π0: A Vision-Language-Action Flow Model for General Robot Control2024 arXiv
2024OpenVLA: An Open-Source Vision-Language-Action Model2024 arXivProjectCode

Policies, Efficiency, and Applied Navigation

YearPaperVenueResources
2026Embodied Agents Take Control: Minimal-Interface Zero-Shot Agents Rival Industrial-Scale Policies in Vision-and-Language Navigation2026 arXiv
2026P3Nav: End-to-End Perception, Prediction and Planning for Vision-and-Language Navigation2026 arXiv
2026Beyond Imitation: Reinforcement Learning Fine-Tuning for Adaptive Diffusion Navigation Policies2026 arXiv
2026NavDP: Learning sim-to-real navigation diffusion policy with privileged information guidance2026 ICRACode
2025Nav-EE: Navigation-Guided Early Exiting for Efficient Vision-Language Models in Autonomous Driving2025 arXiv
2025OmniVLA: An omni-modal vision-language-action model for robot navigation2025 arXivProjectCode
2025EgoPrune: Efficient token pruning for egomotion video reasoning in embodied agent2025 arXiv

🚀 Future Directions

Representative outlook papers explicitly discussed in the manuscript are listed below.

Toward Action-Aware Policies

  • Action grounding beyond static image-text reasoning.

Toward Visual Predictive CoT

  • Reasoning grounded in predicted visual futures.

Towards Unified Visual Geometry Navigation Model

  • Tighter coupling between geometry and policy learning.

Towards Dynamic Environment

YearPaperVenueResources
2026Walk With Me: Long-Horizon Social Navigation for Human-Centric Outdoor Assistance2026 arXiv
2025HA-VLN: A Benchmark for Human-Aware Navigation in Discrete-Continuous Environments with Dynamic Multi-Human Interactions, Real-World Validation, and an Open Leaderboard2025 arXivCode
2025From Cognition to Precognition: A Future-Aware Framework for Social Navigation2025 ICRA
2025VLN-ChEnv: Vision-language Navigation in Changeable Environments2025 ACM MM
2024VLM-Social-Nav: Socially Aware Robot Navigation through Scoring using Vision-Language Models2024 RA-LProject

Towards Collaborative VLN

YearPaperVenueResources
2026Benchmarking Interaction, Beyond Policy: a Reproducible Benchmark for Collaborative Instance Object Navigation2026 arXiv
2026DeCoNav: Dialog enhanced Long-Horizon Collaborative Vision-Language Navigation2026 arXiv
2026MA-CoNav: A Master-Slave Multi-Agent Framework with Hierarchical Collaboration and Dual-Level Reflection for Long-Horizon Embodied VLN2026 arXiv
2026FSUNav: A Cerebrum-Cerebellum Architecture for Fast, Safe, and Universal Zero-Shot Goal-Oriented Navigation2026 arXiv
2026CA-VLN: Collaborative Agents in MLLM-Powered Visual-Language Navigation2026 Sensors
2025Enhancing Multi-Robot Semantic Navigation through Multimodal Chain-of-Thought Score Collaboration2025 AAAI
2023Co-NavGPT: Multi-Robot Cooperative Visual Semantic Navigation Using Vision Language Models2023 arXiv

About

End-to-End Visual Language Navigation with Limited Sensing: A Survey

Resources

Stars

41 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages