Skip to content

Repository files navigation

LookStep

简体中文MODEL

LookStep is a self-contained package for reproducing the main experiments in LookStep: Efficient Vision-Language Navigation with Linguistic Foresight and Event Driven Memory. After downloading the paper checkpoint and preparing the licensed Matterport3D and VLN-CE data, this directory can be used on its own to evaluate the R2R-CE and RxR-CE Val-Unseen main results. It also supports reconstructing the supervision data from expert trajectories and training the model.

Three modules

LookStep/
├── data_construction/ # 1. R2R/RxR expert trajectories -> LFS/EDRM supervision JSONL
├── model_training/ # 2. ms-swift conversion and Qwen3-VL full fine-tuning
└── simulation/ # 3. Online Habitat main-experiment evaluation for R2R/RxR

The complete pipeline is:

R2R-CE + RxR-CE expert trajectories
-> deterministic LookStep labels
-> ms-swift multimodal JSONL
-> Qwen3-VL-8B full fine-tuning
-> Habitat R2R-CE / RxR-CE Val-Unseen evaluation

The training target uses the following fixed order:

progress -> event -> memory_write -> memory_role
-> outcomes for four candidate actions -> action

Reproduce the main results with a downloaded model

First prepare the pinned environment and licensed data by following the reproduction guide, then run:

bash LookStep/reproduce_paper.sh test
conda activate lookstep-simulation
MODEL_PATH=/path/to/downloaded/lookstep-checkpoint \
PROCESSOR_PATH=/path/to/Qwen3-VL-8B-Instruct \
DATA_ROOT=/path/to/data \
bash LookStep/reproduce_paper.sh check-sim
MODEL_PATH=/path/to/downloaded/lookstep-checkpoint \
PROCESSOR_PATH=/path/to/Qwen3-VL-8B-Instruct \
DATA_ROOT=/path/to/data \
bash LookStep/reproduce_paper.sh smoke-r2r
MODEL_PATH=/path/to/downloaded/lookstep-checkpoint \
PROCESSOR_PATH=/path/to/Qwen3-VL-8B-Instruct \
DATA_ROOT=/path/to/data \
CUDA_VISIBLE_DEVICES=0,1,2,3,4,5,6,7 \
bash LookStep/reproduce_paper.sh eval-all
bash LookStep/reproduce_paper.sh verify

Main entry points

ModuleEntry pointPurpose
Data constructiondata_construction/build_short_label_dataset.pyDerive main-experiment labels from R2R/RxR expert actions
Model trainingmodel_training/prepare_ms_swift_data.pySplit by episode and convert to ms-swift JSONL
Model trainingmodel_training/scripts/train_qwen3vl_full.shFull-fine-tuning recipe for the paper checkpoint
Simulationsimulation/evaluate_short_label_sim.pyOnline EventFIFO navigation in Habitat
Orchestratorreproduce_paper.shStaged check/build/prepare/train/eval/verify workflow

Citation

@inproceedings{lookstep,
title = {LookStep: Efficient Vision-Language Navigation with Linguistic Foresight and Event Driven Memory},
author = {Kun-Yang Yu and Yingzhe Li and Hongyu Xu and Shi-Yu Tian and Zhi Zhou and Yang Chen and Ming Yang and Sheng Wang and Qing Yu and Lan-Zhe Guo and Yu-Feng Li},
booktitle = {The 2026 Conference on Empirical Methods in Natural Language Processing},
year = {2026}
}

If you have any questions, please contact yuky@lamda.nju.edu.cn.

About

Code for EMNLP 2026 paper LookStep: efficient vision-language navigation with linguistic foresight and event-driven memory.

Resources

Stars

2 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages