Data Download
Official HA-VLN repository: https://github.com/JostarXiong/HA-VLN
This page explains how to acquire and prepare all datasets, scene meshes, and model weights required for HA-VLN.
1. Directory Layout
All scene meshes, human activities, episodes, and baseline checkpoints reside under the Data/ directory:
Data/
├── HA-R2R/
│ ├── train/train_bertidx.json.gz
│ ├── val_seen/val_seen_bertidx.json.gz
│ └── val_unseen/val_unseen_bertidx.json.gz
├── HA-R2R-tools/
│ ├── collision_num_val_seen.json
│ └── collision_num_val_unseen.json
├── Multi-Human-Annotations/
│ └── human_motion.json
├── HAPS2_0/
│ └── <dynamic_human_motion_glbs>
├── scene_datasets/
│ └── mp3d/<scan>/<scan>.glb
└── ddppo-models/
└── gibson-2plus-resnet50.pth # optional for depth encoder
2. Matterport3D Scene Meshes (Data/scene_datasets)
Matterport3D is licensed separately by Matterport and cannot be redistributed.
- Request access on the Matterport3D Project Page and submit the signed Terms of Use to receive your personal
download_mp.pyscript. - Download and extract the Habitat scene meshes:
python3 /path/to/download_mp.py -o Data/scene_datasets --task_data habitat
# After task-data download finishes, press Ctrl-C at the prompt for the main dataset.
unzip Data/scene_datasets/v1/tasks/mp3d_habitat.zip -d Data/scene_datasets
Verify that scene meshes reside at Data/scene_datasets/mp3d/<scan>/<scan>.glb.
3. HA-VLN Simulation Assets & Annotations
Simulation assets (HAPS 2.0 dynamic human motion models, HA-R2R episodes, and baseline weights) are officially hosted on Hugging Face (fly1113/HA-VLN). Multi-human placement metadata (human_motion.json) and collision evaluation baselines are fetched from the repository.
Option A: Official Downloader Script (Recommended)
The verified downloader automatically fetches all assets, unzips and normalizes HAPS 2.0 meshes, and retrieves collision baselines and human annotations:
python scripts/download_hf.py --destination Data --target all
Option B: Helper Shell Script (Google Drive Mirror)
pip install gdown
bash scripts/download_data.sh
[!NOTE] The
--destinationand--dry-runCLI options belong to the challenge toolkit (roboworld2026-track2/scripts/download_data.sh), which supports dry-run inspection before downloading challenge assets (bash scripts/download_data.sh --destination /path/to/Data --dry-run). In the coreHA-VLNrepository,scripts/download_data.shdownloads directly intoData/via Google Drive without command-line flags; usepython scripts/download_hf.py --destination Data --target allfor configurable command-line downloading.
Option C: Manual Hugging Face Hub CLI
pip install huggingface-hub
hf download fly1113/HA-VLN --repo-type dataset --local-dir Data
Note: Raw HF Hub download leaves HAPS2_0.zip unextracted. If using this option, you must manually unpack HAPS2_0.zip (flattening human_motion_glbs_v3 into Data/HAPS2_0), and retrieve Multi-Human-Annotations/human_motion.json and Data/HA-R2R-tools/collision_num_val_*.json from the repository.
4. Pretrained Depth Encoder Weights (Optional for CMA)
If training or evaluating CMA from scratch with depth observations, download the pretrained DD-PPO ResNet-50 weights:
curl -fL https://dl.fbaipublicfiles.com/habitat/data/baselines/v1/ddppo/ddppo-models.zip -o ddppo-models.zip
# Extract directly into ddppo-models without creating nested directories:
unzip -j ddppo-models.zip "data/ddppo-models/gibson-2plus-resnet50.pth" -d Data/ddppo-models
rm ddppo-models.zip
5. GroundingDINO Weights (Optional for Human Counting)
If HUMAN_COUNTING is enabled in HASimulator/detector.py:
mkdir -p HASimulator/GroundingDINO/weights
curl -fL --retry 3 \
https://github.com/IDEA-Research/GroundingDINO/releases/download/v0.1.0-alpha/groundingdino_swint_ogc.pth \
-o HASimulator/GroundingDINO/weights/groundingdino_swint_ogc.pth
6. Vocabulary Expansion (for static vocab agents)
If your agent uses a static vocabulary file, expand it from HA-R2R corpora to avoid OOV tokens:
import json
import gzip
import re
from pathlib import Path
def clean_text(text: str) -> list:
text = text.lower()
text = re.sub(r'([.?!,;:/\\()\[\]"\'\-])', r' \1 ', text)
text = re.sub(r'\s+', ' ', text).strip()
return text.split()
def update_vocabulary(ha_r2r_dir: str, existing_vocab_path: str, output_vocab_path: str):
ha_r2r_path = Path(ha_r2r_dir)
with open(existing_vocab_path, 'r', encoding='utf-8') as f:
existing_words = [line.strip() for line in f.readlines()]
vocab_set = set(existing_words)
new_words = []
for split in ['train', 'val_seen', 'val_unseen']:
json_file = ha_r2r_path / split / f"{split}.json.gz"
if not json_file.exists():
continue
with gzip.open(json_file, 'rt', encoding='utf-8') as f:
data = json.load(f)
for item in data:
for instruction in item.get('instructions', []):
tokens = clean_text(instruction)
for token in tokens:
if token not in vocab_set:
vocab_set.add(token)
new_words.append(token)
with open(output_vocab_path, 'w', encoding='utf-8') as f:
for word in existing_words + new_words:
f.write(f"{word}\n")