Home Knowledge Base Vision-and-Language Navigation (VLN)

Vision-and-Language Navigation (VLN) is the embodied AI task requiring an agent to navigate through real 3D environments by following natural language instructions — perceiving visual scenes, grounding linguistic references to observed landmarks, and executing a sequence of movement actions to reach the described goal — serving as the benchmark for testing whether AI systems can truly understand the connection between language and the physical world, integrating visual perception, natural language understanding, spatial reasoning, and sequential decision-making in a single challenging task.

What Is Vision-and-Language Navigation?

Why VLN Matters

Key Benchmarks

BenchmarkEnvironmentInstructionsUnique Challenge
R2R (Room-to-Room)Matterport3D (90 buildings)21K English instructionsStandard VLN benchmark
RxRMatterport3D126K instructions in 3 languagesMultilingual, more detailed paths
SOONMatterport3DObject-goal with room descriptionsTarget is an object, not a viewpoint
REVERIEMatterport3DHigh-level instructions + object groundingMust find and identify target object
R2R-CEHabitat continuous environmentsR2R instructionsContinuous navigation (not graph)
ALFREDAI2-THORMulti-step manipulation instructionsNavigation + object interaction

Architecture Approaches

Key Challenges

Vision-and-Language Navigation is the integration test for embodied AI — the task that demands a machine simultaneously see, read, reason, plan, and act in realistic 3D worlds, making it the most comprehensive benchmark for evaluating whether AI can truly operate at the intersection of language and physical reality.

vision-and-language navigationrobotics

Explore 500+ Semiconductor & AI Topics

From EUV lithography to CUDA optimization — search the full knowledge base or chat with our AI assistant.