PofoliaShared via Pofolia

Advanced Robotics· 2026Q2· Review

A review of open-vocabulary semantic mapping and navigation with foundation models for mobile robots

Yoshinobu Hagiwara, S. Hasegawa

Short summary

Foundation models (LLMs, VLMs) combined with generative 3D representations enable mobile robots to create open-vocabulary semantic maps, allowing natural language understanding and interaction with novel objects.

AI-generated from the title and abstract; the full text is not read.

Key points

  • Foundation models (LLMs, VLMs) overcome limitations of closed-vocabulary semantic mapping for robots.
  • Generative 3D representations (NeRFs, 3D Gaussian splatting) enable dense spatial mapping linked to language.
  • Robots can now acquire semantic representations connecting perception, language, and action.
  • The review categorizes advancements into four taxonomies: feature fusion, object-centric mapping, scene graphs, and 3D language fields.

AI-generated from the title and abstract; the full text is not read.

Abstract

Autonomous mobile robots that coexist with humans must construct not only geometric maps but also semantic maps that can be accessed through natural language. Conventional semantic mapping has mainly focused on assigning labels from predefined closed vocabularies to metric maps, limiting its ability to handle novel objects, open-ended linguistic expressions, and flexible human-robot interaction. Recent advances in large-scale foundation models, particularly LLMs and VLMs, have accelerated research on open-vocabulary semantic mapping. In parallel, generative 3D representations such as neural radiance fields and 3D Gaussian splatting have enabled dense, continuous spatial representations associated with language-derived features. Together, these developments allow robots to acquire spatial semantic representations that connect perception, language, and action. This paper reviews this rapidly evolving field through a four-part taxonomy: (i) fusion of semantic features into 3D metric maps; (ii) object-centric open-vocabulary representations; (iii) hierarchical scene graph representations; and (iv) continuous generative 3D language fields. We also revisit the history of open-vocabulary semantic mapping and provide an overview of foundation model-based navigation using language-accessible maps, ranging from object-goal navigation to LLM-based hierarchical task planning. Finally, we introduce evaluation datasets, simulators, robot platforms, and evaluation metrics, and summarize seven open challenges.

The authors' abstract, as published at the source. Advanced Robotics, 2026 · DOI ↗

TakeawaysPremium
Ask the paperFree account

Continue with a free account

Ask the paper: 3 free questions a day about this paper; save it, get its citation, new summaries every day for your field. Takeaways are Premium.

Continue free on the web

Sign in with Google or Apple; no card needed. You come back to this paper.

On your phone:

Field: Aerospace Engineering

Aerospace EngineeringEngineering