FM-Loc: Using Foundation Models for Improved Vision-based Localization

by   Reihaneh Mirjalili, et al.

Visual place recognition is essential for vision-based robot localization and SLAM. Despite the tremendous progress made in recent years, place recognition in changing environments remains challenging. A promising approach to cope with appearance variations is to leverage high-level semantic features like objects or place categories. In this paper, we propose FM-Loc which is a novel image-based localization approach based on Foundation Models that uses the Large Language Model GPT-3 in combination with the Visual-Language Model CLIP to construct a semantic image descriptor that is robust to severe changes in scene geometry and camera viewpoint. We deploy CLIP to detect objects in an image, GPT-3 to suggest potential room labels based on the detected objects, and CLIP again to propose the most likely location label. The object labels and the scene label constitute an image descriptor that we use to calculate a similarity score between the query and database images. We validate our approach on real-world data that exhibit significant changes in camera viewpoints and object placement between the database and query trajectories. The experimental results demonstrate that our method is applicable to a wide range of indoor scenarios without the need for training or fine-tuning.


page 1

page 4

page 5

page 6


Training Semantic Descriptors for Image-Based Localization

Vision based solutions for the localization of vehicles have become popu...

UAV Visual Teach and Repeat Using Only Semantic Object Features

We demonstrate the use of semantic object detections as robust features ...

Semantic Image Based Geolocation Given a Map

The problem visual place recognition is commonly used strategy for local...

Visual Place Representation and Recognition from Depth Images

This work proposes a new method for place recognition based on the scene...

Dynamic Objects Segmentation for Visual Localization in Urban Environments

Visual localization and mapping is a crucial capability to address many ...

Aladdin: Zero-Shot Hallucination of Stylized 3D Assets from Abstract Scene Descriptions

What constitutes the "vibe" of a particular scene? What should one find ...

Dominating Set Database Selection for Visual Place Recognition

This paper presents an approach for creating a visual place recognition ...

Please sign up or login with your details

Forgot password? Click here to reset