IP Library Granted Patent US 12,526,599
Granted Patent B2
US 12,526,599 · App. 18/213,175 · Granted Jan 13, 2026

Localization using audio and visual data

Inventors: Karren Dai Yang (Medford, MA); Michael David Firman (London, GB); Eric Brachmann (Hanover, DE); Clément Godard (San Francisco, CA)
Assignee: Niantic Spatial, Inc.
H04S7/303G06T7/74G06T2207/20084G06T2207/30196
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,526,599
App. No.
18/213,175
Granted
Jan 13, 2026
Kind
B2
Abstract

A reference image and recorded sound of an environment of a client device are obtained. The recorded sound may be captured by a microphone of the client device in a period of time after generation of a localization sound by the client device. The location of the client device in the environment may be determined using the reference image and the recorded sound.

Claims (66)

1 . A computer-implemented method for localization, the method comprising:

obtaining a reference image of an environment of a client device;

obtaining recorded sound of the environment of the client device, the recorded sound captured by a microphone of the client device in a period of time after generation of a localization sound by the client device; and

determining a location of the client device in the environment using the reference image and the recorded sound, wherein determining the location comprises:

extracting features from the recorded sound and the reference image;

determining candidate poses based on the features extracted from the recorded sound and the reference image; and

selecting between the candidate poses, wherein selecting between the candidate poses comprises:

providing the candidate poses as input to a learned gating function; and

obtaining from the learned gating function an output indicating whether to use a candidate pose based on the recorded sound or a candidate pose based on the reference image.

2 . The method of claim 1 further comprising:

obtaining a second recorded sound, wherein the second recorded sound is a recorded sound of the environment captured by a second microphone in the period of time after generation of the localization sound by a second client device; and

determining a location of the second microphone relative to the location of the client device in the environment using the second recorded sound.

3 . The method of claim 2 further comprising:

obtaining a second reference image of the environment based on a second client device; and

determining a location of the second client device relative to the location of the client device in the environment using the second recorded sound and the second reference image.

4 . The method of claim 1 further comprising:

querying an audio-visual database based on the reference image of the environment and the recorded sound; and

determining the location of the client device based on query of the audio-visual database.

5 . The method of claim 1 further comprising:

receiving, at a server, the reference image and the recorded sound, wherein the determining of the location of the client device in the environment is performed by the server; and

providing, to the client device, the location of the client device.

6 . The method of claim 1 wherein the location of the client device includes a description of an orientation of the client device with respect to the environment.

7 . The method of claim 1 , wherein the determining of the location of the client device in the environment is determined at the client device.

8 . The method of claim 1 , wherein determining candidate poses based on the features extracted from the recorded sound and the reference image further comprises comparing the features extracted from the recorded sound and the reference image to features from an audio-visual database.

9 . A non-transitory computer-readable medium storing computer-executable instructions for localization that, when executed by a computing system, cause the computing system, to perform operations comprising:

obtaining a reference image of an environment of a client device;

obtaining recorded sound of the environment of the client device, the recorded sound captured by a microphone of the client device in a period of time after generation of a localization sound by the client device; and

determining a location of the client device in the environment using the reference image and the recorded sound, wherein determining the location comprises:

extracting features from the recorded sound and the reference image;

determining candidate poses based on the features extracted from the recorded sound and the reference image; and

selecting between the candidate poses, wherein selecting between the candidate poses comprises:

providing the candidate poses as input to a learned gating function; and

obtaining from the learned gating function an output indicating whether to use a candidate pose based on the recorded sound, a candidate pose based on the reference image, or an optimal combination of the candidate poses.

10 . The computer-readable medium of claim 9 wherein the operations further comprise:

obtaining a second recorded sound, wherein the second recorded sound is a recorded sound of the environment captured by a second microphone in the period of time after generation of the localization sound by a second client device; and

determining a location of the second microphone relative to the location of the client device in the environment using the second recorded sound.

11 . The computer-readable medium of claim 10 wherein the operations further comprise:

obtaining a second reference image of the environment based on a second client device;

and determining a location of the second client device relative to the location of the client device in the environment using the second recorded sound and the second reference image.

12 . The computer-readable medium of claim 9 wherein the operations comprise:

querying an audio-visual database based on the reference image of the environment and the recorded sound; and

determining the location of the client device based on query of the audio-visual database.

13 . The computer-readable medium of claim 9 wherein the location of the client device includes a description of an orientation of the client device with respect to the environment.

14 . A computer-implemented method for localization, the method comprising:

obtaining a reference image of an environment of a client device;

obtaining recorded sound of the environment of the client device, the recorded sound captured by a microphone of the client device in a period of time after generation of a localization sound by the client device; and

determining a location of the client device in the environment using the reference image and the recorded sound, wherein determining the location comprises:

extracting features from the recorded sound and the reference image;

determining candidate poses based on the features extracted from the recorded sound and the reference image; and

selecting between the candidate poses, wherein selecting between the candidate poses comprises:

providing the candidate poses as input to a learned gating function; and

obtaining from the learned gating function an output indicating an optimal combination of the candidate poses.

15 . The method of claim 14 further comprising:

obtaining a second recorded sound, wherein the second recorded sound is a recorded sound of the environment captured by a second microphone in the period of time after generation of the localization sound by a second client device; and

determining a location of the second microphone relative to the location of the client device in the environment using the second recorded sound.

16 . The method of claim 15 further comprising:

obtaining a second reference image of the environment based on a second client device; and

determining a location of the second client device relative to the location of the client device in the environment using the second recorded sound and the second reference image.

17 . The method of claim 14 further comprising:

querying an audio-visual database based on the reference image of the environment and the recorded sound; and

determining the location of the client device based on query of the audio-visual database.

18 . The method of claim 14 further comprising:

receiving, at a server, the reference image and the recorded sound, wherein the determining of the location of the client device in the environment is performed by the server; and

providing, to the client device, the location of the client device.

19 . The method of claim 14 wherein the location of the client device includes a description of an orientation of the client device with respect to the environment.

20 . The method of claim 14 , wherein determining candidate poses based on the features extracted from the recorded sound and the reference image further comprises comparing the features extracted from the recorded sound and the reference image to features from an audio-visual database.

Assignments (5)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jun 16, 2025
From: NIANTIC, INC.
To: NIANTIC SPATIAL, INC.
Reel/Frame 071555/0833 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Dec 17, 2024
From: FIRMAN, MICHAEL DAVID; BRACHMANN, ERIC
To: NIANTIC INTERNATIONAL TECHNOLOGY LIMITED
Reel/Frame 069605/0261 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Dec 17, 2024
From: YANG, KARREN DAI; GODARD, CLÉMENT
To: NIANTIC, INC.
Reel/Frame 069605/0264 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jan 22, 2024
From: NIANTIC INTERNATIONAL TECHNOLOGY LIMITED
To: NIANTIC, INC.
Reel/Frame 066197/0211 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jul 27, 2023
From: NIANTIC INTERNATIONAL TECHNOLOGY LIMITED
To: NIANTIC, INC.
Reel/Frame 064408/0956 →
Continuity (2)
Provisional Application 63355086 · Jun 23, 2022
Related Publication 20230421985A1 · Dec 28, 2023
References Cited (19)
US 10356393B1 · Binns · 2019 [cited by examiner]
US 20100189271A1 · Tsujino · 2010 [cited by examiner]
US 20120163610A1 · Sakagami · 2012 [cited by examiner]
US 20170311080A1 · Kolb · 2017 [cited by examiner]
US 20170323472A1 · Barnes · 2017 [cited by examiner]
US 20200058169A1 · Friesenhahn · 2020 [cited by examiner]
US 20210344831A1 · Vilermo · 2021 [cited by examiner]
US 20210375049A1 · Syed · 2021 [cited by examiner]
US 20210400417A1 · Freeman · 2021 [cited by examiner]
US 20220028108A1 · Haapoja · 2022 [cited by examiner]
US 20230011087A1 · Mekler · 2023 [cited by examiner]
US 20240398498A1 · Seyed Vahedein · 2024 [cited by examiner]
CN 110999328A · 2020 [cited by examiner]
CN 110089131B · 2021 [cited by examiner]
EP 3400705B1 · 2021 [cited by examiner]
GB 2606650A · 2022 [cited by examiner]
WO WO2018042770A1 · 2018 [cited by examiner]
WO WO2023229600A1 · 2023 [cited by examiner]
Azizyan et al, “SurroundSense: Mobile Phone Localization via Ambience Fingerprinting”. 12 pages. (Year: 2009). [cited by examiner]