IP Library Patent Application 19422095
Patent Application
App. No. 19/422,095

Localization Using Audio and Visual Data

Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US None
App. No.
19/422,095
Abstract

A reference image and recorded sound of an environment of a client device are obtained. The recorded sound may be captured by a microphone of the client device in a period of time after generation of a localization sound by the client device. The location of the client device in the environment may be determined using the reference image and the recorded sound.

Claims (58)

1 . A computer-implemented method for localization, the method comprising:

obtaining a reference image of an environment of a client device;

obtaining recorded sound of the environment of the client device, the recorded sound captured by a microphone of the client device in a period of time after generation of a localization sound by the client device; and

determining a location of the client device in the environment using the reference image and the recorded sound.

2 . The computer-implemented method of claim 1 further comprising:

obtaining a second recorded sound, wherein the second recorded sound is a recorded sound of the environment captured by a second microphone in the period of time after generation of the localization sound by a second client device; and

determining a location of the second microphone relative to the location of the client device in the environment using the second recorded sound.

3 . The computer-implemented method of claim 2 further comprising:

obtaining a second reference image of the environment based on a second client device; and

determining a location of the second client device relative to the location of the client device in the environment using the second recorded sound and the second reference image.

4 . The computer-implemented method of claim 1 further comprising:

querying an audio-visual database based on the reference image of the environment and the recorded sound; and

determining the location of the client device based on query of the audio-visual database.

5 . The computer-implemented method of claim 1 further comprising:

receiving, at a server, the reference image and the recorded sound, wherein the determining of the location of the client device in the environment is performed by the server; and

providing, to the client device, the location of the client device.

6 . The computer-implemented method of claim 1 , wherein the location of the client device includes a description of an orientation of the client device with respect to the environment.

7 . The computer-implemented method of claim 1 , wherein the determining of the location of the client device in the environment is determined at the client device.

8 . The computer-implemented method of claim 1 , wherein the determining the location comprises:

extracting features from the recorded sound and the reference image;

determining candidate poses based on the features extracted from the recorded sound and the reference image; and

selecting between the candidate poses.

9 . The computer-implemented method of claim 8 , wherein selecting between the candidate poses comprises:

validating the candidate poses based on the reference images; and

selecting the candidate poses based on the recorded sound.

10 . The computer-implemented method of claim 8 , wherein determining candidate poses based on the features extracted from the recorded sound and the reference image comprises:

determining a first set of candidate poses based on the features extracted from the recorded sound;

determining a second set of candidate poses based on the features extracted from the reference image; and

fusing the first set of candidate poses and the second set of candidate poses.

11 . The computer-implemented method of claim 8 , wherein determining candidate poses based on the features extracted from the recorded sound and the reference image further comprises comparing the features extracted from the recorded sound and the reference image to features from an audio-visual database.

12 . The computer-implemented method of claim 8 , wherein determining candidate poses based on the features extracted from the recorded sound and the reference image comprises:

providing as input the features extracted from the recorded sound and the reference image to a machine learning model; and

receiving as output the candidate poses.

13 . A non-transitory computer-readable medium storing computer-executable instructions for localization that, when executed by a computing system, cause the computing system to perform operations comprising:

obtaining a reference image of an environment of a client device;

obtaining recorded sound of the environment of the client device, the recorded sound captured by a microphone of the client device in a period of time after generation of a localization sound by the client device; and

determining a location of the client device in the environment using the reference image and the recorded sound.

14 . The non-transitory computer-readable medium of claim 13 , wherein the operations further comprise:

obtaining a second recorded sound, wherein the second recorded sound is a recorded sound of the environment captured by a second microphone in the period of time after generation of the localization sound by a second client device; and

determining a location of the second microphone relative to the location of the client device in the environment using the second recorded sound.

15 . The non-transitory computer-readable medium of claim 14 , wherein the operations further comprise:

obtaining a second reference image of the environment based on a second client device; and

determining a location of the second client device relative to the location of the client device in the environment using the second recorded sound and the second reference image.

16 . The non-transitory computer-readable medium of claim 13 , wherein the operations comprise:

querying an audio-visual database based on the reference image of the environment and the recorded sound; and

determining the location of the client device based on query of the audio-visual database.

17 . The non-transitory computer-readable medium of claim 13 , wherein the operations further comprise:

receiving, at a server, the reference image and the recorded sound;

determining, at the server, the location of the client device in the environment using the reference image and the recorded sound; and

providing to the client device the location of the client device.

18 . The non-transitory computer-readable medium of claim 13 , wherein the location of the client device includes a description of an orientation of the client device with respect to the environment.

19 . A computer system comprising:

a device having a camera, a speaker, and a microphone; and

a localization system configured to:

receive an image of an environment captured by the camera;

receive recorded sound of the environment, the recorded sound captured by the microphone in a period of time after generation of a localization sound by the speaker; and

determine a location of the device in the environment using the image and the recorded sound.

20 . The computer system of claim 19 , wherein the localization system in part of the device.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jun 26, 2026
From: FIRMAN, MICHAEL DAVID; BRACHMANN, ERIC
To: NIANTIC INTERNATIONAL TECHNOLOGY LIMITED
Reel/Frame 075100/0469 →