IP Library Patent Application 19038488
Patent Application
App. No. 19/038,488

Map-Relative Pose Regression for Visual Relocalization

Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US None
App. No.
19/038,488
Abstract

This disclosure pertains to a scene-agnostic, map-relative pose regression method. The pose regressor is conditioned on a scene-specific map representation such that its pose predictions are relative to the scene map. This allows training of the pose regressor across multiple scenes to learn the generic relation between a scene-specific map representation and the camera pose. The map-relative pose regressor can then be applied to new map representations.

Claims (44)

1 . A computer-implemented method for determining a camera pose, the method comprising:

obtaining an image of a scene captured by a camera;

inputting the image into a first machine-learned model trained to predict a 3D coordinate map of the scene by mapping 2D pixels from the image to 3D scene coordinates;

inputting the 3D coordinate map of the scene into a second machine-learned model trained to estimate a pose of the camera when capturing the image of the scene based on the 3D coordinate map; and

causing display of a virtual element at a position on a display of a client device, the position determined based on the pose of the camera estimated by the second machine-learned model.

2 . The computer-implemented method of claim 1 , wherein the image is an RGB image, and wherein the second machine-learned model estimates the pose of the camera without requiring a depth map or a point-cloud model of the image of the scene captured by the camera.

3 . The computer-implemented method of claim 1 , wherein the first machine-learned model is a scene-specific geometry-based prediction model that processes the captured image to predict the 3D coordinate map, the first machine-learned model having been trained specifically for the scene based on training data associated with the scene.

4 . The computer-implemented method of claim 3 , wherein the first machine-learned model is a convolutional neural network.

5 . The computer-implemented method of claim 1 , further comprising:

inputting a camera intrinsic matrix associated with the captured image into the second machine-learned model, wherein the second machine-learned model is a scene-agnostic pose regressor model that directly regresses a multi-dimensional pose vector for the captured image of the scene based on the 3D coordinate map and the camera intrinsic matrix associated with the captured image.

6 . The computer-implemented method of claim 5 , wherein the second machine-learned model is a transformer.

7 . The computer-implemented method of claim 5 , further comprising:

training the second machine-learned model using training data including 3D coordinate maps of a plurality of historical scenes and corresponding ground truth camera poses, wherein the second machine-learned model is agnostic to the scene.

8 . The computer-implemented method of claim 1 , wherein the pose of the camera indicates a position and an orientation of the camera relative to the 3D coordinate map of the scene, and wherein the method further comprises:

transmitting game data to the client device including the camera, the game data based on the pose of the camera.

9 . The computer-implemented method of claim 1 , wherein the first machine-learned model is a scene-specific geometry-based prediction model trained on images of the scene and the second machine-learned model is a scene-agnostic transformer model trained on images of a plurality of other scenes.

10 . The computer-implemented method of claim 1 , wherein the virtual element is displayed as part of an augmented reality experience in a game.

11 . A non-transitory computer-readable medium storing instructions that, when executed by a computing system, cause the computing system to perform operations comprising:

obtaining an image of a scene captured by a camera;

inputting the image into a first machine-learned model predicts a 3D coordinate map of the scene by mapping 2D pixels from the image to 3D scene coordinates;

inputting the 3D coordinate map of the scene into a second machine-learned model trained to estimate a pose of the camera when capturing the image of the scene based on the 3D coordinate map; and

causing display of a virtual element at a position on a display of a client device, the position determined based on the pose of the camera estimated by the second machine-learned model.

12 . The non-transitory computer-readable medium of claim 11 , wherein the image is an RGB image, and wherein the second machine-learned model estimates the pose of the camera without requiring a depth map or a point-cloud model of the image of the scene captured by the camera.

13 . The non-transitory computer-readable medium of claim 11 , wherein the first machine-learned model is a scene-specific geometry-based prediction model that processes the captured image to predict the 3D coordinate map, the first machine-learned model having been trained specifically for the scene based on training data associated with the scene.

14 . The non-transitory computer-readable medium of claim 13 , wherein the first machine-learned model is a convolutional neural network.

15 . The non-transitory computer-readable medium of claim 1 , wherein the operations further comprise:

inputting a camera intrinsic matrix associated with the captured image into the second machine-learned model, wherein the second machine-learned model is a scene-agnostic pose regressor model that directly regresses a multi-dimensional pose vector for the captured image of the scene based on the 3D coordinate map and the camera intrinsic matrix associated with the captured image.

16 . The non-transitory computer-readable medium of claim 15 , wherein the second machine-learned model is a transformer.

17 . The non-transitory computer-readable medium of claim 15 , wherein the operations further comprise:

training the second machine-learned model using training data including 3D coordinate maps of a plurality of historical scenes and corresponding ground truth camera poses, wherein the second machine-learned model is agnostic to the scene.

18 . The non-transitory computer-readable medium of claim 1 , wherein the pose of the camera indicates a position and an orientation of the camera relative to the 3D coordinate map of the scene, and wherein the operations further comprise:

transmitting game data to the client device including the camera, the game data based on the pose of the camera.

19 . The non-transitory computer-readable medium of claim 11 , wherein the first machine-learned model is a scene-specific geometry-based prediction model trained on images of the scene and the second machine-learned model is a scene-agnostic transformer model trained on images of a plurality of other scenes.

20 . A client device for detecting a camera pose, the client device comprising:

a display;

a camera;

one or more processors; and

memory storing a first and second machine-learned models, the memory further storing instructions that, when executed by the one or more processors, cause the one or more processors to perform operations comprising:

capturing an image of a scene by the camera;

inputting the image into a first machine-learned model predicts a 3D coordinate map of the scene by mapping 2D pixels from the image to 3D scene coordinates;

inputting the 3D coordinate map of the scene into a second machine-learned model trained to estimate a pose of the camera when capturing the image of the scene based on the 3D coordinate map; and

transmitting the estimated pose of the camera to a game server;

receiving game data from the game server, the game data based on transmitted pose of the camera; and

presenting a virtual element in a virtual world of an augmented reality game on the display of the client device based on the received game data.

Assignments (2)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jun 16, 2025
From: NIANTIC, INC.
To: NIANTIC SPATIAL, INC.
Reel/Frame 071555/0833 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Apr 4, 2025
From: CHEN, SHUAI; CAVALLARI, TOMMASO; PRISACARIU, VICTOR ADRIAN; BRACHMANN, ERIC
To: NIANTIC, INC.
Reel/Frame 070735/0122 →