IP Library Granted Patent US 12,080,010
Granted Patent B2
US 12,080,010 · App. 17/545,201 · Granted Sep 3, 2024

Self-supervised multi-frame monocular depth estimation model

Inventors: James Watson (London, GB); Oisin MacAodha (Edinburgh, GB); Victor Adrian Prisacariu (London, GB); Gabriel J. Brostow (London, GB); Michael David Firman (London, GB)
Assignee: NIANTIC, INC.
G06T7/55G01B11/22G06T3/18G06T7/73G06T11/00G06T2207/10016G06T2207/20081G06T2207/20084
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,080,010
App. No.
17/545,201
Granted
Sep 3, 2024
Kind
B2
Abstract

A multi-frame depth estimation model is disclosed. The model is trained and configured to receive an input image and an additional image. The model outputs a depth map for the input image based on the input image and the additional image. The model may extract a feature map for the input image and an additional feature map for the additional image. For each of a plurality of depth planes, the model warps the feature map to the depth plane based on relative pose between the input image and the additional image, the depth plane, and camera intrinsics. The model builds a cost volume from the warped feature maps for the plurality of depth planes. A decoder of the model inputs the cost volume and the input image to output the depth map.

Claims (52)

1. A computer-implemented method comprising:

receiving a time series of images of a scene including a primary image and an additional image from an earlier time than the primary image, wherein the time series of images are monocular images derived from monocular video;

inputting the time series of images into a depth estimation model;

receiving, as output from the depth estimation model, a depth map of the primary image, the depth map generated based on a cost volume concatenating differences between a primary feature map of the primary image and a plurality of warped feature maps of the additional image for each of a plurality of depth planes, wherein receiving the depth map as output from the depth estimation model comprises;

generating a primary feature map for the primary image and an additional feature map for the additional image;

generating a warped feature map comprising a plurality of warped feature map layers, each warped feature map layer generated by warping the additional feature map to a plurality of depth planes based on (1) a depth plane of the plurality to which the feature map is being warped, (2) a relative pose between the primary image and the additional image, and (3) intrinsics of a camera used to capture the primary image and the additional image;

for each warped feature map layer, calculating a difference between the warped feature map layer and the primary feature map; and

building the cost volume by concatenating the differences between layers of the warped feature map and the primary feature map;

wherein the output is based on the cost volume and the primary feature map;

generating virtual content using the depth map; and

displaying an image of scene augmented with the virtual content.

2. The computer-implemented method of claim 1 , wherein the relative pose is determined by a convolutional neural network separately trained to determine the relative pose between two images.

3. The computer-implemented method of claim 1 , wherein the difference is an absolute difference.

4. The computer-implemented method of claim 1 , wherein the depth estimation model is trained by:

tuning a minimum depth plane and a maximum depth plane in the plurality of depth planes.

5. The computer-implemented method of claim 1 , wherein the depth estimation model is trained by:

inputting a training secondary image into a secondary depth estimation model to predict an estimated depth of the training secondary image;

generating a binary mask based on the estimated depth;

applying the binary mask to filter out unreliable pixels from the training secondary image; and

training the depth estimation model based on the training secondary image with filtered out unreliable pixels and a training input image.

6. The computer-implemented method of claim 1 , wherein the depth estimation model is trained by:

randomly, according to a set probability, replacing a training secondary image with a color augmented version of a training input image; and

training the depth estimation model based on the color augmented version of the training input image compared to the training input image.

7. The computer-implemented method of claim 1 , wherein the depth estimation model is trained by:

randomly, according to a set probability, setting a cost volume for a training secondary image to a constant value to generate a blank cost volume; and

training the depth estimation model based on the blank cost volume and a training input image.

8. A non-transitory computer-readable storage medium storing instructions that, when executed by a processor, cause the processor to perform operations comprising:

receiving a time series of images of a scene including a primary image and an additional image from an earlier time than the primary image;

inputting the time series of images into a depth estimation model, wherein the time series of images are monocular images derived from monocular video;

receiving, as output from the depth estimation model, a depth map of the primary image, the depth map generated based on a cost volume concatenating differences between a primary feature map of the primary image and a plurality of warped feature maps of the additional image for each of a plurality of depth planes, wherein receiving the depth map as output from the depth estimation model comprises;

generating a primary feature map for the primary image and an additional feature map for the additional image;

generating a warped feature map comprising a plurality of warped feature map layers, each warped feature map layer generated by warping the additional feature map to a plurality of depth planes based on (1) a depth plane of the plurality to which the feature map is being warped, (2) a relative pose between the primary image and the additional image, and (3) intrinsics of a camera used to capture the primary image and the additional image;

for each warped feature map layer, calculating a difference between the warped feature map layer and the primary feature map; and

building the cost volume by concatenating the differences between layers of the warped feature map and the primary feature map;

wherein the output is based on the cost volume and the primary feature map;

generating virtual content using the depth map; and

displaying an image of scene augmented with the virtual content.

9. The non-transitory computer-readable storage medium of claim 8 , wherein the relative pose is determined by a convolutional neural network separately trained to determine the relative pose between two images.

10. The non-transitory computer-readable storage medium of claim 8 , wherein the difference is an absolute difference.

11. The non-transitory computer-readable storage medium of claim 8 , wherein the depth estimation model is trained by:

tuning a minimum depth plane and a maximum depth plane in the plurality of depth planes.

12. The non-transitory computer-readable storage medium of claim 8 , wherein the depth estimation model is trained by:

inputting a training secondary image into a secondary depth estimation model to predict an estimated depth of the training secondary image;

generating a binary mask based on the estimated depth;

applying the binary mask to filter out unreliable pixels from the training secondary image; and

training the depth estimation model based on the training secondary image with filtered out unreliable pixels and a training input image.

13. The non-transitory computer-readable storage medium of claim 8 , wherein the depth estimation model is trained by:

randomly, according to a set probability, replacing a training secondary image with a color augmented version of a training input image; and

training the depth estimation model based on the color augmented version of the training input image compared to the training input image.

14. The non-transitory computer-readable storage medium of claim 8 , wherein the depth estimation model is trained by:

randomly, according to a set probability, setting a cost volume for a training secondary image to a constant value to generate a blank cost volume; and

training the depth estimation model based on the blank cost volume and a training input image.

Assignments (3)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jun 16, 2025
From: NIANTIC, INC.
To: NIANTIC SPATIAL, INC.
Reel/Frame 071555/0833 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jul 13, 2023
From: NIANTIC INTERNATIONAL TECHNOLOGY LIMITED
To: NIANTIC, INC.
Reel/Frame 064249/0011 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Sep 16, 2022
From: WATSON, JAMES; AODHA, OISIN MAC; PRISACARIU, VICTOR ADRIAN; BROSTOW, GABRIEL J.; FIRMAN, MICHAEL DAVID
To: NIANTIC INTERNATIONAL TECHNOLOGY LIMITED
Reel/Frame 061124/0646 →
Continuity (2)
Provisional Application 63124757 · Dec 12, 2020
Related Publication 20220189049A1 · Jun 16, 2022
Cited By (2)
US 12,428,022 US 12,567,162