IP Library › Granted Patent US 12,530,788
Granted Patent B2
US 12,530,788 · App. 17/786,065 · Granted Jan 20, 2026

System and methods for depth estimation

Inventors: Orly Liba (Mountain View, CA); Rahul Garg (Mountain View, CA); Neal Wadhwa (Mountain View, CA); Jon Barron (Mountain View, CA); Hayato Ikoma (Mountain View, CA)
Assignee: Google LLC
G06T7/50G06T2207/10028G06T2207/20081G06T2207/20084
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,530,788
App. No.
17/786,065
Granted
Jan 20, 2026
Kind
B2
Abstract

A system includes a computing device. The computing device is configured to perform a set of functions. The set of functions includes receiving an image, wherein the image comprises a two-dimensional array of data. The set of functions includes extracting, by a two-dimensional neural network, a plurality of two-dimensional features from the two-dimensional array of data. The set of functions includes generating a linear combination of the plurality of two-dimensional features to form a single three-dimensional input feature. The set of functions includes extracting, by a three-dimensional neural network, a plurality of three-dimensional features from the single three-dimensional input feature. The set of functions includes determining a two-dimensional depth map. The two-dimensional depth map contains depth information corresponding to the plurality of three-dimensional features.

Claims (44)

1 . A system for determining a two-dimensional depth map from a single monocular image, the system comprising:

a computing device, wherein the computing device is configured to perform a set of functions comprising:

receiving a single monocular image, wherein the image comprises a two-dimensional array of data;

extracting, by a two-dimensional neural network, a plurality of two-dimensional features from the two-dimensional array of data;

generating a linear combination of the plurality of two-dimensional features to form a single three-dimensional input feature;

extracting, by a three-dimensional neural network, a plurality of three-dimensional features from the single three-dimensional input feature; and

determining a two-dimensional depth map, wherein the two-dimensional depth map contains depth information corresponding to the plurality of three-dimensional features.

2 . The system of claim 1 , wherein the computing device is a first computing device of a plurality of computing devices, and wherein the two-dimensional neural network and the three-dimensional neural network correspond to at least a second computing device of the plurality of computing devices.

3 . The system of claim 1 , wherein the two-dimensional neural network comprises a two-dimensional convolutional neural network, wherein the three-dimensional neural network comprises a three-dimensional convolutional neural network, and wherein extracting the plurality of two-dimensional features comprises using the two-dimensional convolutional neural network as a two-dimensional filter that operates in two directions within the two-dimensional array of data to output the plurality of two-dimensional features.

4 . The system of claim 3 , the set of functions further comprising:

prior to extracting the plurality of two-dimensional features, training the two-dimensional convolutional neural network using a plurality of images representing objects such that different nodes within the two-dimensional convolutional neural network operate to output different types of two-dimensional features corresponding to different objects.

5 . The system of claim 1 , wherein generating the linear combination of the plurality of two-dimensional features to form the single three-dimensional input feature comprises:

classifying a two-dimensional feature of the plurality of two-dimensional features in accordance with an object associated with training the two-dimensional convolutional neural network; and

generating the linear combination of the plurality of two-dimensional features based on classifying the two-dimensional feature.

6 . The system of claim 1 , wherein the two-dimensional neural network comprises a two-dimensional convolutional neural network, wherein the three-dimensional neural network comprises a three-dimensional convolutional neural network, and wherein extracting the plurality of three-dimensional features comprises using the three-dimensional convolutional neural network as a three-dimensional filter that operates in three directions within the three-dimensional input feature to output the plurality of three-dimensional features.

7 . The system of claim 1 , wherein the two-dimensional neural network comprises a two-dimensional convolutional neural network, wherein the three-dimensional neural network comprises a three-dimensional convolutional neural network, and wherein extracting the plurality of three-dimensional features from the single three-dimensional input feature comprises extracting a plurality of sets of voxels, wherein each voxel indicates a level of opaqueness.

8 . A method for determining a two-dimensional depth map from a single monocular image, the method comprising:

receiving a single monocular image, wherein the image comprises a two-dimensional array of data;

extracting, by a two-dimensional neural network, a plurality of two-dimensional features from the two-dimensional array of data;

generating a linear combination of the plurality of two-dimensional features to form a single three-dimensional input feature;

extracting, by a three-dimensional neural network, a plurality of three-dimensional features from the single three-dimensional input feature; and

determining a two-dimensional depth map, wherein the two-dimensional depth map contains depth information corresponding to the plurality of three-dimensional features.

9 . The method of claim 8 , further comprising:

determining, based on the plurality of three-dimensional features extracted by the three-dimensional neural network, a three-dimensional array of voxels, each voxel of the array indicating a respective level of opaqueness, and

wherein determining a two-dimensional depth map comprises, for a plurality of pixels of the two-dimensional depth map, determining respective distances between a capture device location and a respective closest opaque voxel of the array of voxels along a respective different path from the capture device location.

10 . The method of claim 8 , wherein each two-dimensional feature extracted by the two-dimensional neural network represents a respective different objects in the image.

11 . The method of claim 10 , wherein generating the linear combination of the plurality of two-dimensional features to form the single three-dimensional input feature comprises ordering the two-dimensional features based on overlap between the respective different objects in the image.

12 . The method of claim 8 , wherein the two-dimensional neural network comprises a two-dimensional convolutional neural network, wherein the three-dimensional neural network comprises a three-dimensional convolutional neural network, and wherein extracting the plurality of two-dimensional features comprises using the two-dimensional convolutional neural network as a two-dimensional filter that operates in two directions within the two-dimensional array of data to output the plurality of two-dimensional features.

13 . The method of claim 12 , further comprising:

prior to extracting the plurality of two-dimensional features, training the two-dimensional convolutional neural network using a plurality of images representing objects such that different nodes within the two-dimensional convolutional neural network operate to output different types of two-dimensional features corresponding to different objects.

14 . The method of claim 13 , wherein generating the linear combination of the plurality of two-dimensional features to form the single three-dimensional input feature comprises:

classifying a two-dimensional feature of the plurality of two-dimensional features in accordance with an object associated with training the two-dimensional convolutional neural network; and

generating the linear combination of the plurality of two-dimensional features based on classifying the two-dimensional feature.

15 . The method of claim 8 , wherein the two-dimensional neural network comprises a two-dimensional convolutional neural network, wherein the three-dimensional neural network comprises a three-dimensional convolutional neural network, and wherein extracting the plurality of three-dimensional features comprises using the three-dimensional convolutional neural network as a three-dimensional filter that operates in three directions within the three-dimensional input feature to output the plurality of three-dimensional features.

16 . The method of claim 8 , wherein extracting the plurality of three-dimensional features from the single three-dimensional input feature comprises extracting a plurality of sets of voxels, wherein each voxel indicates a level of opaqueness.

17 . The method of claim 16 , wherein determining the two-dimensional depth map comprises determining a plurality of path lengths, wherein each path length represents a distance between a focal point and an opaque voxel.

18 . The method of claim 8 , wherein generating the linear combination of the plurality of two-dimensional features to form the single three-dimensional input feature corresponds to ordering two-dimensional features.

19 . The method of claim 8 , wherein the plurality of two-dimensional features corresponds to a multi-channel output of the two-dimensional neural network, and wherein generating the linear combination of the plurality of two-dimensional features to form the single three-dimensional input feature comprises transforming the multi-channel output of the two-dimensional neural network into a single-channel input of the three-dimensional neural network.

20 . A non-transitory computer readable medium having instructions stored thereon that when executed by a processor cause performance of a set of functions, wherein the set of functions is for determining a two-dimensional depth map from a single monocular image and comprises:

receiving a single monocular image, wherein the image comprises a two-dimensional array of data;

extracting, by a two-dimensional neural network, a plurality of two-dimensional features from the two-dimensional array of data;

generating a linear combination of the plurality of two-dimensional features to form a single three-dimensional input feature;

extracting, by a three-dimensional neural network, a plurality of three-dimensional features from the single three-dimensional input feature; and

determining a two-dimensional depth map, wherein the two-dimensional depth map contains depth information corresponding to the plurality of three-dimensional features.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Nov 21, 2025
From: LIBA, ORLY; GARG, RAHUL; WADHWA, NEAL; BARRON, JON; IKOMA, HAYATO
To: GOOGLE LLC
Reel/Frame 072999/0913 →
Continuity (2)
Provisional Application 62954392 · Dec 27, 2019
Related Publication 20230037958A1 · Feb 9, 2023
References Cited (14)
US 10790056B1 · Accomazzi · 2020 [cited by examiner]
US 10929654B2 · Iqbal · 2021 [cited by examiner]
US 20170132758A1 · Paluri · 2017 [cited by examiner]
US 20190026956A1 · Gausebeck et al. · 2019 [cited by applicant]
US 20190220992A1 · Li et al. · 2019 [cited by applicant]
US 20200273192A1 · Cheng · 2020 [cited by examiner]
US 20200327718A1 · Saragih · 2020 [cited by examiner]
CN 108898669A · 2018 [cited by applicant]
EP 3166075A1 · 2017 [cited by applicant]
International Search Report for International Application No. PCT/US2020/067044 dated Apr. 1, 2021, 2 pages. [cited by applicant]
Cheng et al., “Learning Depth with Convolutional Spatial propagation Network”, Oct. 4, 2019, pp. 1-18. [cited by applicant]
MWITI: “Research Guide for Depth Estimation with Deep Learning”, Sep. 25, 2019, pp. 1-25. [cited by applicant]
Zhao et al., “Pyramid Scene Parsing Network”, Dec. 4, 2016, 11 pages. [cited by applicant]
Shi et al., “Feature Enhanced Fully Convolutional Networks for Monocular Depth Estimation”, 2019 IEEE International Conference on Data Science and Advanced Analytics (DSAA), 7 pages. [cited by applicant]