IP Library Granted Patent US 11,017,542
Granted Patent B2
US 11,017,542 · App. 16/229,808 · Granted May 25, 2021

Systems and methods for determining depth information in two-dimensional images

Inventors: Zafar Takhirov (Santa Clara, CA); Yun Jiang (Mountain View, CA)
Assignee: BEIJING VOYAGER TECHNOLOGY CO., LD.
G06T7/50G06K9/4604G06N5/046G06T2207/20081
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,017,542
App. No.
16/229,808
Granted
May 25, 2021
Kind
B2
Abstract

Embodiments of the disclosure provide systems and methods for determining depth information in a two-dimensional (2D) image. An exemplary system may include a processor and a non-transitory memory storing instructions that, when executed by the processor, cause the system to perform the various operations. The operations may include receiving a first feature map based on the 2D image and applying an extraction network having a convolution operation and a pooling operation to the first feature map to obtain a second feature map. The operations may also include applying a reconstruction network having a deconvolution operation to the second feature map to obtain a depth map.

Claims (56)

1. A system for determining depth information in a two-dimensional (2D) image, comprising:

at least one processor; and

at least one non-transitory memory storing instructions that, when executed by the at least one processor, cause the system to perform operations comprising:

receiving a first feature map based on the 2D image;

applying an extraction network comprising at least one convolution operation and at least one pooling operation to the first feature map to obtain a second feature map; and

applying a reconstruction network comprising at least one deconvolution operation and at least one unpooling operation to the second feature map, wherein an output of the deconvolution operation and the unpooling operation is a 3D depth map.

2. The system of claim 1 , comprising an image capturing component configured to capture the 2D image.

3. The system of claim 1 , comprising a feature extractor configured to extract at least one feature from the 2D image to generate the first feature map.

4. The system of claim 1 , wherein features in the first or second feature map include at least one of lines, edges, curves, circles, squares, corners, or texture.

5. The system of claim 1 , wherein:

the extraction network comprises a plurality of layers, each layer comprising at least one convolution operation and one pooling operation; and

wherein the operations comprise:

applying the plurality of layers of the extraction network to the first feature map sequentially to obtain intermediate results after each of the plurality of layers.

6. The system of claim 5 , wherein:

the reconstruction network comprises a plurality of layers, each layer comprising at least one deconvolution operation; and

where the operations comprise:

applying the plurality of layers of the reconstruction network to the second feature map sequentially to obtain intermediate results after each of the plurality of layers.

7. The system of claim 6 , wherein the operations comprise:

applying a convolution operation to the corresponding intermediate results obtained by the extraction network and the reconstruction network to obtain preprocessors; and

concatenating multiple preprocessors.

8. The system of claim 7 , wherein the operations comprise:

classifying objects in one or more cells of the 2D image based on the concatenation of the preprocessors.

9. The system of claim 7 , wherein the operations comprise:

estimating bounding boxes of objects in one or more cells of the 2D image based on the concatenation of the preprocessors.

10. The system of claim 7 , wherein the operations comprise:

applying a further convolution operation to the preprocessors; and

estimating 3D parameters of objects in the 2D image based on the further convoluted preprocessors.

11. The system of claim 1 , wherein the operations comprise:

training the extraction network and the reconstruction network using a training dataset.

12. The system of claim 1 , wherein a dimension of the second feature map is smaller than a dimension of the first feature map.

13. A method for determining depth information in a two-dimensional (2D) image, comprising:

receiving, from a feature extractor, a first feature map based on the 2D image;

applying, by a processor, an extraction network comprising at least one convolution operation and at least one pooling operation to the first feature map to obtain a second feature map; and

applying, by the processor, a reconstruction network comprising at least one deconvolution operation and at least one unpooling operation to the second feature map, wherein an output of the deconvolution operation and the unpooling operation is a 3D depth map.

14. The method of claim 13 , wherein:

the extraction network comprises a plurality of layers, each layer comprising at least one convolution operation and one pooling operation;

the reconstruction network comprises a plurality of layers, each layer comprising at least one deconvolution operation; and

the method further comprises:

applying, by the processor, the plurality of layers of the extraction network to the first feature map sequentially to obtain intermediate results after each of the plurality of layers;

applying, by the processor, the plurality of layers of the reconstruction network to the second feature map sequentially to obtain intermediate results after each of the plurality of layers;

applying, by the processor, a convolution operation to the corresponding intermediate results obtained by the extraction network and the reconstruction network to obtain preprocessors; and

concatenating, by the processor, multiple preprocessors.

15. The method of claim 14 , further comprising:

classifying, by the processor, objects in one or more cells of the 2D image based on the concatenation of the preprocessors.

16. The method of claim 14 , further comprising:

estimating, by the processor, bounding boxes of objects in one or more cells of the 2D image based on the concatenation of the preprocessors.

17. The method of claim 14 , further comprising:

applying, by the processor, a further convolution operation to the preprocessors; and

estimating, by the processor, 3D parameters of objects in the 2D image based on the further convoluted preprocessors.

18. The method of claim 13 , further comprising:

training the extraction network and the reconstruction network using a training dataset.

19. The method of claim 13 , wherein a dimension of the second feature map is smaller than a dimension of the first feature map.

20. A non-transitory computer-readable medium having instructions stored thereon that, when executed by one or more processors, causes the one or more processors to perform a method for determining depth information in a two-dimensional (2D) image, the method comprising:

receiving, from a feature extractor, a first feature map based on the 2D image;

applying, by the one or more processors, an extraction network comprising at least one convolution operation and at least one pooling operation to the first feature map to obtain a second feature map; and

applying, by the one or more processors, a reconstruction network comprising at least one deconvolution operation and at least one unpooling operation to the second feature map, wherein an output of the deconvolution operation and the unpooling operation is a 3D depth map.

Assignments (3)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Mar 20, 2020
From: DIDI RESEARCH AMERICA, LLC
To: VOYAGER (HK) CO., LTD.
Reel/Frame 052182/0481 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Mar 20, 2020
From: VOYAGER (HK) CO., LTD.
To: BEIJING VOYAGER TECHNOLOGY CO., LTD.
Reel/Frame 052182/0896 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Dec 21, 2018
From: TAKHIROV, ZAFAR; JIANG, YUN
To: DIDI RESEARCH AMERICA, LLC
Reel/Frame 047844/0091 →
Continuity (1)
Related Publication 20200202542A1 · Jun 25, 2020