IP Library › Granted Patent US 12,243,260
Granted Patent B2
US 12,243,260 · App. 17/879,307 · Granted Mar 4, 2025

Producing a depth map from a monocular two-dimensional image

Inventors: Vitor Guizilini (Santa Clara, CA); Rares A. Ambrus (San Francisco, CA); Dian Chen (Mountain View, CA); Adrien David Gaidon (Mountain View, CA); Sergey Zakharov (San Francisco, CA)
Assignees: Toyota Research Institute, Inc.; Toyota Jidosha Kabushiki Kaisha
G06T7/73G06T7/55G06T7/593G06T2207/10012G06T2207/10028G06T2207/20076G06T2207/20081G06T2207/20084
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,243,260
App. No.
17/879,307
Granted
Mar 4, 2025
Kind
B2
Abstract

A system for producing a depth map can include a processor and a memory. The memory can store a neural network. The neural network can include an encoding portion module, a multi-frame feature matching portion module, and a decoding portion module. The encoding portion module can include instructions that, when executed by the processor, cause the processor to encode an image to produce single-frame features. The multi-frame feature matching portion module can include instructions that, when executed by the processor, cause the processor to process the single-frame features to produce information. The decoding portion module can include instructions that, when executed by the processor, cause the processor to decode the information to produce the depth map. A first training dataset, used to train the multi-frame feature matching portion module, can be different from a second training dataset used to train the encoding portion module and the decoding portion module.

Claims (48)

1. A system, comprising:

a first processor; and

a first memory storing:

a neural network, the neural network including:

a first encoding portion module including instructions that, when executed by the first processor, cause the first processor to encode an image to produce single-frame features;

a multi-frame feature matching portion module including instructions that, when executed by the first processor, cause the first processor to process the single-frame features to produce information; and

a decoding portion module including instructions that, when executed by the first processor, cause the first processor to decode the information to produce a depth map, wherein a first training dataset, used to train the multi-frame feature matching portion module, is different from a second training dataset used to train the encoding portion module and the decoding portion module.

2. The system of claim 1 , wherein:

the image is of a scene that includes an object, and

the first memory further stores a distance determination module including instructions that, when executed by the first processor, cause the first processor to determine, using the depth map, a distance to the object.

3. The system of claim 1 , wherein:

the information comprises an input information and an output information,

the instructions to process the single-frame features to produce the information include instructions to process the single-frame features to produce the input information,

the instructions to decode the information to produce the depth map include instructions to decode the output information to produce the depth map, and

the neural network further includes a second encoding portion module including instructions that, when executed by the first processor, cause the first processor to encode the input information to produce the output information.

4. The system of claim 1 , wherein the multi-frame feature matching portion module includes a cost volume of similarity measures between features of a first training image and features of a second training image, the first training image being a first frame in a sequence of frames, and the second training image being a second frame in the sequence of frames.

5. The system of claim 4 , wherein for a pixel of the first training image, a similarity measure, of the similarity measures, is determined using a first attention technique with respect to features, associated with a corresponding pixel of the second training image, and features associated with a set of candidate depths determined from the first training image.

6. The system of claim 5 , wherein the set of candidate depths is associated with pixels, in the first training image, along an epipolar line associated with the corresponding pixel in the second training image.

7. The system of claim 6 , wherein:

members of the set of candidate depths include a first member, a second member, and a third member, the third member being directly adjacent to the second member, and the second member being directly adjacent to the first member, and

a relationship between a first distance and a second distance is a logarithmic relationship, a first distance being a distance between the first member and the second member, the second distance being a distance between the second member and the third member.

8. The system of claim 5 , wherein for the pixel of the first training image, the similarity measure, of the similarity measures, is further determined using a second attention technique with respect to the features, associated with one member of the set of candidate depths, and the features associated with other members of the set of candidate depths.

9. The system of claim 1 , further comprising:

a second processor; and

a second memory storing a first training operation module including instructions that, when executed by the second processor, cause the second processor to cause a first geometry-based task training operation to be performed on the multi-frame feature matching portion module.

10. The system of claim 9 , wherein:

the second processor is the first processor; and

the second memory is the first memory.

11. The system of claim 9 , wherein the first memory further stores an installation module including instructions that, when executed by the first processor, cause the first processor to cause, after a completion of the first geometry-based task training operation, the multi-frame feature matching portion module to be included in the neural network.

12. The system of claim 11 , further comprising:

a third processor; and

a third memory storing a second training operation module including instructions that, when executed by the third processor, cause the third processor to cause, using the second training dataset, an image-based task training operation to be performed on the first encoding portion module and the decoding portion module.

13. The system of claim 12 , wherein:

the third processor is the first processor; and

the third memory is the first memory.

14. The system of claim 12 , wherein the instructions to cause the image-based task training operation to be performed on the first encoding portion module and the decoding portion module include instructions to cause, after the multi-frame feature matching portion module has been caused to be included in the neural network, the image-based task training operation to be performed on the first encoding portion module and the decoding portion module.

15. The system of claim 12 , wherein the instructions to cause the image-based task training operation to be performed on the first encoding portion module and the decoding portion module include instructions to cause, in a manner that causes weights associated with the multi-frame feature matching portion module to remain unchanged, the image-based task training operation to be performed on the first encoding portion module and the decoding portion module.

16. The system of claim 12 , wherein the second training operation module further includes instructions that, when executed by the third processor, cause the third processor to cause, using the second training dataset, a second geometry-based task training operation to be performed on the multi-frame feature matching portion module.

17. The system of claim 16 , wherein the second geometry-based task training operation is performed concurrently with the image-based task training operation.

18. A method, comprising:

encoding, by a processor operating an encoding portion of a neural network, an image to produce single-frame features;

processing, by the processor operating a multi-frame feature matching portion of the neural network, the single-frame features to produce information; and

decoding, by the processor operating a decoding portion of the neural network, the information to produce a depth map, wherein a first training dataset, used to train the multi-frame feature matching portion, is different from a second training dataset used to train the encoding portion and the decoding portion.

19. The method of claim 18 , wherein the depth map comprises a set of depth maps, a scale of a first depth map, in the set of depth maps, being different from a scale of a second depth map in the set of depth maps.

20. A non-transitory computer-readable medium for producing a depth map, the non-transitory computer-readable medium including instructions that, when executed by one or more processors, cause the one or more processors to:

encode, operating an encoding portion of a neural network, an image to produce single-frame features;

process, operating a multi-frame feature matching portion of the neural network, the single-frame features to produce information; and

decode, operating a decoding portion of the neural network, the information to produce the depth map, wherein a first training dataset, used to train the multi-frame feature matching portion, is different from a second training dataset used to train the encoding portion and the decoding portion.

Assignments (2)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded May 7, 2025
From: TOYOTA RESEARCH INSTITUTE, INC.
To: TOYOTA JIDOSHA KABUSHIKI KAISHA
Reel/Frame 071044/0244 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Oct 12, 2022
From: GUIZILINI, VITOR; AMBRUS, RARES A.; CHEN, DIAN; GAIDON, ADRIEN DAVID; ZAKHAROV, SERGEY
To: TOYOTA RESEARCH INSTITUTE, INC.; TOYOTA JIDOSHA KABUSHIKI KAISHA
Reel/Frame 061385/0408 →
Continuity (3)
Provisional Application 63279823 · Nov 16, 2021
Provisional Application 63279404 · Nov 15, 2021
Related Publication 20230154024A1 · May 18, 2023
References Cited (41)
US 11176709B2 · Pillai · 2021 [cited by examiner]
US 11803950B2 · Li · 2023 [cited by examiner]
US 11948309B2 · Guizilini · 2024 [cited by examiner]
US 11948310B2 · Guizilini · 2024 [cited by examiner]
US 20170337711A1 · Ratner · 2017 [cited by examiner]
US 20180295375A1 · Ratner · 2018 [cited by examiner]
US 20190347538A1 · Lee et al. · 2019 [cited by applicant]
US 20200372660A1 · Li et al. · 2020 [cited by applicant]
US 20210118184A1 · Pillai · 2021 [cited by examiner]
US 20210150757A1 · Mustikovela · 2021 [cited by examiner]
US 20220383530A1 · Giryes · 2022 [cited by examiner]
US 20220392083A1 · Guizilini · 2022 [cited by examiner]
US 20220392089A1 · Guizilini · 2022 [cited by examiner]
US 20240020710A1 · Ettl · 2024 [cited by examiner]
CN 109554559A · 2019 [cited by applicant]
CN 110120049A · 2019 [cited by applicant]
Wang et al. “Deep Visual Domain Adaptation: A Survey,” Neurocomputing, vol. 312, No. 27, 2018, pp. 135-153. [cited by applicant]
Islam et al., “A new algorithm to design compact two-hidden-layer artificial neural networks,” Neural Networks, vol. 14, No. 9, pp. 1265-1278, 2001. [cited by applicant]
Watson et al., “The Temporal Opportunist: Self-Supervised Multi-Frame Monocular Depth,” 2021 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2021, pp. 1164-1174. [cited by applicant]
Isikodogan et al., “SemifreddoNets: Partially Frozen Neural Networks for Efficient Computer Vision Systems,” Retrieved from arxiv.org/abs/2006.06888.v1 [cs.CV] Jun. 12, 2020 pp. 1-16. [cited by applicant]
Zhao et al., “Monocular Depth Estimation Based on Deep Learning: An Overview,” Sci. China Technol. Sci. No. 63, Mar. 14, 2020, pp. 1612-1627. [cited by applicant]
Evain et al., “A Lightweight Neural Network for Monocular View Generation with Occlusion Handling,” IEEE Transactions on Pattern Analysis & Machine Intelligence, vol. 43, No. 6, Jul. 24, 2020, pp. 1832-1844. [cited by applicant]
Guizilini et al., “Geometric Unsupervised Domain Adaptation for Semantic Segmentation,” Retrieved from arXiv:2103.16694v2 [cs.CV] Aug. 18, 2021, pp. 1-15. [cited by applicant]
Antequera et al., “Mapillary Planet-Scale Depth Dataset,” In European Conference on Computer Vision, Aug. 23, 2020, pp. 589-604. [cited by applicant]
Huynh et al., “Guiding Monocular Depth Estimation Using Depth-Attention Volume,” In ECCV, Aug. 23-28, 2020, pp. 1-17. [cited by applicant]
Sadek et al., “Self-Supervised Attention Learning for Depth and Ego-Motion Estimation,” 2020 IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS), Apr. 27, 2020, pp. 1-7. [cited by applicant]
Lee et al., “Patch-wise attention network for monocular depth estimation,” Association for the Advancement of Artificial Intelligence Conference, Feb. 2021, pp. 1873-1881. [cited by applicant]
Ranftl et al., “Vision Transformers for Dense Prediction,” Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV), Mar. 24, 2021, pp. 12179-12188. [cited by applicant]
Johnston et al., “Self-supervised Monocular Trained Depth Estimation using Self-attention and Discrete Disparity Volume,” CVPR 2020, Computer Vision Foundation, pp. 4756-4765. [cited by applicant]
Li et al., “Revisiting Stereo Depth Estimation From a Sequence-to-Sequence Perspective with Transformers,” Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV), Aug. 25, 2021, pp. 1-15. [cited by applicant]
Ruhkamp et al., “Attention meets Geometry: Geometry Guided Spatial-Temporal Attention for Consistent Self-Supervised Monocular Depth Estimation,” 2021 International Conference on 3D Vision (3DV), Oct. 15, 2021, pp. 1-11. [cited by applicant]
Im et al., “DPSNet: End-to-end Deep Plane Sweep Stereo,” Retrieved from: https://arXiv:1905.00538v1 [cs.CV] May 2, 2019, pp. 1-12. [cited by applicant]
Wu et al., “Semantic Stereo Matching with Pyramid Cost Volumes,” 2019 IEEE/CVF International Conference on Computer Vision (ICCV), 2019, pp. 7484-7493. [cited by applicant]
Xu et al., “Structured Attention Guided Convolutional Neural Fields for Monocular Depth Estimation,” 2018 IEEE/CVF Conference on Computer Vision and Pattern Recognition, Mar. 29, 2018, pp. 3917-3925. [cited by applicant]
Wimbauer et al., “MonoRec: Semi-Supervised Dense Reconstruction in Dynamic Environments from a Single Moving Camera,” In CVPR, 2021, Computer Vision Foundation, pp. 6112-6122. [cited by applicant]
Unknown, “Bridging the Domain Gap for Neural Models,” Computer vision, Jun. 2019, 9 pages. [cited by applicant]
Unknown, “Diversity? Accuracy? Important Properties ofyour Dataset,” Jun. 20, 2020, 11 pages. [cited by applicant]
Unknown, “Domain adaptation,” last accessed Jan. 10, 2022, 5 pages, found at https://en.wikipedia.org/wiki/Domain_adaptation. [cited by applicant]
Zhang et al., “Joint Geometrical and Statistical Alignment for Visual Domain Adaptation,” Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 2017, pp. 1859-1867. [cited by applicant]
Nam et al., “Reducing Domain Gap by Reducing Style Bias,” Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2021, pp. 8690-8699. [cited by applicant]
Ram Sagar, “What Does Freezing A Layer Mean And How Does It HelpIn Fine Tuning Neural Networks,” May 25, 2019, 18 pages. [cited by applicant]