IP Library › Granted Patent US 12,020,489
Granted Patent B2
US 12,020,489 · App. 17/333,537 · Granted Jun 25, 2024

Network architecture for monocular depth estimation and object detection

Inventors: Dennis Park (Fremont, CA); Rares A. Ambrus (San Francisco, CA); Vitor Guizilini (Santa Clara, CA); Jie Li (Los Altos, CA); Adrien David Gaidon (Mountain View, CA)
Assignee: Toyota Research Institute, Inc.
G06V20/58G01S17/42G01S17/89G01S17/931G06F18/2113G06F18/2155G06F18/217G06F18/251G06N3/04G06N3/08G06N20/00G06T7/10G06T7/11G06T7/50G06V10/462G06V10/757G06V20/56G06T2207/10024G06T2207/10028G06T2207/20016G06T2207/20081G06T2207/20084G06T2207/30248
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,020,489
App. No.
17/333,537
Filed
May 28, 2021
Granted
Jun 25, 2024
Kind
B2
Art Unit
2663
USPC
382/100
Abstract

Systems, methods, and other embodiments described herein relate to performing depth estimation and object detection using a common network architecture. In one embodiment, a method includes generating, using a backbone of a combined network, a feature map at multiple scales from an input image. The method includes decoding, using a top-down pathway of the combined network, the feature map to provide features at the multiple scales. The method includes generating, using a head of the combined network, a depth map from the features for a scene depicted in the input image, and bounding boxes identifying objects in the input image.

Claims (35)

1. A depth system, comprising:

one or more processors; and

a memory communicably coupled to the one or more processors and storing:

a network module including instructions that, when executed by the one or more processors, cause the one or more processors to:

generate, using a backbone of a combined network, a feature map at multiple scales from an input image;

decode, using a top-down pathway of the combined network, the feature map to provide features at the multiple scales;

generate, using a head of the combined network, a depth map from the features for a scene depicted in the input image and bounding boxes identifying objects in the input image, wherein the head includes multiple sub-heads for each of the multiple scales that perform 3D object detection, 2D object detection, depth estimation, and classification;

train, in a first stage, the combined network by using a supervised depth loss derived from the depth map; and

train, in a second stage, the combined network by using the bounding boxes and ground-truth data to compute a detection loss.

2. The depth system of claim 1 , wherein the network module includes instructions to decode including instructions to provide, using lateral connections between the backbone and the top-down pathway, the multiple scales of the feature map in addition to an output of a prior level from within the top-down pathway.

3. The depth system of claim 1 , wherein the network module includes instructions to generate the feature map including instructions to generate the feature map at the multiple scales as a feature hierarchy, and

wherein the network module includes instructions to generate the feature map to encode features of the input image to provide a common reference for generating the depth map and the bounding boxes.

4. The depth system of claim 1 , wherein the network module includes instructions to generate the depth map and the bounding boxes including instructions to use the head among separate layers of the top-down pathway at the multiple scales to generate the bounding boxes at the multiple scales and the depth map at one of the multiple scales.

5. The depth system of claim 1 , wherein the input image is a monocular image in RGB.

6. A non-transitory computer-readable medium including instructions that, when executed by one or more processors, cause the one or more processors to:

generate, using a backbone of a combined network, a feature map at multiple scales from an input image;

decode, using a top-down pathway of the combined network, the feature map to provide features at the multiple scales; and

generate, using a head of the combined network, a depth map from the features for a scene depicted in the input image and bounding boxes identifying objects in the input image, wherein the head includes multiple sub-heads for each of the multiple scales that perform 3D object detection, 2D object detection, depth estimation, and classification;

train, in a first stage, the combined network by using a supervised depth loss derived from the depth map; and

train, in a second stage, the combined network by using the bounding boxes and ground-truth data to compute a detection loss.

7. The non-transitory computer-readable medium of claim 6 , wherein the instructions to decode include instructions to provide, using lateral connections between the backbone and the top-down pathway, the multiple scales of the feature map in addition to an output of a prior level from within the top-down pathway.

8. The non-transitory computer-readable medium of claim 6 , wherein the instructions to generate the feature map including instructions to generate the feature map at the multiple scales as a feature hierarchy, and

wherein the instructions to generate the feature map encode features of the input image to provide a common reference for generating the depth map and the bounding boxes.

9. A method, comprising:

generating, using a backbone of a combined network, a feature map at multiple scales from an input image;

decoding, using a top-down pathway of the combined network, the feature map to provide features at the multiple scales; and

generating, using a head of the combined network, a depth map from the features for a scene depicted in the input image and bounding boxes identifying objects in the input image, wherein the head includes multiple sub-heads for each of the multiple scales that perform 3D object detection, 2D object detection, depth estimation, and classification;

training, in a first stage, the combined network by using a supervised depth loss derived from the depth map; and

training, in a second stage, the combined network by using the bounding boxes and ground-truth data to compute a detection loss.

10. The method of claim 9 , wherein decoding includes providing, using lateral connections between the backbone and the top-down pathway, the multiple scales of the feature map in addition to an output of a prior level from within the top-down pathway.

11. The method of claim 9 , wherein generating the feature map includes generating the feature map at the multiple scales as a feature hierarchy, and

wherein generating the feature map encodes features of the input image to provide a common reference for generating the depth map and the bounding boxes.

12. The method of claim 9 , wherein generating the depth map and the bounding boxes includes using the head among separate layers of the top-down pathway at the multiple scales to generate the bounding boxes at the multiple scales and the depth map at one of the multiple scales.

13. The method of claim 9 , further comprising:

providing the depth map and the bounding boxes to cause navigation of a device according to the depth map and the bounding boxes.

Assignments (2)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Dec 6, 2024
From: TOYOTA RESEARCH INSTITUTE, INC.
To: TOYOTA JIDOSHA KABUSHIKI KAISHA
Reel/Frame 069506/0202 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jun 8, 2021
From: PARK, DENNIS; AMBRUS, RARES A.; GUIZILINI, VITOR; LI, JIE; GAIDON, ADRIEN DAVID
To: TOYOTA RESEARCH INSTITUTE, INC.
Reel/Frame 056517/0454 →
Continuity (2)
Provisional Application 63161735 · Mar 16, 2021
Related Publication 20220301202A1 · Sep 22, 2022
Cited By (2)
US 12,272,120 US 12,573,168