IP Library › Granted Patent US 12,008,818
Granted Patent B2
US 12,008,818 · App. 17/384,121 · Granted Jun 11, 2024

Systems and methods to train a prediction system for depth perception

Inventors: Rares A. Ambrus (San Francisco, CA); Dennis Park (Fremont, CA); Vitor Guizilini (Santa Clara, CA); Jie Li (Los Altos, CA); Adrien David Gaidon (Mountain View, CA)
Assignee: Toyota Research Institute, Inc.
G06V20/58G01S17/42G01S17/89G01S17/931G06F18/2113G06F18/2155G06F18/217G06F18/251G06N3/04G06N3/08G06N20/00G06T7/10G06T7/11G06T7/50G06V10/462G06V10/757G06V20/56G06T2207/10024G06T2207/10028G06T2207/20016G06T2207/20081G06T2207/20084G06T2207/30248
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,008,818
App. No.
17/384,121
Filed
Jul 23, 2021
Granted
Jun 11, 2024
Kind
B2
Art Unit
2671
USPC
382/100
Abstract

System, methods, and other embodiments described herein relate to a manner of training a depth prediction system using bounding boxes. In one embodiment, a method includes segmenting an image to mask areas beyond bounding boxes and identify unmasked areas within the bounding boxes. The method also includes training a depth model using depth losses from comparing weighted points associated with pixels of the image within the unmasked areas to ground-truth depth. The method also includes providing the depth model for object detection.

Claims (31)

1. A prediction system for training a depth model, comprising:

a processor; and

a memory storing instructions that, when executed by the processor, cause the processor to:

segment an image to mask areas beyond bounding boxes and identify unmasked areas within the bounding boxes;

train the depth model using depth losses from comparing weighted points associated with ranked pixels and a size of objects associated with a pixel count within the unmasked areas to ground-truth depth; and

provide the depth model for object detection.

2. The prediction system of claim 1 , wherein the instructions to train the depth model further include instructions to calculate the depth losses by adaptively decaying the ranked pixels relative to centers of the bounding boxes and selecting the weighted points according to the decaying.

3. The prediction system of claim 1 , wherein the instructions to train the depth model further include instructions to calculate the depth losses for the ranked pixels within the bounding boxes according to a statistical distribution for selecting the weighted points.

4. The prediction system of claim 1 , wherein the weighted points are over-weighted for the ranked pixels that overlap between the bounding boxes and the bounding boxes represent the objects identified within the image.

5. The prediction system of claim 1 , wherein the instructions to train the depth model further include instructions to calculate the depth losses by selecting the weighted points according to the size identified within the unmasked areas.

6. The prediction system of claim 1 , wherein the weighted points are proximate to focal points of the objects within the unmasked areas.

7. The prediction system of claim 1 , wherein the bounding boxes are locations within the image having the objects in a scene and the weighted points are associated with a point cloud generated from the image.

8. A non-transitory computer-readable medium for training a depth model including instructions that, when executed by a processor, cause the processor to:

segment an image to mask areas beyond bounding boxes and identify unmasked areas within the bounding boxes;

train the depth model using depth losses from comparing weighted points associated with ranked pixels and a size of objects associated with a pixel count within the unmasked areas to ground-truth depth; and

provide the depth model for object detection.

9. The non-transitory computer-readable medium of claim 8 , wherein the instructions to train the depth model further include instructions to calculate the depth losses by adaptively decaying the ranked pixels relative to centers of the bounding boxes and selecting the weighted points according to the decaying.

10. The non-transitory computer-readable medium of claim 8 , wherein the instructions to train the depth model further include instructions to calculate the depth losses the ranked pixels within the bounding boxes according to a statistical distribution for selecting the weighted points.

11. The non-transitory computer-readable medium of claim 8 , wherein the weighted points are over-weighted for the ranked pixels that overlap between the bounding boxes and the bounding boxes represent the objects identified within the image.

12. The non-transitory computer-readable medium of claim 8 , wherein the instructions to train the depth model further include instructions to calculate the depth losses by selecting the weighted points according to the sizes identified within the unmasked areas.

13. The non-transitory computer-readable medium of claim 8 , wherein the weighted points are proximate to focal points of the objects within the unmasked areas.

14. A method, comprising:

segmenting an image to mask areas beyond bounding boxes and identify unmasked areas within the bounding boxes;

training a depth model using depth losses from comparing weighted points associated with ranked pixels and a size of objects associated with a pixel count within the unmasked areas to ground-truth depth; and

providing the depth model for object detection.

15. The method of claim 14 , wherein training the depth model further includes calculating the depth losses by adaptively decaying the ranked pixels relative to centers of the bounding boxes and selecting the weighted points according to the decaying.

16. The method of claim 14 , wherein training the depth model further includes calculating the depth losses for the ranked pixels within the bounding boxes according to a statistical distribution for selecting the weighted points.

17. The method of claim 14 , wherein the weighted points are over-weighted for the ranked pixels that overlap between the bounding boxes and the bounding boxes represent the objects identified within the image.

18. The method of claim 14 , wherein training the depth model further includes calculating the depth losses by selecting the weighted points according to the size identified within the unmasked areas.

19. The method of claim 14 , wherein the weighted points are proximate to focal points of the objects within the unmasked areas.

20. The method of claim 14 , wherein the bounding boxes are locations within the image having the objects in a scene and the weighted points are associated with a point cloud generated from the image.

Assignments (2)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Nov 5, 2024
From: TOYOTA RESEARCH INSTITUTE, INC.
To: TOYOTA JIDOSHA KABUSHIKI KAISHA
Reel/Frame 069142/0710 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Aug 6, 2021
From: AMBRUS, RARES A.; PARK, DENNIS; GUIZILINI, VITOR; LI, JIE; GAIDON, ADRIEN DAVID
To: TOYOTA RESEARCH INSTITUTE, INC.
Reel/Frame 057112/0137 →
Continuity (2)
Provisional Application 63161735 · Mar 16, 2021
Related Publication 20220301203A1 · Sep 22, 2022