IP Library Granted Patent US 12,153,439
Granted Patent B2
US 12,153,439 · App. 16/936,415 · Granted Nov 26, 2024

Monocular 3D object detection from image semantics network

Inventors: Oscar Olof Beijbom (Santa Monica, CA); Varun Kumar Reddy Bankiti (Los Angeles, CA); Donghyeon Won (Los Angeles, CA)
Assignee: Motional AD LLC
G05D1/0251G05D1/0088G05D1/0219G05D1/0221G06F18/214G06T7/80G06V20/588
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,153,439
App. No.
16/936,415
Granted
Nov 26, 2024
Kind
B2
Abstract

Techniques are provided for monocular 3D object detection from an image semantics network. An image semantics network (ISN) is a single stage, single image object detection network that is based on single shot detection (SSD). In an embodiment, the ISN augments the SSD outputs to provide encoded 3D properties of the object along with a 2D bounding box and classification scores. For each priorbox, a 3D bounding box is generated for the object using the dimensions and location of the priorbox, the encoded 3D properties and camera intrinsic parameters.

Claims (41)

1. A method comprising:

receiving, using one or more processors of a vehicle, images from a camera of the vehicle;

generating, using an object detection network with the images as input, two-dimensional (2D) positions, dimensions and center offsets of 2priorboxes and corresponding classification scores for each object detected in the images, wherein the object detection network is a single stage, single image network with a single shot detector detection head;

for each detected object:

generating encoded three-dimensional (3D) properties of the detected object, the encoded 3D properties including a center projection of a 3D bounding box for the detected object, the center projection determined by the position, dimensions and center offsets of a 2D bounding box corresponding to the detected object;

generating a 3D bounding box for the detected object using the encoded 3D properties and camera intrinsic parameters;

computing a route or trajectory for the vehicle using at least in part the generated 3D bounding boxes; and

causing, using a controller of the vehicle, the vehicle to travel along the route or trajectory.

2. The method of claim 1 , wherein the encoded 3D properties include dimensions, radial distance, viewing angle and center projection offsets.

3. The method of claim 1 , wherein generating, using the one or more processors, encoded three-dimensional (3D) properties of the object, further comprises: for each priorbox, estimating a vector of parameters that include a set of offsets from a bottom center of the priorbox, width, length and height of the object, radial distance from a center of the camera to the object center and viewing angle.

4. The method of claim 1 , wherein the object detection network outputs six groups of parameters including classification, 2D localization, dimensional, orientation, radial distance and center projection, and the method further comprises:

computing a loss function for each group of parameters individually; and

using the loss functions to train the object detection network.

5. A system comprising:

one or more processors;

memory storing instructions that when executed by the one or more processors, cause the one or more processors to perform operations comprising:

receiving images from a camera of a vehicle;

generating, using an object detection network with the images as input, two-dimensional (2D) positions, dimensions and center offsets of the 2D priorboxes and corresponding classification scores for each object detected in the images, wherein the object detection network is a single stage, single image network with a single shot detector detection head;

for each detected object:

generating encoded three-dimensional (3D) properties of the detected object, the encoded 3D properties including a center projection of a 3D bounding box for the detected object, the center projection determined by the position, dimensions and center offsets of a corresponding 2D bounding box;

generating a 3D bounding box for the detected object using the encoded 3D properties and camera intrinsic parameters;

computing a route or trajectory for the vehicle using at least in part the generated 3D bounding boxes; and

causing, using a controller of the vehicle, the vehicle to travel along the route or trajectory.

6. The system of claim 5 , wherein the encoded 3D properties include dimensions, radial distance, viewing angle and center projection offsets for a 3D bounding box.

7. The system of claim 5 , wherein generating encoded three-dimensional (3D) properties of the object, further comprises: for each priorbox, estimating a vector of parameters that include a set of offsets from a bottom center of the priorbox, width, length and height of the object, radial distance from a center of the camera to the object center and viewing angle.

8. The system of claim 5 , wherein the object detection network outputs six groups of parameters including classification, 2D localization, dimensional, orientation, radial distance and center projection of a 3D bounding box, and the method further comprise:

computing a loss function for each group of parameters individually; and

using the loss functions to train the object detection network.

9. A non-transitory, computer-readable storage medium having instructions stored thereon, that when executed by at least one processor, cause the at least one processor to perform operations comprising:

receiving images from a camera of a vehicle;

generating, using an object detection network with the images as input, two-dimensional (2D) positions, dimensions and center offsets of 2D priorboxes and corresponding classification scores for each object detected in the images, wherein the object detection network is a single stage, single image network with a single shot detector detection head;

for each object:

generating encoded three-dimensional (3D) properties of the object, the encoded 3D properties including a center projection of a 3D bounding box for the detected object, the center projection determined by the position, dimensions and center offsets of a corresponding 2D bounding box;

generating a 3D bounding box for the detected object using the encoded 3D properties and camera intrinsic parameters;

computing a route or trajectory for the vehicle using at least in part the generated 3D bounding boxes; and

causing, using a controller of the vehicle, the vehicle to travel along the route or trajectory.

10. The non-transitory, computer-readable storage medium of claim 9 , wherein the encoded 3D properties include dimensions, radial distance, viewing angle and center projection offsets of a 3D bounding box.

11. The non-transitory, computer-readable storage medium of claim 9 , wherein generating, using the one or more processors, encoded three-dimensional (3D) properties of the object, further comprises: for each priorbox, estimating a vector of parameters that include a set of offsets from a bottom center of the priorbox, width, length and height of the object, radial distance from a center of the camera to the object center and viewing angle.

12. The non-transitory, computer-readable storage medium of claim 9 , wherein the object detection network outputs six groups of parameters including classification, 2D localization, dimensional, orientation, radial distance and center projection, and the operations further comprises:

computing a loss function for each group of parameters individually; and

using the loss functions to train the object detection network.

Assignments (2)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Sep 23, 2020
From: APTIV TECHNOLOGIES LIMITED
To: MOTIONAL AD LLC
Reel/Frame 053861/0847 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jul 23, 2020
From: BEIJBOM, OSCAR OLOF; BANKITI, VARUN KUMAR REDDY; WON, DONGHYEON
To: APTIV TECHNOLOGIES LIMITED
Reel/Frame 053289/0396 →
Continuity (1)
Related Publication 20220026917A1 · Jan 27, 2022