IP Library Granted Patent US 10,670,416
Granted Patent B2
US 10,670,416 · App. 15/859,182 · Granted Jun 2, 2020

Traffic sign feature creation for high definition maps used for navigating autonomous vehicles

Inventors: Mark Damon Wheeler (Saratoga, CA); Lin Yang (San Carlos, CA); Derek Thomas Miller (Palo Alto, CA); Yu Zhang (Mountain View, CA); Lenord Melvix Joseph Stephen Max (Mountain View, CA)
Assignee: DEEPMAP INC.
G01C21/3638B60W40/04G01C21/32G01C21/3635G05D1/0088G06K9/00798G06K9/00818G06K9/44G06K9/4638G06T17/00G06T17/05B60W2420/42B60W2420/52
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 10,670,416
App. No.
15/859,182
Granted
Jun 2, 2020
Kind
B2
Abstract

An HD map system represents landmarks on a high definition map for autonomous vehicle navigation, including describing spatial location of lanes of a road and semantic information about each lane, and along with traffic signs and landmarks. The system generates lane lines designating lanes of roads based on, for example, mapping of camera image pixels with high probability of being on lane lines into a three-dimensional space, and locating/connecting center lines of the lane lines. The system builds a large connected network of lane elements and their connections as a lane element graph. The system also represents traffic signs based on camera images and detection and ranging sensor depth maps. These landmarks are used in building a high definition map that allows autonomous vehicles to safely navigate through their environments.

Claims (141)

1. A method of representing a traffic sign in a three-dimensional map comprising:

receiving an image captured by a camera mounted on an autonomous vehicle, the image including the traffic sign;

identifying a portion of the image corresponding to the traffic sign;

receiving a depth map including the traffic sign captured by a detection and ranging sensor, the depth map comprising a plurality of points with each point describing distance;

constructing the three-dimensional map by mapping the plurality of points describing distances from the depth map into the three-dimensional map;

identifying a subset of at least three points in the depth map that correspond to the traffic sign;

fitting a plane in the three-dimensional map based at least in part on the subset of at least three points corresponding to the traffic sign;

projecting the portion of the image onto the fitted plane in the three-dimensional map

determining a type of the traffic sign based on a neural network analysis of the portion of the image;

identifying one or more legal requirements associated with the type of the traffic sign; and

storing the type of the traffic sign and the one or more legal requirements as attributes of the traffic sign, wherein the attributes of the traffic sign aid in navigation of autonomous vehicles.

2. The method of claim 1 , wherein identifying a portion of the image corresponding to the traffic sign comprises:

identifying a plurality of vertices of the traffic sign on the image; and

determining a polygon with the plurality of vertices as the portion of the image corresponding to the traffic sign.

3. The method of claim 2 , wherein projecting the portion of the image onto the plane in the three-dimensional map comprises for each vertex:

determining a ray from an origin of the camera through the vertex;

determining an intersection of the ray with the plane in the three-dimensional map; and

mapping the vertex to the intersection in the three-dimensional map.

4. The method of claim 1 , wherein the detection and ranging sensor is a light detection and ranging sensor (LIDAR) mounted on the vehicle.

5. The method of claim 4 , wherein the depth map captured by the LIDAR comprises merging a plurality of scans taken by the LIDAR so as to increase a total number of points in the plurality of points describing distances in the depth map.

6. The method of claim 1 , wherein constructing the three-dimensional map further comprises:

receiving a second depth map from a second detection and ranging sensor mounted on a second vehicle, the second depth map comprising a second plurality of points with each point describing distance; and

mapping the second plurality of points from the second depth map into the three-dimensional map.

7. The method of claim 1 , wherein identifying a subset of at least three points in the depth map that correspond to the traffic sign further comprises:

determining a bounding box on the depth map based in part on the portion of the image corresponding to the traffic sign;

determining a minimum depth and a maximum depth of the depth map based at least in part on a size of the portion of the image corresponding to the traffic sign, wherein the traffic sign is within the minimum depth and the maximum depth;

determining a frustum produced by the bounding box with the minimum depth and the maximum depth; and

identifying the subset of at least three points from points in the depth map which reside within the frustum.

8. The method of claim 7 , wherein identifying a subset of at least three points in the depth map that correspond to the traffic sign further comprises:

determining a first point of the plurality of points in the depth map which resides within the frustum at a minimum depth; and

selecting two or more points of the plurality of points in the depth map within a threshold depth of the first point of the plurality of points in the depth map which reside within the frustum.

9. The method of claim 1 further comprising:

determining text on the traffic sign based on a neural network analysis of characters on the portion of the image corresponding to the traffic sign; and

storing the text on the traffic sign as an attribute of the traffic sign in the three-dimensional map.

10. The method of claim 1 , wherein fitting a plane in the three-dimensional map inclusive of the subset of at least three points corresponding to the traffic sign comprises utilizing random sample consensus (RANSAC) to determine the plane which has a high probability of fitting the subset of at least three points.

11. A method of representing a traffic sign in a three-dimensional map comprising:

receiving an image captured by a camera mounted on a vehicle, the image including the traffic sign;

identifying a portion of the image corresponding to the traffic sign;

receiving a depth map including the traffic sign captured by a detection and ranging sensor, the depth map comprising a plurality of points with each point describing distance;

constructing the three-dimensional map by mapping the plurality of points describing distances from the depth map into the three-dimensional map;

identifying a subset of at least three points in the depth map that correspond to the traffic sign;

fitting a plane in the three-dimensional map based at least in part on the subset of at least three points corresponding to the traffic sign;

projecting the portion of the image onto the fitted plane in the three-dimensional map;

determining a type of the traffic sign based on a neural network analysis of the portion of the image;

identifying one or more legal requirements associated with the type of the traffic sign;

storing the type of the traffic sign and the one or more legal requirements as attributes of the traffic sign, wherein the attributes of the traffic sign aid in navigation of one or more vehicles; and

providing for display the three-dimensional map including the projected portion of the image corresponding to the traffic sign and one or more attributes of the traffic sign through one or more graphical user interfaces on the one or more vehicles.

12. The method of claim 11 , wherein identifying a portion of the image corresponding to the traffic sign comprises:

identifying a plurality of vertices of the traffic sign on the image; and

determining a polygon with the plurality of vertices as the portion of the image corresponding to the traffic sign.

13. The method of claim 11 , wherein projecting the portion of the image onto the plane in the three-dimensional map comprises for each vertex:

determining a ray from an origin of the camera through the vertex;

determining an intersection of the ray with the plane in the three-dimensional map; and

mapping the vertex to the intersection in the three-dimensional map.

14. The method of claim 11 , wherein constructing the three-dimensional map further comprises:

receiving a second depth map from a second detection and ranging sensor mounted on a second vehicle, the second depth map comprising a second plurality of points with each point describing distance; and

mapping the second plurality of points from the second depth map into the three-dimensional map.

15. The method of claim 11 , wherein identifying a subset of at least three points in the depth map that correspond to the traffic sign further comprises:

determining a bounding box on the depth map based in part on the portion of the image corresponding to the traffic sign;

determining a minimum depth and a maximum depth of the depth map based at least in part on a size of the portion of the image corresponding to the traffic sign, wherein the traffic sign is within the minimum depth and the maximum depth;

determining a frustum produced by the bounding box with the minimum depth and the maximum depth; and

identifying the subset of at least three points from points in the depth map which reside within the frustum.

16. The method of claim 15 , wherein identifying a subset of at least three points in the depth map that correspond to the traffic sign further comprises:

determining a first point of the plurality of points in the depth map which resides within the frustum at a minimum depth; and

selecting two or more points of the plurality of points in the depth map within a threshold depth of the first point of the plurality of points in the depth map which reside within the frustum.

17. The method of claim 11 further comprising:

determining text on the traffic sign based on a neural network analysis of characters on the portion of the image corresponding to the traffic sign; and

storing the text on the traffic sign as an attribute of the traffic sign in the three-dimensional map.

18. The method of claim 11 , wherein fitting a plane in the three-dimensional map inclusive of the subset of at least three points corresponding to the traffic sign comprises utilizing random sample consensus (RANSAC) to determine the plane which has a high probability of fitting the subset of at least three points.

19. A non-transitory computer-readable storage medium for representing a traffic sign in a three-dimensional map, the non-transitory computer-readable storage medium storing instructions that, when executed by a processor, cause a system to perform operations comprising:

receiving an image captured by a camera mounted on an autonomous vehicle, the image including the traffic sign;

identifying a portion of the image corresponding to the traffic sign;

receiving a depth map including the traffic sign captured by a detection and ranging sensor, the depth map comprising a plurality of points with each point describing distance;

constructing the three-dimensional map by mapping the plurality of points describing distances from the depth map into the three-dimensional map;

identifying a subset of at least three points in the depth map that correspond to the traffic sign;

fitting a plane in the three-dimensional map based at least in part on the subset of at least three points corresponding to the traffic sign;

projecting the portion of the image onto the fitted plane in the three-dimensional map;

determining a type of the traffic sign based on a neural network analysis of the portion of the image;

identifying one or more legal requirements associated with the type of the traffic sign; and

storing the type of the traffic sign and the one or more legal requirements as attributes of the traffic sign, wherein the attributes of the traffic sign aid in navigation of autonomous vehicles.

20. The storage medium of claim 19 , wherein identifying a portion of the image corresponding to the traffic sign comprises:

identifying a plurality of vertices of the traffic sign on the image; and

determining a polygon with the plurality of vertices as the portion of the image corresponding to the traffic sign.

21. The storage medium of claim 20 , wherein projecting the portion of the image onto the plane in the three-dimensional map comprises, for each vertex:

determining a ray from an origin of the camera through the vertex;

determining an intersection of the ray with the plane in the three-dimensional map; and

mapping the vertex to the intersection in the three-dimensional map.

22. The storage medium of claim 19 , wherein the detection and ranging sensor is a light detection and ranging sensor (LIDAR) mounted on the vehicle.

23. The storage medium of claim 22 , wherein the depth map captured by the LIDAR comprises merging a plurality of scans taken by the LIDAR so as to increase total number of points in the plurality of points describing distances in the depth map.

24. The storage medium of claim 19 , wherein constructing the three-dimensional map further comprises:

receiving a second depth map from a second detection and ranging sensor mounted on a second vehicle, the second depth map comprising a second plurality of points with each point describing distance; and

mapping the second plurality of points from the second depth map into the three-dimensional map.

25. The storage medium of claim 19 , where identifying a subset of at least three points in the depth map that correspond to the traffic sign further comprises:

determining a bounding box on the depth map based in part on the portion of the image corresponding to the traffic sign;

determining a minimum depth and a maximum depth of the depth map based at least in part on a size of the portion of the image corresponding to the traffic sign, wherein the traffic sign is within the minimum depth and the maximum depth;

determining a frustum produced by the bounding box with the minimum depth and the maximum depth; and

identifying the subset of at least three points from points in the depth map which reside within the frustum.

26. The storage medium of claim 25 , where identifying a subset of at least three points in the depth map that correspond to the traffic sign further comprises:

determining a first point of the plurality of points in the depth map which resides within the frustum at a minimum depth; and

selecting two or more points of the plurality of points in the depth map within a threshold depth of the first point of the plurality of points in the depth map which reside within the frustum.

27. The storage medium of claim 19 , the operations further comprising:

determining text on the traffic sign based on a neural network analysis of characters on the portion of the image corresponding to the traffic sign; and

storing the text on the traffic sign as an attribute of the traffic sign in the three-dimensional map.

28. The storage medium of claim 19 , wherein fitting a plane in the three-dimensional map inclusive of the subset of at least three points corresponding to the traffic sign comprises utilizing random sample consensus (RANSAC) to determine the plane which has a high probability of fitting the subset of at least three points.

29. A system comprising:

a processor; and

a non-transitory computer-readable storage medium for representing a traffic sign in a three-dimensional map, the non-transitory computer-readable storage medium storing instructions that, when executed by the processor, cause the system to perform operations comprising:

receiving an image captured by a camera mounted on an autonomous vehicle, the image including the traffic sign;

identifying a portion of the image corresponding to the traffic sign;

receiving a depth map including the traffic sign captured by a detection and ranging sensor, the depth map comprising a plurality of points with each point describing distance;

constructing the three-dimensional map by mapping the plurality of points describing distances from the depth map into the three-dimensional map;

identifying a subset of at least three points in the depth map that correspond to the traffic sign;

fitting a plane in the three-dimensional map based at least in part on the subset of at least three points corresponding to the traffic sign;

projecting the portion of the image onto the fitted plane in the three-dimensional map;

determining a type of the traffic sign based on a neural network analysis of the portion of the image;

identifying one or more legal requirements associated with the type of the traffic sign; and

storing the type of the traffic sign and the one or more legal requirements as attributes of the traffic sign, wherein the attributes of the traffic sign aid in navigation of autonomous vehicles.

30. The system of claim 29 , wherein identifying a portion of the image corresponding to the traffic sign comprises:

identifying a plurality of vertices of the traffic sign on the image; and

determining a polygon with the plurality of vertices as the portion of the image corresponding to the traffic sign.

31. The system of claim 30 , wherein projecting the portion of the image onto the plane in the three-dimensional map comprises, for each vertex:

determining a ray from an origin of the camera through the vertex;

determining an intersection of the ray with the plane in the three-dimensional map; and

mapping the vertex to the intersection in the three-dimensional map.

32. The system of claim 29 , wherein the detection and ranging sensor is a light detection and ranging sensor (LIDAR) mounted on the vehicle.

33. The system of claim 32 , wherein the depth map captured by the LIDAR comprises merging a plurality of scans taken by the LIDAR so as to increase total number of points in the plurality of points describing distances in the depth map.

34. The system of claim 29 , wherein constructing the three-dimensional map further comprises:

receiving a second depth map from a second detection and ranging sensor mounted on a second vehicle, the second depth map comprising a second plurality of points with each point describing distance; and

mapping the second plurality of points from the second depth map into the three-dimensional map.

35. The system of claim 29 , where identifying a subset of at least three points in the depth map that correspond to the traffic sign further comprises:

determining a bounding box on the depth map based in part on the portion of the image corresponding to the traffic sign;

determining a minimum depth and a maximum depth of the depth map based at least in part on a size of the portion of the image corresponding to the traffic sign, wherein the traffic sign is within the minimum depth and the maximum depth;

determining a frustum produced by the bounding box with the minimum depth and the maximum depth; and

identifying the subset of at least three points from points in the depth map which reside within the frustum.

36. The system of claim 35 , where identifying a subset of at least three points in the depth map that correspond to the traffic sign further comprises:

determining a first point of the plurality of points in the depth map which resides within the frustum at a minimum depth; and

selecting two or more points of the plurality of points in the depth map within a threshold depth of the first point of the plurality of points in the depth map which reside within the frustum.

37. The system of claim 29 , the operations further comprising:

determining text on the traffic sign based on a neural network analysis of characters on the portion of the image corresponding to the traffic sign; and

storing the text on the traffic sign as an attribute of the traffic sign in the three-dimensional map.

38. The system of claim 29 , wherein fitting a plane in the three-dimensional map inclusive of the subset of at least three points corresponding to the traffic sign comprises utilizing random sample consensus (RANSAC) to determine the plane which has a high probability of fitting the subset of at least three points.

Assignments (4)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jul 20, 2022
From: DEEPMAP INC.
To: NVIDIA CORPORATION
Reel/Frame 061038/0311 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Sep 5, 2019
From: DEEPMAP CAYMAN LIMITED
To: DEEPMAP INC.
Reel/Frame 050281/0787 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jun 26, 2018
From: DEEPMAP INC.
To: DEEPMAP CAYMAN LIMITED
Reel/Frame 046208/0934 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Feb 4, 2018
From: WHEELER, MARK DAMON; YANG, LIN; MILLER, DEREK THOMAS; ZHANG, YU; JOSEPH STEPHEN MAX, LENORD MELVIX
To: DEEPMAP INC.
Reel/Frame 044825/0771 →
Continuity (3)
Provisional Application 62441065 · Dec 30, 2016
Provisional Application 62441080 · Dec 30, 2016
Related Publication 20180188060A1 · Jul 5, 2018
Cited By (3)
US 12,384,410 US 12,625,926 US 12,700,308