IP Library › Granted Patent US 11,308,639
Granted Patent B2
US 11,308,639 · App. 16/692,901 · Granted Apr 19, 2022

Tool and method for annotating a human pose in 3D point cloud data

Inventors: Saudin Botonjic (Gothenburg, SE); Sihao Ding (Mountain View, CA); Andreas Wallin (Billdal, SE)
Assignee: Volvo Car Corporation
G06T7/73G06T15/20G06T19/00G06T2200/24G06T2207/10028G06T2207/20081G06T2207/20084G06T2207/20101G06T2207/20132G06T2207/30196G06T2219/004
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,308,639
App. No.
16/692,901
Granted
Apr 19, 2022
Kind
B2
Abstract

A method and apparatus for annotating point cloud data. An apparatus may be configured to cause display of the point cloud data, label points in the point cloud data with a plurality of annotation points, the plurality of annotation points corresponding to points on a human body, move, in response to a user input, one or more of the annotation points to define a human pose and create annotated point cloud data, and output the annotated point cloud data.

Claims (68)

1. A method comprising:

causing, by one or more processors, display of at least one frame of multiple frames of point cloud data;

labeling, by the one or more processors, points in the at least one frame of the point cloud data with a plurality of annotation points, the plurality of annotation points corresponding to points on a human body;

causing, by the one or more processors, a video display of the multiple frames of the point cloud data, wherein the video display may be controlled to be forward, fast forward, pause, and reverse;

moving, by the one or more processors, and in response to a user input, one or more of the annotation points to define a human pose and create annotated point cloud data; and

outputting, by the one or more processors, the annotated point cloud data for training by a neural network.

2. The method of claim 1 , wherein labeling, by the one or more processors, points in the at least one frame of the point cloud data with the plurality of annotation points comprises:

estimating, by the one or more processors, a position of a potential human pose in the at least one frame of the point cloud data; and

labeling, by the one or more processors, the annotation points to correspond with the estimated position of the potential human pose.

3. The method of claim 1 , wherein moving, by the one or more processors, and in response to the user input, one or more of the annotation points to define the human pose and create annotated point cloud data comprises:

receiving, by the one or more processors, a selection of a single annotation point of the plurality of annotation points;

moving, by the one or more processors, and in response to the user input, only the single annotation point to define a portion of the human pose.

4. The method of claim 1 , wherein moving, by the one or more processors, and in response to the user input, one or more of the annotation points to define the human pose and create annotated point cloud data comprises:

receiving, by the one or more processors, a selection of two or more annotation points of the plurality of annotation points;

moving, by the one or more processors, and in response to the user input, only the two or more annotation points to define a portion of the human pose.

5. The method of claim 1 , further comprising:

displaying an image that corresponds to the at least one frame of the point cloud data at the same time as displaying the point cloud data;

cropping, by the one or more processors, the point cloud data and the image around a region of a potential human pose in the point cloud data; and

causing, by the one or more processors, display of the cropped region.

6. The method of claim 5 , further comprising:

causing, by the one or more processors, display of the cropped region at a plurality of perspectives.

7. The method of claim 1 , wherein the plurality of annotation points includes annotation points corresponding to a top of a head, a center of a neck, a right hip, a left hip, a right shoulder, a right elbow, a right hand, a right knee, a right foot, a left shoulder, a left elbow, a left hand, a left knee, and a left foot, and wherein groups of the annotation points correspond to limbs of a person, the method further comprising:

causing, by the one or more processors, display of lines between the annotation points to define the limbs, including displaying different limbs using different colors.

8. The method of claim 1 , further comprising:

adding, by the one or more processors, an action label for each of the multiple frames of the point cloud data and to the human pose.

9. The method of claim 1 , further comprising:

training, by the one or more processors, the neural network with the annotated point cloud data, wherein the neural network is configured to estimate a pose of a person from LiDAR point cloud data.

10. An apparatus comprising:

a memory configured to store point cloud data; and

one or more processors in communication with the memory, the one or more processors configured to:

cause display of at least one frame of multiple frames of the point cloud data;

label points in the at least one frame of the point cloud data with a plurality of annotation points, the plurality of annotation points corresponding to points on a human body;

cause a video display of the multiple frames of the point cloud data, wherein the video display may be controlled to be forward, fast forward, pause, and reverse;

move, in response to a user input, one or more of the annotation points to define a human pose and create annotated point cloud data; and

output the annotated point cloud data for training by a neural network.

11. The apparatus of claim 10 , wherein to label points in the at least one frame of the point cloud data with the plurality of annotation points, the one or more processors are further configured to:

estimate a position of a potential human pose in the at least one frame of the point cloud data; and

label the annotation points to correspond with the estimated position of the potential human pose.

12. The apparatus of claim 10 , wherein to move, in response to the user input, one or more of the annotation points to define the human pose and create annotated point cloud data, the one or more processors are further configured to:

receive a selection of a single annotation point of the plurality of annotation points;

move, in response to the user input, only the single annotation point to define a portion of the human pose.

13. The apparatus of claim 10 , wherein to move, in response to the user input, one or more of the annotation points to define the human pose and create annotated point cloud data, the one or more processors are further configured to:

receive a selection of two or more annotation points of the plurality of annotation points;

move, in response to the user input, only the two or more annotation points to define a portion of the human pose.

14. The apparatus of claim 10 , wherein the one or more processors are further configured to:

display an image that corresponds to the at least one frame of the point cloud data at the same time as displaying the point cloud data;

crop the point cloud data and the image around a region of a potential human pose in the point cloud data; and

cause display of the cropped region.

15. The apparatus of claim 14 , wherein the one or more processors are further configured to:

cause display of the cropped region at a plurality of perspectives.

16. The apparatus of claim 10 , wherein the plurality of annotation points include annotation points corresponding to a top of a head, a center of a neck, a right hip, a left hip, a right shoulder, a right elbow, a right hand, a right knee, a right foot, a left shoulder, a left elbow, a left hand, a left knee, and a left foot, and wherein groups of the annotation points correspond to limbs of a person, and wherein the one or more processors are further configured to:

cause display of lines between the annotation points to define the limbs, including cause display of different limbs using different colors.

17. The apparatus of claim 10 , wherein the one or more processors are further configured to:

add an action label for each of the multiple frames of the point cloud data and to the human pose.

18. The apparatus of claim 10 , wherein the one or more processors are further configured to:

train the neural network with the annotated point cloud data, wherein the neural network is configured to estimate a pose of a person from LiDAR point cloud data.

19. An apparatus comprising:

means for causing display of at least one frame of multiple frames of point cloud data;

means for labeling points in the at least one frame of the point cloud data with a plurality of annotation points, the plurality of annotation points corresponding to points on a human body;

means for causing a video display of the multiple frames of the point cloud data, wherein the video display may be controlled to be forward, fast forward, pause, and reverse;

means for moving, in response to a user input, one or more of the annotation points to define a human pose and create annotated point cloud data; and

means for outputting the annotated point cloud data for training by a neural network.

20. A non-transitory computer-readable storage medium storing instructions that, when executed, cause one or more processors to:

cause display of at least one frame of multiple frames of point cloud data;

label points in the at least one frame of the point cloud data with a plurality of annotation points, the plurality of annotation points corresponding to points on a human body;

cause a video display of the multiple frames of the point cloud data, wherein the video display may be controlled to be forward, fast forward, pause, and reverse;

move, in response to a user input, one or more of the annotation points to define a human pose and create annotated point cloud data; and

output the annotated point cloud data for training by a neural network.

Assignments (2)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Apr 16, 2020
From: BOTONJIC, SAUDIN; DING, SIHAO; WALLIN, ANDREAS
To: VOLVO CAR CORPORATION
Reel/Frame 052422/0253 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Nov 22, 2019
From: BOTONJIC, SAUDIN; DING, SIHAO
To: VOLVO CAR CORPORATION
Reel/Frame 051092/0010 →
Continuity (2)
Provisional Application 62817400 · Mar 12, 2019
Related Publication 20200294266A1 · Sep 17, 2020