Systems and methods for point cloud annotation using kinematic models
A system including at least one memory and at least one processor coupled to the at least one memory is disclosed. The at least one processor is configured to: receive, from a sensor on an autonomous vehicle, sensor data; receive, from a camera on the autonomous vehicle, camera data; provide a frame of the camera data overlaid with a respective frame of the sensor data to a user; receive a user annotation indicative of an actor in the frame, where the user annotation encloses a first set of points associated with the actor; generate an a virtual annotation enclosing a second set of points associated with the actor; determine that a point in the second set of points is older than a reference time; project the point to a new position based on a kinematic model; generate a temporally aggregated actor to determine a quality of the user input.
1 . A system comprising:
at least one memory configured to store machine executable instructions; and
at least one processor coupled to the at least one memory and configured to execute the machine executable instructions to perform operations comprising:
receiving, from a sensor mounted on an autonomous vehicle, sensor data;
receiving, from a camera mounted on the autonomous vehicle, camera data;
providing a respective frame of a set of frames of the camera data overlaid with a corresponding frame of the sensor data, via an interface of a client device, to a user;
receiving, via the interface, a user annotation indicating an actor in the respective frame, wherein the user annotation is drawn by the user based on the camera data and encloses a first set of points of the sensor data associated with the actor;
generating a virtual annotation enclosing a second set of points associated with the actor, wherein the second set of points comprises the first set of points;
determining that a point in the second set of points is associated with a time that is older than a reference time;
generating an actor compensated set of points by projecting the point to a new position based on a kinematic model associated with the actor;
generating a temporally aggregated actor by overlaying the actor compensated sets of points of the respective frames of the set of frames to determine a quality associated with user annotations.
2 . The system of claim 1 , wherein the operations further comprise:
generating the kinematic model for the actor by comparing a position of a first point of the first set of points in the respective frame with the position of the first point in a subsequent frame.
3 . The system of claim 1 , wherein the kinematic model comprises a velocity and an acceleration associated with the actor.
4 . The system of claim 1 , wherein the user input is annotated with an identifier that is associated with the actor across frames of the set of frames.
5 . The system of claim 1 , wherein the operations further comprise:
based on the quality being less than a threshold quality, refining the first set of points by:
determining an annotation error of a point of the first set of points; and
compensating the point based on a difference in position of the point in a reference frame and a subsequent frame.
6 . The system of claim 5 , wherein the operations further comprise:
updating, based on the compensation, the kinematic model; and
re-generating the temporally aggregated actor using the updated kinematic model.
7 . The system of claim 6 , wherein the operations further comprise:
iteratively optimizing the temporally aggregated actor until the quality is greater than or equal to the threshold quality.
8 . The system of claim 1 , wherein the first sensor data is generated by a raster LiDAR mounted on the autonomous vehicle.
9 . The system of claim 1 , wherein the user annotation is a three-dimensional cuboid.
10 . A computer-implemented method comprising:
receiving, from a sensor mounted on an autonomous vehicle, sensor data;
receiving, from a camera mounted on the autonomous vehicle, camera data;
providing a respective frame of a set of frames of the camera data overlaid with a corresponding frame of the sensor data, via an interface of a client device, to a user;
receiving, via the interface, a user annotation indicating an actor in the respective frame, wherein the user annotation is drawn by the user based on the camera data and encloses a first set of points of the sensor data associated with the actor;
generating a virtual annotation enclosing a second set of points associated with the actor, wherein the second set of points comprises the first set of points;
determining that a point in the second set of points is associated with a time that is older than a reference time;
generating an actor compensated set of points by projecting the point to a new position based on a kinematic model associated with the actor;
generating a temporally aggregated actor by overlaying the actor compensated sets of points of the respective frames of the set of frames to determine a quality associated with user annotations.
11 . The method of claim 10 , further comprising:
generating the kinematic model for the actor by comparing a position of a first point of the first set of points in the respective frame with the position of the first point in a subsequent frame.
12 . The method of claim 10 , wherein the kinematic model comprises a velocity and an acceleration associated with the actor.
13 . The method of claim 10 , wherein the user input is annotated with an identifier that is associated with the actor across frames of the set of frames.
14 . The method of claim 10 , further comprising:
based on the quality being less than a threshold quality, refining the first set of points by:
determining an annotation error of a point of the first set of points; and
compensating the point based on a difference in position of the point in a reference frame and a subsequent frame.
15 . The method of claim 14 , further comprising:
updating, based on the compensation, the kinematic model; and
re-generating the temporally aggregated actor using the updated kinematic model.
16 . The method of claim 15 , further comprising:
iteratively optimizing the temporally aggregated actor until the quality is greater than or equal to the threshold quality.
17 . The method of claim 10 , wherein the first sensor data is generated by a raster LiDAR mounted on the autonomous vehicle.
18 . The method of claim 10 , wherein the user annotation is a three-dimensional cuboid.
19 . A system comprising:
an autonomous vehicle hosting a sensor and a camera; and
a computing system comprising:
at least one interface;
at least one memory configured to store machine executable instructions; and
at least one processor coupled to the at least one memory and configured to execute the instructions to perform operations comprising:
receiving, from the autonomous vehicle, sensor data and camera data;
providing a respective frame of a set of frames of the camera data overlaid with a corresponding frame of the sensor data, via an interface of a client device, to a user;
receiving, via the interface, a user annotation indicating an actor in the respective frame, wherein the user annotation is drawn by the user based on the camera data and encloses a first set of points of the sensor data associated with the actor;
generating a virtual annotation enclosing a second set of points associated with the actor, wherein the second set of points comprises the first set of points;
determining that a point in the second set of points is associated with a time that is older than a reference time;
generating an actor compensated set of points by projecting the point to a new position based on a kinematic model associated with the actor;
generating a temporally aggregated actor by overlaying the actor compensated sets of points of the respective frames of the set of frames to determine a quality associated with user annotations.
20 . The system of claim 19 , wherein the operations further comprise:
generating the kinematic model for the actor by comparing a position of a first point of the first set of points in the respective frame with the position of the first point in a subsequent frame.