Trajectory prediction on top-down scenes and associated model
Techniques are discussed for determining prediction probabilities of an object based on a top-down representation of an environment. Data representing objects in an environment can be captured. Aspects of the environment can be represented as map data. A multi-channel image representing a top-down view of object(s) in the environment can be generated based on the data representing the objects and map data. The multi-channel image can be used to train a machine learned model by minimizing an error between predictions from the machine learned model and a captured trajectory associated with the object. Once trained, the machine learned model can be used to generate prediction probabilities of objects in an environment, and the vehicle can be controlled based on such prediction probabilities.
1 . A system comprising:
one or more processors; and
one or more non-transitory computer-readable media storing instructions executable by the one or more processors, wherein the instructions, when executed, cause the system to perform operations comprising:
receiving map data associated with an environment;
receiving sensor data from a sensor associated with a vehicle in the environment;
determining, based at least in part on the sensor data, object data associated with an object in the environment, the object data comprising at least one of a semantic label of the object, a class associated with the object, a bounding box representing the object, a velocity of the object, or an acceleration of the object;
determining, based at least in part on the object data, the map data, and the sensor data, a multi-channel data structure, wherein channels of the multi-channel data structure encode different information, the information encoded by an individual channel comprising at least one of the map data, the object data, or the sensor data;
inputting the multi-channel data structure into a machine learned model;
receiving, from the machine learned model and based at least in part on the object data and the map data, a prediction probability associated with movement of the object in the environment; and
controlling, based at least in part on the prediction probability, the vehicle to traverse the environment.
2 . The system of claim 1 , wherein the map data includes semantic information associated with the environment, the semantic information comprising at least one of road network information or a traffic light status.
3 . The system of claim 1 , wherein the prediction probability comprises at least one of:
a multi modal Gaussian trajectory; or
an occupancy grid associated with a future time, wherein a cell of the occupancy grid is indicative of a probability of the object being in a region associated with the cell at the future time.
4 . The system of claim 1 , wherein the machine learned model comprises an encoder and a decoder.
5 . The system of claim 4 , wherein the decoder comprises one or more of:
a recurrent neural network;
a network configured to regress a plurality of prediction probabilities substantially simultaneously; or
a network comprising a two dimensional convolutional-transpose network.
6 . The system of claim 1 , the operations further comprising determining, based on the prediction probability and a vehicle dynamics model associated with the object, a predicted trajectory associated with the object.
7 . The system of claim 6 , wherein the vehicle dynamics model includes at least a velocity cost, a position cost, an acceleration cost, and rules of the road.
8 . One or more non-transitory computer-readable media storing instructions executable by one or more processors, wherein the instructions, when executed, cause the one or more processors to perform operations comprising:
receiving map data associated with an environment;
receiving sensor data from a sensor associated with a vehicle in the environment;
determining, based at least in part on the sensor data, object data associated with an object in the environment, the object data comprising at least one of a semantic label of the object, a class associated with the object, a bounding box representing the object, a velocity of the object, or an acceleration of the object;
determining, based at least in part on the object data, the map data, and the sensor data, a multi-channel data structure, wherein channels of the multi-channel data structure encode different information, the information encoded by an individual channel comprising at least one of the map data, the object data, or the sensor data;
determining, based at least in part on the object data and the map data multi-channel data structure, a prediction probability associated with movement of the object in the environment; and
controlling, based at least in part on the prediction probability, the vehicle to traverse the environment.
9 . The one or more non-transitory computer-readable media of claim 8 , wherein the map data includes semantic information associated with the environment, the semantic information comprising at least one of road network information or a traffic light status.
10 . The one or more non-transitory computer-readable media of claim 8 , wherein the prediction probability comprises at least one of:
a multi modal Gaussian trajectory; or
an occupancy grid associated with a future time, wherein a cell of the occupancy grid is indicative of a probability of the object being in a region associated with the cell at the future time.
11 . The one or more non-transitory computer-readable media of claim 8 , wherein determining the prediction probability comprises inputting the multi-channel data structure to a machine learned model, and wherein the machine learned model comprises an encoder and a decoder.
12 . The one or more non-transitory computer-readable media of claim 11 , wherein the decoder comprises one or more of:
a recurrent neural network;
a network configured to regress a plurality of prediction probabilities substantially simultaneously; or
a network comprising a two dimensional convolutional-transpose network.
13 . The one or more non-transitory computer-readable media of claim 8 , the operations further comprising determining, based on the prediction probability and a vehicle dynamics model associated with the object, a predicted trajectory associated with the object.
14 . The one or more non-transitory computer-readable media of claim 8 , wherein the object data comprises the bounding box representing the object, and wherein the bounding box representing the object is a three-dimensional bounding box.
15 . A method comprising:
receiving map data associated with an environment;
receiving sensor data from a sensor associated with a vehicle in the environment;
determining, based at least in part on the sensor data, object data associated with an object in the environment, the object data comprising at least one of a semantic label of the object, a class associated with the object, a bounding box representing the object, a velocity of the object, or an acceleration of the object;
determining, based at least in part on the object data, the map data, and the sensor data, a multi-channel data structure, wherein channels of the multi-channel data structure encode different information, the information encoded by an individual channel comprising at least one of the map data, the object data, or the sensor data;
determining, based at least in part on the multi-channel data structure, a prediction probability associated with movement of the object in the environment; and
controlling, based at least in part on the prediction probability, the vehicle to traverse the environment.
16 . The method of claim 15 , wherein the map data includes semantic information associated with the environment, the semantic information comprising at least one of road network information or a traffic light status.
17 . The method of claim 15 , wherein the prediction probability comprises at least one of:
a multi modal Gaussian trajectory; or
an occupancy grid associated with a future time, wherein a cell of the occupancy grid is indicative of a probability of the object being in a region associated with the cell at the future time.
18 . The method of claim 15 , wherein determining the prediction probability comprises inputting the multi-channel data structure to a machine learned model, and wherein the machine learned model comprises an encoder and a decoder.
19 . The method of claim 18 , wherein the decoder comprises one or more of:
a recurrent neural network;
a network configured to regress a plurality of prediction probabilities substantially simultaneously; or
a network comprising a two dimensional convolutional-transpose network.
20 . The method of claim 15 , further comprising determining, based on the prediction probability and a vehicle dynamics model associated with the object, a predicted trajectory associated with the object.