Predictive models for autonomous vehicles based on object interactions
Techniques are discussed herein for training and executing machine learning (ML) prediction models used to control autonomous vehicles in driving environments. In various examples, ML prediction models configured to output joint trajectory predictions for multiple objects in an environment may be trained by evaluating the interactions between the objects represented by the predicted trajectories. A training component may train an ML prediction model using a standard loss function based on the accuracy of the predicted trajectories relative to the ground truth trajectories, and based on an auxiliary loss determined by the agent-to-agent interactions represented by the predicted trajectories. The auxiliary loss may be determined by various techniques, including using a classification model trained to receive and classify sets of object trajectories in a generative adversarial network (GAN), and/or determining a divergence loss based on an alternative ML prediction model that masks object interactions, thereby increasing reliance on object interactions in the training of the ML prediction model.
1 . A system comprising:
one or more processors; and
one or more computer-readable media storing computer-executable instructions that, when executed, cause the one or more processors to perform operations comprising:
receiving data representing an environment, the data including a first trajectory associated with a first object in the environment and a second trajectory associated with a second object in the environment;
providing at least a portion of the data as a first input to a machine learning model, wherein the machine learning model is a joint prediction model configured to output a set of compatible trajectory predictions within the environment;
determining, based at least in part on a first output of the machine learning model associated with the first input, a first predicted trajectory for the first object and a second predicted trajectory for the second object;
determining a first loss value associated with the machine learning model, based at least in part on:
a first difference between the first trajectory and the first predicted trajectory; and
a second difference between the second trajectory and the second predicted trajectory;
determining modified data representing the environment by modifying, in the data representing the environment, relative object data between the first object and the second object, wherein:
the modifying relative object data comprises at least one of masking interaction data representing an interaction between the first object and the second object, removing the interaction data, or down-weighting the interaction data, and
wherein the interaction data subject to modification represents at least one of a relative position between the first object and the second object, a relative orientation between the first object and the second object, and a relative motion difference between the first object and the second object;
providing the modified data as a second input to a second machine learning model, and receiving a second output of the second machine learning model associated with the second input;
determining, based at least in part on a difference between the first output of the machine learning model and the second output of the second machine learning model being lower than a threshold value, a second loss value associated with the difference;
training the machine learning model based at least in part on the first loss value and the second loss value, to determine a trained machine learning model; and
transmitting the trained machine learning model to a computing device associated with a vehicle, wherein operation of the vehicle is based at least in part on executing the trained machine learning model.
2 . The system of claim 1 , wherein the first output of the machine learning model comprises a first set of predicted trajectories including the first predicted trajectory and the second predicted trajectory, the operations further comprising: determining, based at least in part on the second output of the second machine learning model, a second set of predicted trajectories including a third predicted trajectory for the first object and a fourth predicted trajectory for the second object; and
determining the second loss value based at least in part by comparing the first set of predicted trajectories to the second set of predicted trajectories.
3 . The system of claim 2 , wherein determining the second loss value comprises:
increasing a prediction loss value associated with the machine learning model, based at least in part on determining that a difference between the first set of predicted trajectories and the second set of predicted trajectories is less than a difference threshold.
4 . A method comprising:
receiving data representing an environment at a first time, the environment including a first object and a second object;
providing at least a first portion of the data as a first input to a first machine learning model;
determining, based at least in part on an output of the first machine learning model, a first set of predicted trajectories including a first predicted trajectory for the first object and a second predicted trajectory for the second object;
determining modified data representing the environment by modifying, in the data representing the environment, relative object data between the first object and the second object, wherein the modifying relative object data comprises at least one of masking interaction data representing an interaction between the first object and the second object, removing the interaction data, or down-weighting the interaction data;
providing the modified data as a second input to a second machine learning model, wherein the second machine learning model is configured to modify interaction data representing an interaction between the first object and the second object;
determining, based at least in part on a second output of the second machine learning model, a second set of predicted trajectories including a third predicted trajectory for the first object and a fourth predicted trajectory for the second object;
determining a loss value based at least in part on a difference between the output of the first machine learning model and the second output of the second machine learning model, the difference being lower than a threshold value;
training the first machine learning model based at least in part on the loss value, to determine a trained machine learning model; and
controlling operation of a vehicle, based at least in part on the trained machine learning model.
5 . The method of claim 4 , wherein determining the loss value comprises:
determining the loss value based at least in part by comparing the first set of predicted trajectories to the second set of predicted trajectories.
6 . The method of claim 5 , wherein determining the loss value comprises:
increasing a prediction loss value associated with the first machine learning model, based at least in part on determining that a difference between the first set of predicted trajectories and the second set of predicted trajectories is less than a difference threshold.
7 . The method of claim 4 , further comprising:
determining, based at least in part on the data, a first ground truth trajectory associated with the first object and a second ground truth trajectory associated with the second object;
determining a second difference between the first ground truth trajectory and the first predicted trajectory, and a third difference between the second ground truth trajectory and the second predicted trajectory;
determining a second loss value based at least in part on the second difference and the third difference; and
training the first machine learning model based at least in part on the loss value and the second loss value, to determine the trained machine learning model.
8 . The method of claim 4 , wherein determining the loss value comprises:
determining a driving context variable, based at least in part on at least one of:
an object density of the environment;
a location of the environment;
a velocity of the first object in the environment; or
a velocity of the second object in the environment; and
weighting the loss value based at least in part on the driving context variable.
9 . The method of claim 4 , wherein controlling the operation of the vehicle comprises:
transmitting the trained machine learning model to a computing device associated with the vehicle, wherein the operation of the vehicle is based at least in part on executing the trained machine learning model.
10 . One or more non-transitory computer-readable media storing instructions executable by a processor, wherein the instructions, when executed, cause the processor to perform operations comprising:
receiving data representing an environment at a first time, the environment including a first object and a second object;
providing at least a first portion of the data as a first input to a first machine learning model;
determining, based at least in part on an output of the first machine learning model, a first set of predicted trajectories including a first predicted trajectory for the first object and a second predicted trajectory for the second object;
determining modified data representing the environment by modifying, in the data representing the environment, relative object data between the first object and the second object, wherein:
the modifying relative object data comprises at least one of masking interaction data representing an interaction between the first object and the second object, removing the interaction data, or down-weighting the interaction data, and
wherein the interaction data subject to modification represents at least one of a relative position between the first object and the second object, a relative orientation between the first object and the second object, or a relative motion difference between the first object and the second object;
providing the modified data as a second input to a second machine learning model, wherein the second machine learning model is configured to modify interaction data representing an interaction between the first object and the second object;
determining, based at least in part on a second output of the second machine learning model, a second set of predicted trajectories including a third predicted trajectory for the first object and a fourth predicted trajectory for the second object;
determining a loss value based at least in part on a difference between the output of the first machine learning model and the second output of the second machine learning model, the difference being lower than a threshold value;
training the first machine learning model based at least in part on the loss value, to determine a trained machine learning model; and
controlling operation of a vehicle, based at least in part on the trained machine learning model.
11 . The one or more non-transitory computer-readable media of claim 10 , wherein determining the loss value comprises:
determining the loss value based at least in part by comparing the first set of predicted trajectories to the second set of predicted trajectories.
12 . The one or more non-transitory computer-readable media of claim 11 , wherein determining the loss value comprises:
increasing a prediction loss value associated with the first machine learning model, based at least in part on determining that a difference between the first set of predicted trajectories and the second set of predicted trajectories is less than a difference threshold.
13 . The one or more non-transitory computer-readable media of claim 10 , the operations further comprising:
determining, based at least in part on the data, a first ground truth trajectory associated with the first object and a second ground truth trajectory associated with the second object;
determining a second difference between the first ground truth trajectory and the first predicted trajectory, and a third difference between the second ground truth trajectory and the second predicted trajectory;
determining a second loss value based at least in part on the second difference and the third difference; and
training the first machine learning model based at least in part on the loss value and the second loss value, to determine the trained machine learning model.
14 . The one or more non-transitory computer-readable media of claim 10 , wherein determining the loss value comprises:
determining a driving context variable, based at least in part on at least one of:
an object density of the environment;
a location of the environment;
a velocity of the first object in the environment; or
a velocity of the second object in the environment; and
weighting the loss value based at least in part on the driving context variable.
15 . The method of claim 4 , wherein:
the second machine learning model is identical to the first machine learning model; and
providing the second input to the second machine learning model comprises:
determining modified data representing the environment by modifying, in the data representing the environment, relative object data between the first object and the second object; and
providing the modified data as the second input to the second machine learning model.
16 . The method of claim 15 , wherein:
the data representing the environment comprises graph neural network (GNN) representation of the environment; and
determining the modified data comprises modifying or removing, within the GNN representation, an edge feature representing a relative state difference between the first object and the second object.
17 . The method of claim 4 , wherein:
the first machine learning model is a first predictive model configured to output the first set of predicted trajectories, based at least in part on an indication within the data representing the environment of a potential interaction between the first object and the second object; and
the second machine learning model is a second predictive model configured to output the second set of predicted trajectories, based at least in part on masking out the indication of the potential interaction between the first object and the second object.
18 . The one or more non-transitory computer-readable media of claim 10 , wherein:
the second machine learning model is identical to the first machine learning model; and
providing the second input to the second machine learning model comprises:
determining modified data representing the environment by modifying, in the data representing the environment, relative object data between the first object and the second object; and
providing the modified data as the second input to the second machine learning model.
19 . The system of claim 1 , wherein the second loss value is a divergent loss value and is increased based at least in part on a magnitude associated with the difference decreasing.
20 . The method of claim 17 , wherein the trained machine learning model is, based at least in part on using the loss value, associated with an increased responsiveness to the potential interaction compared to the first machine learning model.