Systems and methods for assisting a driver using a foundation model in a shared-autonomy driving mode of a vehicle
Systems and methods for assisting a driver using a foundation model in a shared-autonomy driving mode of a vehicle are disclosed herein. One embodiment of a shared-autonomy assistance subsystem processes, in a vehicle operating in a shared-autonomy driving mode, inputs including vehicle state information, external-road-agent state information, vehicle environmental sensor data, and map data using one or more encoder neural networks that have been trained to extract features for a large language model (LLM). The subsystem inputs the extracted features to the LLM. The subsystem predicts, using the LLM, an objective of a driver of the vehicle. The subsystem then executes, based on an output from the LLM, one or more actions to assist the driver in meeting the predicted objective. The one or more actions include controlling, at least in part, operation of the vehicle.
1 . A system, comprising:
a processor; and
a memory storing machine-readable instructions that, when executed by the processor, cause the processor to:
process, in a vehicle operating in a shared-autonomy driving mode, inputs including a driver vehicle-control input that includes at least one of turning a steering wheel of the vehicle, operating a throttle of the vehicle, or operating a brake of the vehicle, vehicle state information, external-road-agent state information, vehicle environmental sensor data, and map data using one or more encoder neural networks that have been trained to extract features for a large language model (LLM);
input the extracted features to the LLM;
predict, using the LLM based on the extracted features, an objective of a driver of the vehicle; and
execute, based on an output from the LLM, one or more actions to assist the driver in meeting the predicted objective, wherein the one or more actions include controlling, at least in part, operation of the vehicle.
2 . The system of claim 1 , wherein the machine-readable instructions include further instructions that, when executed by the processor, cause the processor to input, to the LLM, a driver language input that includes at least one of speech or text.
3 . The system of claim 1 , wherein the predicted objective is one of changing lanes, remaining in a current lane, merging, exiting a roadway, overtaking another vehicle, and executing a turn at an intersection.
4 . The system of claim 1 , wherein the controlling, at least in part, the operation of the vehicle, includes controlling at least one of steering, acceleration, or braking while retaining the predicted objective of the driver.
5 . The system of claim 1 , wherein the LLM selects the one or more actions based, at least in part, on learned past driving behavior of the driver and the one or more actions include at least one of the LLM advising the driver, the LLM warning the driver, the LLM activating a turn signal, the LLM controlling headlight high beams, the LLM activating a horn, the LLM controlling hazard lights, or the LLM controlling windshield wipers.
6 . The system of claim 1 , wherein the LLM detects, based on the extracted features, that the driver is distracted and the one or more actions compensate for the driver being distracted.
7 . The system of claim 1 , wherein the machine-readable instructions include further instructions that, when executed by the processor, cause the processor to output, from the LLM, a question to the driver and to process, via the LLM, a reply from the driver to confirm the predicted objective before executing the one or more actions.
8 . A non-transitory computer-readable medium storing instructions that, when executed by a processor, cause the processor to:
process, in a vehicle operating in a shared-autonomy driving mode, inputs including a driver vehicle-control input that includes at least one of turning a steering wheel of the vehicle, operating a throttle of the vehicle, or operating a brake of the vehicle, vehicle state information, external-road-agent state information, vehicle environmental sensor data, and map data using one or more encoder neural networks that have been trained to extract features for a large language model (LLM);
input the extracted features to the LLM;
predict, using the LLM based on the extracted features, an objective of a driver of the vehicle; and
execute, based on an output from the LLM, one or more actions to assist the driver in meeting the predicted objective, wherein the one or more actions include controlling, at least in part, operation of the vehicle.
9 . The non-transitory computer-readable medium of claim 8 , wherein the controlling, at least in part, the operation of the vehicle, includes controlling at least one of steering, acceleration, or braking while retaining the predicted objective of the driver.
10 . The non-transitory computer-readable medium of claim 8 , wherein the LLM selects the one or more actions based, at least in part, on learned past driving behavior of the driver and the one or more actions include at least one of the LLM advising the driver, the LLM warning the driver, the LLM activating a turn signal, the LLM controlling headlight high beams, the LLM activating a horn, the LLM controlling hazard lights, or the LLM controlling windshield wipers.
11 . The non-transitory computer-readable medium of claim 8 , wherein the instructions include further instructions that, when executed by the processor, cause the processor to output, from the LLM, a question to the driver and to process, via the LLM, a reply from the driver to confirm the predicted objective before executing the one or more actions.
12 . A method, comprising:
processing, in a vehicle operating in a shared-autonomy driving mode, inputs including a driver vehicle-control input that includes at least one of turning a steering wheel of the vehicle, operating a throttle of the vehicle, or operating a brake of the vehicle, vehicle state information, external-road-agent state information, vehicle environmental sensor data, and map data using one or more encoder neural networks that have been trained to extract features for a large language model (LLM) and inputting the extracted features to the LLM;
predicting, using the LLM based on the extracted features, an objective of a driver of the vehicle; and
executing, based on an output from the LLM, one or more actions to assist the driver in meeting the predicted objective, wherein the one or more actions include controlling, at least in part, operation of the vehicle.
13 . The method of claim 12 , further comprising inputting, to the LLM, a driver language input that includes at least one of speech or text.
14 . The method of claim 12 , wherein the predicted objective is one of changing lanes, remaining in a current lane, merging, exiting a roadway, overtaking another vehicle, and executing a turn at an intersection.
15 . The method of claim 12 , wherein the controlling, at least in part, the operation of the vehicle, includes controlling at least one of steering, acceleration, or braking while retaining the predicted objective of the driver.
16 . The method of claim 12 , wherein the LLM selects the one or more actions based, at least in part, on learned past driving behavior of the driver and the one or more actions include at least one of the LLM advising the driver, the LLM warning the driver, the LLM activating a turn signal, the LLM controlling headlight high beams, the LLM activating a horn, the LLM controlling hazard lights, or the LLM controlling windshield wipers.
17 . The method of claim 12 , wherein the LLM detects, based on the extracted features, that the driver is distracted and the one or more actions compensate for the driver being distracted.
18 . The method of claim 12 , further comprising outputting, from the LLM, a question to the driver and processing, via the LLM, a reply from the driver to confirm the predicted objective before executing the one or more actions.