Method and apparatus for controlling traffic light
A method for controlling a traffic light, a method and apparatus for navigating an unmanned vehicle and a method and apparatus for training a model are provided. An implementation comprises: generating a reinforced traffic light state parameter according to vehicle state representation information of an unmanned vehicle currently contained in a preset area of a target traffic light and a current traffic light state parameter of the target traffic light; and generating a traffic light control action according to the reinforced traffic light state parameter; where the reinforced traffic light state parameter is used to cause an unmanned vehicle navigation end to generate a reinforced vehicle state parameter according to a reinforced traffic light state and a current vehicle state parameter of a target unmanned vehicle, and generate an unmanned vehicle navigation action according to the reinforced vehicle state parameter.
1 . A method for controlling a traffic light, applied to a traffic light control end communicating with an unmanned vehicle navigation end, the method comprising:
generating a reinforced traffic light state parameter according to vehicle state representation information of an unmanned vehicle currently contained in a preset area of a target traffic light and a current traffic light state parameter of the target traffic light; and
generating, according to the reinforced traffic light state parameter, a traffic light control action matching the reinforced traffic light state parameter, wherein the generating the traffic light control action matching the reinforced traffic light state parameter according to the reinforced traffic light state parameter comprises:
inputting the reinforced traffic light state parameter into a first reinforcement learning model, to obtain the traffic light control action matching the reinforced traffic light state parameter;
wherein the reinforced traffic light state parameter is used to cause the unmanned vehicle navigation end to:
generate a reinforced vehicle state parameter according to the reinforced traffic light state parameter and a current vehicle state parameter of the unmanned vehicle, and generate an unmanned vehicle navigation action matching the reinforced vehicle state parameter according to the reinforced vehicle state parameter,
wherein the method further comprises training the first reinforcement learning model, and the training comprises:
generating a sample reinforced traffic light state parameter according to sample vehicle state representation information of a sample unmanned vehicle currently contained in a sample preset area of a sample target traffic light and a sample current traffic light state parameter of the sample target traffic light;
inputting the sample reinforced traffic light state parameter into the first reinforcement learning model, to obtain a sample traffic light control action matching the sample reinforced traffic light state parameter;
performing the sample traffic light control action, to obtain a new traffic light state parameter and a first reward parameter; and
determining a first loss value, based on the first reward parameter, the new traffic light state parameter, and the sample reinforced traffic light state parameter; and
training the first reinforcement learning model according to the first loss value.
2 . The method according to claim 1 , wherein the vehicle state representation information is generated by the unmanned vehicle navigation end according to a current vehicle state parameter of the unmanned vehicle contained in the preset area and historical vehicle state representation information at a plurality of previous moments.
3 . The method according to claim 1 , wherein the generating a reinforced traffic light state parameter according to vehicle state representation information of an unmanned vehicle currently contained in a preset area of a target traffic light and a current traffic light state parameter of the target traffic light comprises:
stitching the vehicle state representation information and the current traffic light state parameter into hybrid environment information; and
inputting the hybrid environment information into a first encoder, to obtain the reinforced traffic light state parameter.
4 . The method according to claim 1 , wherein the generating, according to the reinforced traffic light state parameter, a traffic light control action matching the reinforced traffic light state parameter comprises:
acquiring associated traffic light state aggregation information of associated traffic lights associated with the target traffic light; and
generating the traffic light control action matching the reinforced traffic light state parameter, according to the reinforced traffic light state parameter and the associated traffic light state aggregation information.
5 . The method according to claim 4 , wherein the associated traffic light state aggregation information is generated by:
generating an associated traffic light state matrix according to current traffic light state parameters of the associated traffic lights; and
generating the associated traffic light state aggregation information, according to the associated traffic light state matrix, a connectivity parameter of the target traffic light, and a weight matrix of the target traffic light.
6 . The method according to claim 5 , wherein the generating the associated traffic light state aggregation information, according to the associated traffic light state matrix, a connectivity parameter of the target traffic light, and a weight matrix of the target traffic light comprises:
generating, through a first graph neural network, the associated traffic light state aggregation information according to the associated traffic light state matrix, the connectivity parameter of the target traffic light, and the weight matrix of the target traffic light.
7 . The method according to claim 1 , wherein, after the generating a reinforced traffic light state parameter according to vehicle state representation information of an unmanned vehicle currently contained in a preset area of a target traffic light and a current traffic light state parameter of the target traffic light, the method further comprises:
inputting the reinforced traffic light state parameter into a pre-trained goal network to obtain a goal vector,
wherein the unmanned vehicle navigation end generates, through a second reinforcement learning model, the unmanned vehicle navigation action matching the reinforced vehicle state parameter according to the reinforced vehicle state parameter; and the goal vector is used to cause the unmanned vehicle navigation end to adjust the second reinforcement learning model according to the goal vector.
8 . The method according to claim 1 , wherein the method further comprises navigating the unmanned vehicle, applied to the unmanned vehicle navigation end communicating with the traffic light control end, the method comprising:
generating the reinforced vehicle state parameter according to the current reinforced traffic light state parameter of the target traffic light that is acquired from the traffic light control end and the current vehicle state parameter of the unmanned vehicle; and
generating, according to the reinforced vehicle state parameter, the unmanned vehicle navigation action matching the reinforced vehicle state parameter.
9 . The method according to claim 8 , wherein, before the generating the reinforced vehicle state parameter according to the current reinforced traffic light state parameter of the target traffic light that is acquired from the traffic light control end and a current vehicle state parameter of the unmanned vehicle, the method further comprises:
generating vehicle state aggregation information, according to the vehicle state parameter of an unmanned vehicle currently contained in a preset area of the target traffic light; and
generating the current vehicle state representation information according to the vehicle state aggregation information and historical vehicle state representation information at a plurality of previous moments,
wherein the vehicle state representation information is used to cause the traffic light control end to generate the reinforced traffic light state parameter according to the vehicle state representation information of the unmanned vehicle currently contained in the preset area of the target traffic light and the current traffic light state parameter of the target traffic light.
10 . The method according to claim 9 , wherein the generating vehicle state aggregation information according to a vehicle state parameter of an unmanned vehicle currently contained in a preset area of the target traffic light comprises:
generating, through a second graph neural network, the vehicle state aggregation information according to the vehicle state parameter of the unmanned vehicle currently contained in the preset area of the target traffic light.
11 . The method according to claim 9 , wherein the generating the current vehicle state representation information according to the vehicle state aggregation information and historical vehicle state representation information at a plurality of previous moments comprises:
inputting the vehicle state aggregation information and the historical vehicle state representation information at the plurality of previous moments into a recurrent neural network, to obtain the current vehicle state representation information; or
constructing a linear function according to the historical vehicle state representation information at the plurality of previous moments, and obtaining, through the linear function, the current vehicle state representation information according to the vehicle state aggregation information.
12 . The method according to claim 8 , wherein the generating, according to the reinforced vehicle state parameter, an unmanned vehicle navigation action matching the reinforced vehicle state parameter comprises:
inputting the reinforced vehicle state parameter into a second reinforcement learning model, to obtain the unmanned vehicle navigation action matching the reinforced vehicle state parameter.
13 . The method according to claim 12 , further comprising:
adjusting the second reinforcement learning model according to a goal vector, wherein the goal vector is generated by the traffic light control end by inputting the reinforced traffic light state parameter into a pre-trained goal network.
14 . The method according to claim 1 , wherein the inputting the sample reinforced traffic light state parameter into the first reinforcement learning model, to obtain the sample traffic light control action matching the sample reinforced traffic light state parameter comprises:
acquiring associated traffic light state aggregation information of associated traffic lights associated with the sample target traffic light; and
inputting the sample reinforced traffic light state parameter and the associated traffic light state aggregation information into the first reinforcement learning model, to obtain the sample traffic light control action matching the sample reinforced traffic light state parameter,
wherein the associated traffic light state aggregation information is generated by:
generating an associated traffic light state matrix according to current traffic light state parameters of associated traffic light states; and
generating the associated traffic light state aggregation information, according to the associated traffic light state matrix, a connectivity parameter of the sample target traffic light, and a weight matrix of the sample target traffic light, and
the determining a first loss value, based on the first reward parameter, the new traffic light state parameter, and the sample reinforced traffic light state parameter comprises:
determining the first loss value based on the first reward parameter, the new traffic light state parameter, a weight matrix corresponding to the new traffic light state parameter, the reinforced traffic light state parameter, and a weight matrix corresponding to the sample reinforced traffic light state parameter, wherein the weight matrices are learned and obtained during the training of the first reinforcement learning model.
15 . The method according to claim 7 , wherein the method further comprises training the second reinforcement learning model, the training comprising:
generating a sample reinforced vehicle state parameter, according to a current reinforced traffic light state parameter of a sample target traffic light that is acquired from a sample traffic light control end and a sample current vehicle state parameter of a sample target-unmanned vehicle;
inputting the sample reinforced vehicle state parameter into the second reinforcement learning model, to obtain a sample unmanned vehicle navigation action matching the sample reinforced vehicle state parameter;
performing the sample unmanned vehicle navigation action, to obtain a new vehicle state parameter and a second reward parameter;
determining a second loss value, based on the second reward parameter, the new vehicle state parameter and the sample reinforced vehicle state parameter; and
training the second reinforcement learning model according to the second loss value.
16 . The method according to claim 15 , wherein, after the performing the sample unmanned vehicle navigation action to obtain a new vehicle state parameter and a second reward parameter, and before the determining a second loss value based on the second reward parameter, the new vehicle state parameter and the sample reinforced vehicle state parameter, the method further comprises:
determining an additional reward parameter, according to a goal vector, the sample current vehicle state parameter, and an ideal vehicle state parameter predicted by the second reinforcement learning model; and
updating the second reward parameter according to the additional reward parameter,
wherein the goal vector is generated by the sample traffic light control end by inputting the sample reinforced traffic light state parameter into a pre-trained goal network.
17 . An apparatus for controlling a traffic light according to claim 1 , comprising:
at least one processor; and
a memory, communicating with the at least one processor, wherein
the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor, to enable the at least one processor to perform operations comprising:
generating a reinforced traffic light state parameter according to vehicle state representation information of an unmanned vehicle currently contained in a preset area of a target traffic light and a current traffic light state parameter of the target traffic light; and
generating, according to the reinforced traffic light state parameter, a traffic light control action matching the reinforced traffic light state parameter, wherein the generating the traffic light control action matching the reinforced traffic light state parameter according to the reinforced traffic light state parameter comprises: inputting the reinforced traffic light state parameter into a first reinforcement learning model, to obtain the traffic light control action matching the reinforced traffic light state parameter;
wherein the reinforced traffic light state parameter is used to cause the unmanned vehicle navigation end to: generate a reinforced vehicle state parameter according to the reinforced traffic light state parameter and a current vehicle state parameter of the unmanned vehicle, and generate an unmanned vehicle navigation action matching the reinforced vehicle state parameter according to the reinforced vehicle state parameter,
wherein the method further comprises training the first reinforcement learning model, and the training comprises:
generating a sample reinforced traffic light state parameter according to sample vehicle state representation information of a sample unmanned vehicle currently contained in a sample preset area of a sample target traffic light and a sample current traffic light state parameter of the sample target traffic light;
inputting the sample reinforced traffic light state parameter into the first reinforcement learning model, to obtain a sample traffic light control action matching the sample reinforced traffic light state parameter;
performing the sample traffic light control action, to obtain a new traffic light state parameter and a first reward parameter; and
determining a first loss value, based on the first reward parameter, the new traffic light state parameter, and the sample reinforced traffic light state parameter; and
training the first reinforcement learning model according to the first loss value.
18 . An apparatus for navigating an unmanned vehicle according to claim 8 , comprising:
at least one processor; and
a memory, communicating with the at least one processor, wherein
the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor, to enable the at least one processor to perform the method according to claim 8 .