IP Library Granted Patent US 12,654,741
Granted Patent B2
US 12,654,741 · App. 18/521,128 · Granted Jun 16, 2026

Driving policy determining method and apparatus, device, and vehicle

Inventors: Shixiong Kai (Beijing, CN); Bin Wang (Beijing, CN); Wulong Liu (Montreal, CA)
Assignee: YINWANG INTELLIGENT TECHNOLOGIES CO., LTD.
B60W60/0011B60W30/09B60W30/0956B60W30/18159B60W40/04B60W40/06B60W40/105B60W60/00274B60W2554/4041B60W2554/4045B60W2554/406B60W2556/10
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,654,741
App. No.
18/521,128
Granted
Jun 16, 2026
Kind
B2
Abstract

In a driving policy determining method, for each object, a first target motion trajectory of the object is calculated on a premise that the object does not collide with another object that moves based on an initial motion trajectory. Then, for the ego vehicle, a second target motion trajectory of the ego vehicle is calculated on a premise that the ego vehicle does not collide with another object that moves based on a first target motion trajectory. Then, a driving policy is determined based on the second target motion trajectory of the ego vehicle and a first target motion trajectory of at least one game object. The foregoing operations are repeated until the determined driving policy matches an initial driving policy.

Claims (63)

1 . A method, comprising:

determining, for each of a plurality of objects, a first target sequence that satisfies a first condition, wherein the first target sequence comprises values of a motion parameter at a plurality of first moments after an object of the plurality of objects starts moving from a current location of the object, wherein the first condition is that the object does not collide, in a process of moving from the current location of the object based on a current speed and the first target sequence, with other objects that are of the plurality of objects and that move based on an initial motion trajectory of the other objects, wherein the initial motion trajectory comprises locations of the other objects at a plurality of second moments, and wherein the plurality of objects comprises an ego vehicle and at least two game objects;

calculating a first target motion trajectory generated when each object moves from the current location of the object based on the current speed and the first target sequence, wherein the first target motion trajectory comprises locations of the object at the plurality of second moments;

determining, for the ego vehicle, a second target sequence that satisfies a second condition, wherein the second target sequence comprises values of the motion parameter at the plurality of first moments after the ego vehicle starts from a current location of the ego vehicle, and wherein the second condition is that the ego vehicle does not collide, in a process of moving from the current location of the ego vehicle based on a current speed and the second target sequence, with each of the at least two game objects that moves based on the first target motion trajectory of each of the at least two game objects;

calculating a second target motion trajectory generated when the ego vehicle moves from the current location of the ego vehicle based on the current speed and the second target sequence, wherein the second target motion trajectory comprises locations of the ego vehicle at the plurality of second moments;

determining a driving policy of the ego vehicle in a current iterative calculation based on the second target motion trajectory of the ego vehicle and the first target motion trajectory of each of the at least two game objects;

repeatedly performing the foregoing steps until the driving policy of the ego vehicle in the current iterative calculation matches an initial driving policy,

outputting, when the driving policy matches the initial driving policy, action instruction information of the ego vehicle based on the driving policy and a road topology; and

controlling, based on the action instruction information, the ego vehicle to complete self-driving.

2 . The method of claim 1 , wherein the motion parameter is acceleration.

3 . The method of claim 1 , wherein a type of the at least two game objects comprises at least one of a pedestrian, a motor vehicle, or a non-motor vehicle.

4 . The method of claim 1 , further comprising:

planning an initial motion trajectory of the ego vehicle based on the current location of the ego vehicle, a destination, and the road topology; and

determining an initial motion trajectory of each of the at least two game objects based on a current location of each of the at least two game objects and the road topology.

5 . The method of claim 1 , wherein respective initial motion trajectories of the plurality of objects are respective first target motion trajectories of a plurality of target objects from a previous iterative calculation of a current iterative calculation.

6 . The method of claim 1 , wherein the driving policy is at least one of non-yielding, yielding, or car-following.

7 . The method of claim 1 , wherein the initial driving policy is a driving policy in a previous iterative calculation of the current iterative calculation.

8 . The method of claim 1 , wherein determining the first target sequence that satisfies the first condition comprises:

obtaining, for each of the plurality of objects, a plurality of first action sequences, wherein each of the plurality of first action sequences comprises the values of the motion parameter at the plurality of first moments after the object starts from the current location of the object; and

selecting, from the plurality of first action sequences, a first action sequence that satisfies the first condition as the first target sequence.

9 . The method of claim 8 , wherein selecting the first action sequence that satisfies the first condition as the first target sequence comprises:

calculating a score of each of the plurality of first action sequences; and

selecting, from the plurality of first action sequences, the first action sequence that satisfies the first condition and whose score is greater than a first threshold as the first target sequence.

10 . The method of claim 9 , wherein calculating the score of each of the plurality of first action sequences comprises:

calculating, for each of the plurality of first action sequences and based on the first action sequence, a difference between accelerations of the object at two adjacent first moments of the plurality of first moments; and

calculating the score of each of the plurality of first action sequences based on the difference between the accelerations at the two adjacent first moments of the plurality of first moments.

11 . The method of claim 9 , wherein the motion parameter is acceleration, and wherein calculating the score of each of the plurality of first action sequences comprises calculating the score of each of the plurality of first action sequences based on a value of the acceleration that is comprised in each of the plurality of first action sequences and that is at each first moment.

12 . The method of claim 8 , wherein obtaining the plurality of first action sequences comprises obtaining, for each of the plurality of objects and based on a plurality of first reference values of the motion parameter, the plurality of first action sequences that satisfy a fourth condition, wherein the value of the motion parameter comprised in each of the plurality of first action sequences belongs to the plurality of first reference values, and wherein the fourth condition comprises at least one of the following conditions:

a value range of acceleration of the object at each of the plurality of first moments;

a value range of a speed of the object at each first moment; or

a range of a difference between the accelerations of the object at two adjacent first moments of the plurality of first moments, wherein the acceleration of the object at each first moment, the speed of the object at each first moment, and the difference between the accelerations of the object at the two adjacent first moments of the plurality of first moments are based on the first action sequence.

13 . The method of claim 12 , wherein the value range of the acceleration of the object at each first moment is based on a type of the object.

14 . The method of claim 12 , further comprising determining the value range of the speed of the object at each first moment based on at least one of a type of the object or intention information of a motion of the object, wherein the intention information of the motion of the object comprises at least one of turning left at an intersection, turning right at the intersection, going straight at the intersection, entering a roundabout, or leaving the roundabout.

15 . The method of claim 1 , wherein determining the second target sequence that satisfies the second condition comprises:

obtaining, for the ego vehicle, a plurality of second action sequences, wherein each of the plurality of second action sequences comprises the values of the motion parameter at the plurality of first moments after the ego vehicle starts from the current location of the ego vehicle; and

selecting, from the plurality of second action sequences, a second action sequence that satisfies the second condition as the second target sequence.

16 . The method of claim 15 , wherein selecting the second action sequence that satisfies the second condition as the second target sequence comprises:

calculating a score of each of the plurality of second action sequences; and

selecting, from the plurality of second action sequences, a second action sequence that satisfies the second condition and whose score is greater than a second threshold as the second target sequence.

17 . The method of claim 16 , wherein calculating the score of each of the plurality of second action sequences comprises:

calculating, for each of the plurality of second action sequences and based on the second action sequence, a difference between accelerations of the ego vehicle at two adjacent first moments of the plurality of first moments; and

calculating the score of each of the plurality of second action sequences based on the difference between the accelerations at the two adjacent first moments of the plurality of first moments.

18 . The method of claim 16 , wherein the motion parameter is acceleration, and wherein calculating the score of each of the plurality of second action sequences comprises calculating the score of each of the plurality of second action sequences based on a value of the acceleration that is comprised in each of the plurality of second action sequences and that is at each first moment.

19 . A policy determining apparatus, comprising:

a memory configured to store programming instructions; and

at least one processor coupled to the memory and configured to execute the programming instructions to cause the policy determining apparatus to:

determine, for each of a plurality of objects, a first target sequence that is of the object and that satisfies a first condition, wherein the first target sequence comprises values of a motion parameter at a plurality of first moments after an object of the plurality of objects starts moving from a current location of the object, wherein the first condition is that the object does not collide, in a process of moving from the current location of the object based on a current speed and the first target sequence, with another object that are of the plurality of objects and that move based on an initial motion trajectory of the another object, wherein the initial motion trajectory comprises locations of the another object at a plurality of second moments, and wherein the plurality of objects comprises an ego vehicle and at least two game objects;

calculate a first target motion trajectory generated when each object moves from the current location of the object based on the current speed and the first target sequence, wherein the first target motion trajectory comprises locations of the object at the plurality of second moments;

determine, for the ego vehicle, a second target sequence that satisfies a second condition, wherein the second target sequence comprises values of the motion parameter at the plurality of first moments after the ego vehicle starts from a current location of the ego vehicle, and wherein the second condition is that the ego vehicle does not collide, in a process of moving from the current location of the ego vehicle based on a current speed and the second target sequence, with each of the at least two game objects that moves based on the first target motion trajectory of each of the at least two game objects;

calculate a second target motion trajectory generated when the ego vehicle moves from the current location of the ego vehicle based on the current speed and the second target sequence, wherein the second target motion trajectory comprises locations of the ego vehicle at the plurality of second moments;

determine a driving policy of the ego vehicle in a current iterative calculation based on the second target motion trajectory of the ego vehicle and the first target motion trajectory of each of the at least two game objects;

repeatedly perform the foregoing steps until the driving policy of the ego vehicle in the current iterative calculation matches an initial driving policy;

output, when the driving policy matches the initial driving policy, action instruction information of the ego vehicle based on the driving policy and a road topology; and

control, based on the action instruction information, the ego vehicle to complete self-driving.

20 . A computer program product comprising computer-executable instructions stored on a non-transitory computer-readable storage medium that, when executed by a processor, cause an apparatus to:

determine, for each of a plurality of objects, a first target sequence that satisfies a first condition, wherein the first target sequence comprises values of a motion parameter at a plurality of first moments after an object of the plurality of objects starts moving from a current location of the object, wherein the first condition is that the object does not collide, in a process of moving from the current location of the object based on a current speed and the first target sequence, with other objects that are of the plurality of objects and that move based on an initial motion trajectory of the other objects, wherein the initial motion trajectory comprises locations of the other objects at a plurality of second moments, and wherein the plurality of objects comprises an ego vehicle and at least two game objects;

calculate a first target motion trajectory generated when each object moves from the current location of the object based on the current speed and the first target sequence, wherein the first target motion trajectory comprises locations of the object at the plurality of second moments;

determine, for the ego vehicle, a second target sequence that satisfies a second condition, wherein the second target sequence comprises values of the motion parameter at the plurality of first moments after the ego vehicle starts from a current location of the ego vehicle, and wherein the second condition is that the ego vehicle does not collide, in a process of moving from the current location of the ego vehicle based on a current speed and the second target sequence, with each of the at least two game objects that moves based on the first target motion trajectory of each of the at least two game objects;

calculate a second target motion trajectory generated when the ego vehicle moves from the current location of the ego vehicle based on the current speed and the second target sequence, wherein the second target motion trajectory comprises locations of the ego vehicle at the plurality of second moments;

determine a driving policy of the ego vehicle in a current iterative calculation based on the second target motion trajectory of the ego vehicle and the first target motion trajectory of each of the at least two game objects;

repeatedly perform the foregoing steps until the driving policy of the ego vehicle in the current iterative calculation matches an initial driving policy;

output, when the driving policy matches the initial driving policy, action instruction information of the ego vehicle based on the driving policy and a road topology; and

control, based on the action instruction information, the ego vehicle to complete self-driving.

Assignments (3)
CHANGE OF NAME Recorded Apr 29, 2026
From: SHENZHEN YINWANG INTELLIGENT TECHNOLOGIES CO., LTD.
To: YINWANG INTELLIGENT TECHNOLOGIES CO., LTD.
Reel/Frame 074513/0969 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Nov 12, 2024
From: HUAWEI TECHNOLOGIES CO., LTD.
To: SHENZHEN YINWANG INTELLIGENT TECHNOLOGIES CO., LTD.
Reel/Frame 069336/0125 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Mar 11, 2024
From: KAI, SHIXIONG; WANG, BIN; LIU, WULONG
To: HUAWEI TECHNOLOGIES CO., LTD.
Reel/Frame 066713/0127 →