IP Library › Granted Patent US 12,444,244
Granted Patent B2
US 12,444,244 · App. 18/145,557 · Granted Oct 14, 2025

Driving decision-making method and apparatus and chip

Inventors: Dong Li (Beijing, CN); Bin Wang (Beijing, CN); Wulong Liu (Montreal, CA); Yuzheng Zhuang (Shenzhen, CN)
Assignee: Huawei Technologies Co., Ltd.
G07C5/02B60W50/06G06N7/01B60W2050/0018B60W60/001
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,444,244
App. No.
18/145,557
Granted
Oct 14, 2025
Kind
B2
Abstract

The present disclosure relates to driving decision-making methods, apparatuses, and chips. One example method includes building a Monte Carlo tree based on a current driving environment state, where the Monte Carlo tree includes a root node and N−1 non-root nodes, each node represents one driving environment state, and a driving environment state represented by any non-root node is predicted by a stochastic model of driving environments. Based on at least one of an access count or a value function of each node in the Monte Carlo tree, a node sequence that starts from the root node and ends at a leaf node is determined, and a driving action sequence is determined based on a driving action corresponding to each node in the node sequence.

Claims (73)

1. A driving decision-making method, comprising:

obtaining, by an autonomous driving vehicle, information of a current driving environment state of the autonomous driving vehicle, wherein the autonomous driving vehicle includes one or more sensors;

constructing, by the autonomous driving vehicle, a Monte Carlo tree based on the current driving environment state, wherein the Monte Carlo tree comprises N nodes, each node represents a corresponding driving environment state, the N nodes comprise a root node and N−1 non-root nodes, the root node represents the current driving environment state, a first driving environment state represented by a first node is predicted by using a stochastic model of driving environments based on a second driving environment state represented by a parent node of the first node and based on a driving action, the driving action is determined by the parent node of the first node in a process of obtaining the first node through expansion, the first node is any node of the N−1 non-root nodes, and N is a positive integer greater than or equal to 2;

determining, by the autonomous driving vehicle, in the Monte Carlo tree based on at least one of an access count or a value function of each node in the Monte Carlo tree, a node sequence, wherein the node sequence comprises a plurality of nodes that starts from the root node and ends at a leaf node;

in response to determining the node sequence, determining, by the autonomous driving vehicle, a driving action sequence of a plurality of future driving steps, wherein each future driving step in the driving action sequence comprises a driving action corresponding to each node comprised in the node sequence, and wherein the driving action sequence is used by the autonomous driving vehicle for driving decision-making;

autonomously driving, by the autonomous driving vehicle, the autonomous driving vehicle based on a first driving action in the driving action sequence;

obtaining, by the autonomous driving vehicle, an actual driving environment state after the first driving action is executed; and

updating, by the autonomous driving vehicle, the stochastic model of driving environments based on the current driving environment state, the first driving action, and the actual driving environment state, wherein

the access count of each node is determined based on access counts of subnodes of the each node and an initial access count of the each node, the value function of the each node is determined based on value functions of subnodes of the each node and an initial value function of the each node, the initial access count of the each node is 1, and the initial value function of the each node is determined based on a value function that matches the corresponding driving environment state represented by the each node.

2. The method according to claim 1 , wherein that the first driving environment state represented by the first node is predicted by using the stochastic model of driving environments based on the second driving environment state represented by the parent node of the first node and based on the driving action comprises:

predicting, through dropout-based forward propagation by using the stochastic model of driving environments, a probability distribution of a driving environment state after the driving action is executed based on the second driving environment state represented by the parent node of the first node; and

obtaining the first driving environment state represented by the first node through sampling from the probability distribution.

3. The method according to claim 1 , wherein that the initial value function of the node is determined based on the value function that matches the driving environment state represented by the node comprises:

selecting, from an episodic memory, a first quantity of target driving environment states that have a highest matching degree with the driving environment state represented by the node; and

determining the initial value function of the node based on value functions respectively corresponding to the first quantity of target driving environment states.

4. The method according to claim 3 , wherein the method further comprises:

when a driving episode ends, determining a cumulative reward return value corresponding to an actual driving environment state after each driving action in the driving episode is executed; and

updating the episodic memory by using, as a value function corresponding to the actual driving environment state, the cumulative reward return value corresponding to the actual driving environment state after each driving action is executed.

5. The method according to claim 1 , wherein the node sequence is determined:

based on the access count of the each node in the Monte Carlo tree according to a maximum access count rule;

based on the value function of the each node in the Monte Carlo tree according to a maximum value function rule; or

based on the access count and the value function of the each node in the Monte Carlo tree according to a “maximum access count first, maximum value function next” rule.

6. The method according to claim 1 , wherein obtaining, by the autonomous driving vehicle, information of the current driving environment state of the autonomous driving vehicle comprises:

receiving, from a vehicle velocity sensor of the autonomous driving vehicle, a velocity of the autonomous driving vehicle.

7. The method according to claim 1 , obtaining, by the autonomous driving vehicle, information of the current driving environment state of the autonomous driving vehicle comprises:

receiving, from an acceleration sensor of the autonomous driving vehicle, an acceleration of the autonomous driving vehicle.

8. The method according to claim 1 , wherein obtaining, by the autonomous driving vehicle, information of the current driving environment state of the autonomous driving vehicle comprises:

receiving, from a distance sensor of the autonomous driving vehicle, a relative distance between the autonomous driving vehicle and another vehicle.

9. A driving decision-making apparatus in an autonomous driving vehicle, comprising:

at least one processor; and

one or more memories coupled to the at least one processor and storing programming instructions for execution by the at least one processor to:

obtain information of a current driving environment state of the autonomous driving vehicle, wherein the autonomous driving vehicle includes one or more sensors;

construct a Monte Carlo tree based on the current driving environment state, wherein the Monte Carlo tree comprises N nodes, each node represents a corresponding driving environment state, the N nodes comprise a root node and N−1 non-root nodes, the root node represents the current driving environment state, a first driving environment state represented by a first node is predicted by using a stochastic model of driving environments based on a second driving environment state represented by a parent node of the first node and based on a driving action, the driving action is determined by the parent node of the first node in a process of obtaining the first node through expansion, the first node is any node of the N−1 non-root nodes, and N is a positive integer greater than or equal to 2;

determine, in the Monte Carlo tree based on at least one of an access count or a value function of each node in the Monte Carlo tree, a node sequence, wherein the node sequence comprises a plurality of nodes that starts from the root node and ends at a leaf node;

in response to determining the node sequence, determine a driving action sequence of a plurality of future driving steps, wherein each driving future step in the driving action sequence comprises a driving action corresponding to each node comprised in the node sequence, and wherein the driving action sequence is used for driving decision-making;

autonomously drive the autonomous driving vehicle based on a first driving action in the driving action sequence;

obtain an actual driving environment state after the first driving action is executed; and

update the stochastic model of driving environments based on the current driving environment state, the first driving action, and the actual driving environment state, wherein

the access count of each node is determined based on access counts of subnodes of the each node and an initial access count of the each node, the value function of each node is determined based on value functions of subnodes of the each node and an initial value function of the each node, the initial access count of the each node is 1, and the initial value function of the each node is determined based on a value function that matches the corresponding driving environment state represented by the each node.

10. The apparatus according to claim 9 , wherein the programming instructions are for execution by the at least one processor to:

predict, through dropout-based forward propagation by using the stochastic model of driving environments, a probability distribution of a driving environment state after the driving action is executed based on the second driving environment state represented by the parent node of the first node; and

obtain the first driving environment state represented by the first node through sampling from the probability distribution.

11. The apparatus according to claim 9 , wherein the programming instructions are for execution by the at least one processor to:

select, from an episodic memory, a first quantity of target driving environment states that have a highest matching degree with the driving environment state represented by the node; and determine the initial value function of the node based on value functions respectively corresponding to the first quantity of target driving environment states.

12. The apparatus according to claim 11 , wherein the programming instructions are for execution by the at least one processor to:

when a driving episode ends, determine a cumulative reward return value corresponding to an actual driving environment state after each driving action in the driving episode is executed; and

update the episodic memory by using, as a value function corresponding to the actual driving environment state, the cumulative reward return value corresponding to the actual driving environment state after each driving action is executed.

13. The apparatus according to claim 9 , wherein the node sequence is determined:

based on the access count of the each node in the Monte Carlo tree according to a maximum access count rule;

based on the value function of the each node in the Monte Carlo tree according to a maximum value function rule; or

based on the access count and the value function of the each node in the Monte Carlo tree according to a “maximum access count first, maximum value function next” rule.

14. A non-transitory computer-readable storage medium comprising a computer program or instructions which, when executed by a driving decision-making apparatus in an autonomous driving vehicle, cause the apparatus to perform operations comprising:

obtaining information of a current driving environment state of the autonomous driving vehicle, wherein the autonomous driving vehicle includes one or more sensors;

constructing a Monte Carlo tree based on the current driving environment state, wherein the Monte Carlo tree comprises N nodes, each node represents a corresponding driving environment state, the N nodes comprise a root node and N−1 non-root nodes, the root node represents the current driving environment state, a first driving environment state represented by a first node is predicted by using a stochastic model of driving environments based on a second driving environment state represented by a parent node of the first node and based on a driving action, the driving action is determined by the parent node of the first node in a process of obtaining the first node through expansion, the first node is any node of the N−1 non-root nodes, and N is a positive integer greater than or equal to 2;

determining, in the Monte Carlo tree based on at least one of an access count or a value function of each node in the Monte Carlo tree, a node sequence, wherein the node sequence comprises a plurality of nodes that starts from the root node and ends at a leaf node;

in response to determining the node sequence, determining a driving action sequence of a plurality of future driving steps, wherein each future driving step in the driving action sequence comprises a driving action corresponding to each node comprised in the node sequence, and wherein the driving action sequence is used for driving decision-making;

autonomously driving the autonomous driving vehicle based on a first driving action in the driving action sequence;

obtaining an actual driving environment state after the first driving action is executed; and

updating the stochastic model of driving environments based on the current driving environment state, the first driving action, and the actual driving environment state, wherein

the access count of each node is determined based on access counts of subnodes of the each node and an initial access count of the each node, the value function of the each node is determined based on value functions of subnodes of the each node and an initial value function of the each node, the initial access count of the each node is 1, and the initial value function of the each node is determined based on a value function that matches the corresponding driving environment state represented by the each node.

15. The non-transitory computer-readable storage medium according to claim 14 , wherein that the first driving environment state represented by the first node is predicted by using the stochastic model of driving environments based on the second driving environment state represented by the parent node of the first node and based on the first driving action comprises:

predicting, through dropout-based forward propagation by using the stochastic model of driving environments, a probability distribution of a driving environment state after the first driving action is executed based on the second driving environment state represented by the parent node of the first node; and

obtaining the first driving environment state represented by the first node through sampling from the probability distribution.

16. The non-transitory computer-readable storage medium according to claim 14 , wherein that the initial value function of the node is determined based on the value function that matches the driving environment state represented by the node comprises:

selecting, from an episodic memory, a first quantity of target driving environment states that have a highest matching degree with the driving environment state represented by the node; and

determining the initial value function of the node based on value functions respectively corresponding to the first quantity of target driving environment states.

17. The non-transitory computer-readable storage medium according to claim 16 , wherein the operations further comprise:

when a driving episode ends, determining a cumulative reward return value corresponding to an actual driving environment state after each driving action in the driving episode is executed; and

updating the episodic memory by using, as a value function corresponding to the actual driving environment state, the cumulative reward return value corresponding to the actual driving environment state after each driving action is executed.

18. The non-transitory computer-readable storage medium according to claim 14 , wherein the node sequence is determined:

based on the access count of the each node in the Monte Carlo tree according to a maximum access count rule;

based on the value function of the each node in the Monte Carlo tree according to a maximum value function rule; or

based on the access count and the value function of the each node in the Monte Carlo tree according to a “maximum access count first, maximum value function next” rule.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Apr 6, 2023
From: LI, DONG; WANG, BIN; LIU, WULONG; ZHUANG, YUZHENG
To: HUAWEI TECHNOLOGIES CO., LTD.
Reel/Frame 063242/0203 →
Priority Claims (1)
CN 202010584738.2 · Jun 23, 2020 · national
Continuity (2)
Continuation PCTCN2021090365 · Apr 27, 2021
Related Publication 20230162539A1 · May 25, 2023
References Cited (22)
US 20100138368A1 · Stundner et al. · 2010 [cited by applicant]
US 20130185039A1 · Tesauro et al. · 2013 [cited by applicant]
US 20180089563A1 · Redding · 2018 [cited by examiner]
US 20200150672A1 · Naghshvar et al. · 2020 [cited by applicant]
CN 108446727A · 2018 [cited by applicant]
CN 108803609A · 2018 [cited by applicant]
CN 108897928A · 2018 [cited by applicant]
CN 108985458A · 2018 [cited by applicant]
CN 106169188B · 2019 [cited by applicant]
CN 109598934A · 2019 [cited by applicant]
CN 109858553A · 2019 [cited by applicant]
Carl-Johan Hoel, et al., Combining Planning and Deep Reinforcement Learning in Tactical Decision Making for Autonomous Driving, Jun. 2020, IEEE Transactions on intelligent vehicles, vol. 5, No. 2 pp. 294-305 (Year: 2020… [cited by examiner]
Hoel et al., “Combining Planning and Deep Reinforcement Learning in Tactical Decision Making for Autonomous Driving,” IEEE Transactions on Intelligent Vehicles, May 6, 2019, 12 pages. [cited by applicant]
Cai et al., “HyP-DESPOT: A Hybrid Parallel Algorithm for Online Planning under Uncertainty,” CoRR, submitted on Feb. 17, 2018, arXiv:1802.06215v1, 9 pages. [cited by applicant]
Jiang et al., “Feedback-Based Tree Search for Reinforcement Learning,” CoRR, submitted on May 15, 2018, arXiv:1805.05935v1, 19 pages. [cited by applicant]
Xu et al., “Cooperative Driving at Unsignalized Intersections Using Tree Search,” IEEE Transactions on Intelligent Tranportation Systems, vol. 21, No. 11, Nov. 2020, 9 pages. [cited by applicant]
Extended European Search Report in European Appln No. 21829106.0, dated Oct. 16, 2023, 10 pages. [cited by applicant]
Office Action in Japanese Appln. No. 2022-578790, mailed on Apr. 2, 2024, 4 pages (with English translation). [cited by applicant]
Sunberg et al., “The Value of Inferring the Internal State of Traffic Participants for Autonomous Freeway Driving,” Paper, Presented at Proceedings of the 2017 American Control Conference, May 24-26, 2017, 7 pages. [cited by applicant]
Peeta et al., “Adaptability of a Hybrid Route Choice Model to Incorporating Driver Behavior Dynamics Under Information Provision,” IEEE Transactions on Systems, Man, and Cybernetics—Part A: Systems and Humans, Mar. 2004… [cited by applicant]
Schulz et al., “Interaction-Aware Probabilistic Behavior Prediction in Urban Environments,” Paper, Presented at Proceedings of the 2018 IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS), Oct. 1-… [cited by applicant]
Office Action in Indian AppIn. No. 2022-17075925, mailed on Jun. 6, 2025, 10 pages. [cited by applicant]