IP Library › Granted Patent US 12,283,181
Granted Patent B2
US 12,283,181 · App. 17/870,138 · Granted Apr 22, 2025

Apparatus and method for controlling traffic signals of traffic lights in sub-area by using reinforcement learning model

Inventors: Jin Won Yoon (Iksan-si, KR); Seung Eon Baek (Incheon, KR); Seong Jin Lee (Seoul, KR)
Assignee: NOTA, INC.
G08G1/0125G06N3/08G08G1/08G08G1/083
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,283,181
App. No.
17/870,138
Granted
Apr 22, 2025
Kind
B2
Abstract

Provided are methods and apparatuses for controlling traffic signals of traffic lights in a sub-area by using a neural network model. The method according to an embodiment of the present disclosure may configure state information of a sub-area by using downstream information obtained in a current cycle time for each of a plurality of intersections included in the sub-area. In addition, the method may input the state information to a trained reinforcement learning model, and obtain action information of the sub-area including green times and offsets, by using an output from the trained reinforcement learning model. Furthermore, the method may generate coordinated signal values for applying the action to traffic lights in the sub-area in a subsequent cycle time.

Claims (46)

1. A method of controlling traffic signals of a plurality of traffic lights in a sub-area by using a neural network model, the method comprising:

configuring state information of the sub-area by using downstream information for a current cycle time, wherein the downstream information is configured for each of a plurality of intersections included in the sub-area;

obtaining action information including green times and offsets for the sub-area by inputting the state information to a trained neural network model;

determining whether an offset is set to within a preset absolute value range; and

generating a coordinated signal value for applying the action information to the plurality of traffic lights in the sub-area during a transition process configured with a plurality of subsequent cycle times, in response to determining that the offset is set to a value out of the preset absolute value range,

wherein the neural network model is trained using a reinforcement learning algorithm which is based on the action information, the state information, and reward information, and

wherein the reward information for each intersection is defined as an arithmetic mean of stop rate obtained at downstream of each of a plurality of links of intersection, each of the stop rates is a value obtained by diving a processed queue length by a processed traffic volume of the downstream of one of the plurality of links of the intersection.

2. The method of claim 1 , wherein

the downstream information includes the processed traffic volume and the processed queue length,

the processed traffic volume is calculated based on a traffic volume defined as a number of vehicles passing a particular point per hour and a maximum traffic volume not related to a geometric structure of a road, and

the processed queue length is calculated based on a number of waiting vehicles and a length of a downstream.

3. The method of claim 1 , further comprising

configuring state information of the sub-area by using downstream information obtained in a subsequent cycle time after the transition process, in response to determining that the offset is set to a value out of the preset absolute value range.

4. The method of claim 1 , wherein the coordinated signal value is generated to gradually apply the action information to the plurality of traffic lights in the sub-area during the transition process configured with the plurality of subsequent cycle times, in response to determining that the offset is set to a value out of the preset absolute value range.

5. The method of claim 1 , wherein a number of subsequent cycle times in the transition process is determined based on an extent to which the offset exceeds the preset absolute value range.

6. The method of claim 1 , further comprising

generating a coordinated signal value for applying the action information to the plurality of traffic lights in the sub-area in an earliest cycle time among the plurality of subsequent cycle times, in response to determining that the offset is set to a value within the preset absolute value range.

7. The method of claim 1 , wherein

the reinforcement learning algorithm comprises a deep deterministic policy gradient (DDPG) algorithm, and

the neural network model is trained to output the action information, wherein the green times and the offsets minimize a total number of stops of vehicles passing through the sub-area.

8. The method of claim 1 , wherein

a green time is set to a value between a minimum green time and a maximum green time, and

the offset is set to a value within an absolute value range of a fixed cycle time.

9. The method of claim 1 , wherein

the reward information is defined by using a stop rate representing a ratio of vehicles that have experienced a stop among vehicles passing through during a specific cycle time.

10. The method of claim 1 , wherein

the neural network model is an actor network, and is trained based on reinforcement learning including a critic network,

the actor network is trained by using the state information as input data and the action information as output data, to determine an optimal policy for controlling the traffic signals of the plurality of traffic lights in the sub-area, and

the critic network is trained by using, as input data, action information output from the actor network and using, as output data, the reward information for evaluating adequacy of the action information output from the actor network.

11. An apparatus for controlling traffic signals of a plurality of traffic lights in a sub-area by using a neural network model, the apparatus comprising:

a memory storing at least one program; and

at least one processor configured to drive the neural network model by executing the at least one program,

wherein the at least one processor is further configured to

configure state information of the sub-area by using downstream information for each of a plurality of intersections included in the sub-area, the downstream information being obtained in a current cycle time,

input the state information into trained neural network model and obtain an action of the sub-area by using an output of the trained neural network model, the action including green times and offsets,

determine whether an offset is set to within a preset absolute value range; and

generate a coordinated signal value for applying action information to the plurality of traffic lights in the sub-area during a transition process configured with a plurality of subsequent cycle times, in response to determining that the offset is set to a value out of the preset absolute value range,

wherein the neural network model is trained using a reinforcement learning algorithm which is based on the action information, the state information, and reward information, and

wherein the reward information for each intersection is defined as arithmetic mean of stop rates obtained at downstream of each of a plurality of links of the intersection, each of the stop rates is a value obtained by dividing a processed queue length by a processed traffic volume of the downstream of one of the plurality links of the intersection.

12. A non-transitory computer-readable recording medium that stores a program that, when executed by a computer, configures the computer to:

configure state information of a sub-area by using downstream information for a current cycle time, wherein the downstream information is configured for each of a plurality of intersections included in the sub-area,

obtain action information including green times and offsets for the sub-area by inputting the state information to a trained neural network model,

determine whether an offset is within a preset absolute value range; and

generate a coordinated signal value for applying the action information to a plurality of traffic lights in the sub-area during a transition process configured with a plurality of subsequent cycle times, in response to determining that the offset is set to a value out of the preset absolute value range,

wherein the neural network model is trained using a reinforcement learning algorithm which is based on the action information, the state information, and reward information, and

wherein the reward information for each intersection is defined as arithmetic mean of stop rates obtained at downstream of each of a plurality of links of the intersection, each of the stop rates is a value obtained by dividing a processed queue length by a processed traffic volume of the downstream of one of the plurality links of the intersection.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jul 21, 2022
From: YOON, JIN WON; BAEK, SEUNG EON; LEE, SEONG JIN
To: NOTA, INC.
Reel/Frame 060581/0680 →
Priority Claims (1)
KR 10-2022-0084607 · Jul 8, 2022 · national
Continuity (1)
Related Publication 20240013654A1 · Jan 11, 2024
References Cited (52)
US 4167784A · McReynolds · 1979 [cited by examiner]
US 5257194A · Sakita · 1993 [cited by examiner]
US 5357436A · Chiu · 1994 [cited by examiner]
US 6188778B1 · Higashikubo · 2001 [cited by examiner]
US 8298493B2 · Holl · 2012 [cited by examiner]
US 12125380B2 · Eldessouki · 2024 [cited by examiner]
US 20110175753A1 · Free · 2011 [cited by examiner]
US 20130099942A1 · Mantalvanos · 2013 [cited by examiner]
US 20140139358A1 · Lee · 2014 [cited by examiner]
US 20140210646A1 · Subramanya · 2014 [cited by examiner]
US 20150066340A1 · Nelson · 2015 [cited by examiner]
US 20150120175A1 · Vahidi · 2015 [cited by examiner]
US 20160027300A1 · Raamot · 2016 [cited by examiner]
US 20180286228A1 · Xu · 2018 [cited by examiner]
US 20190385447A1 · Marecek · 2019 [cited by examiner]
US 20200035096A1 · Sun · 2020 [cited by examiner]
US 20200042799A1 · Huang · 2020 [cited by examiner]
US 20200096977A1 · Itou · 2020 [cited by examiner]
US 20200139989A1 · Xu · 2020 [cited by examiner]
US 20200143206A1 · Kartal · 2020 [cited by examiner]
US 20200192390A1 · Luo · 2020 [cited by examiner]
US 20200296741A1 · Ayala Romero · 2020 [cited by examiner]
US 20200363814A1 · He · 2020 [cited by examiner]
US 20210125076A1 · Zhang · 2021 [cited by examiner]
US 20210158690A1 · Lau · 2021 [cited by examiner]
US 20210174672A1 · Sakakibara · 2021 [cited by examiner]
US 20210276594A1 · Oh · 2021 [cited by examiner]
US 20210287534A1 · Savla · 2021 [cited by examiner]
US 20210348932A1 · Friedman · 2021 [cited by examiner]
US 20210397949A1 · Wang · 2021 [cited by examiner]
US 20220076569A1 · Choi · 2022 [cited by examiner]
US 20220076571A1 · Choi · 2022 [cited by examiner]
US 20220138568A1 · Smolyanskiy · 2022 [cited by examiner]
US 20220198925A1 · Mohamad Alizadeh Shabestary · 2022 [cited by examiner]
US 20220270480A1 · Lee · 2022 [cited by examiner]
US 20220327925A1 · Jackson · 2022 [cited by examiner]
US 20220398921A1 · Jaggi · 2022 [cited by examiner]
US 20230134029A1 · Ogata · 2023 [cited by examiner]
US 20230222905A1 · Xu · 2023 [cited by examiner]
US 20230351886A1 · Lee · 2023 [cited by examiner]
KR 102029656B1 · 2019 [cited by applicant]
KR 102185101B1 · 2020 [cited by applicant]
KR 1020210122181A · 2021 [cited by applicant]
KR 102356738B1 · 2022 [cited by applicant]
KR 102385575B1 · 2022 [cited by applicant]
KR 102400833B1 · 2022 [cited by applicant]
Mingyu PI, et al., “Reinforcement Learning-based Traffic Signal Control under Real-World Constraints”, Journal of KIISE, Aug. 2021, vol. 48, No. 8, pp. 871-877 (7 pages) https://doi.org/10.5626/JOK.2021.48.8.871. [cited by applicant]
S.Y. Jang, et al., “Research Trends on Deep Reinforcement Learning”, Electronics and Telecommunications Trends, 2019, DOI: https://doi.org/10.22648/ETRI.2019.J.340401, pp. 1-14 (14 Pages). [cited by applicant]
Korean Office Action dated Dec. 5, 2022 in Application No. 10-2022-0084607. [cited by applicant]
Korean Office Action dated Jul. 20, 2023 in Application No. 10-2022-0084607. [cited by applicant]
Korean Office Action dated Jan. 17, 2023 in Application No. 10-2022-0084607. [cited by applicant]
Korean Office Action dated May 17, 2023 in Application No. 10-2022-0084607. [cited by applicant]