IP Library Granted Patent US 12,596,338
Granted Patent B2
US 12,596,338 · App. 18/020,682 · Granted Apr 7, 2026

Method and apparatus for performing optimal control

Inventors: Sungho Joo (Seoul, KR); Jeonghoon Lee (Sejong-Si, KR); Joongjae Kim (Daejeon, KR); Je Yeol Lee (Seoul, KR); Dongmin Lee (Seoul, KR); Minseop Kim (Seoul, KR)
Assignees: MakinaRocks Co., Ltd.; Hanon Systems
G05B13/047G05B13/0265G05B13/041
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,596,338
App. No.
18/020,682
Granted
Apr 7, 2026
Kind
B2
Abstract

According to an exemplary embodiment of the present disclosure, an optimal control method performed by a computing device including at least one processor is disclosed. The method includes acquiring state information including at least one state variable; calculating first control information by inputting the state information to a reinforcement learning control model; calculating second control information from the state information based on a feedback control algorithm; and calculating optimal control information based on the first control information and the second control information.

Claims (47)

1 . An optimal control method performed by a computing device including at least one processor, the optimal control method comprising:

acquiring state information, wherein the state information includes at least one state variable selected from subcool data, heater outlet temperature data, compressor output data, and valve open/close data, the at least one state variable being measured from a mobile heat-pump HVAC loop comprising a compressor, a valve, and a heater mounted in a vehicle;

calculating first control information by inputting the state information to a reinforcement learning control model;

calculating second control information from the state information based on a feedback control algorithm;

calculating optimal control information based on the first control information and the second control information; and

transmitting the optimal control information to one or more physical actuators to drive, in real time, at least one of the compressor, the valve, and the heater.

2 . The optimal control method according to claim 1 , wherein the state information is acquired according to a previously determined time interval during a process of driving the vehicle.

3 . The optimal control method according to claim 1 , wherein the at least one state variable is selected from state variables about temperature data, subcool data, compressor output data, valve open/close data, and heater output data.

4 . The optimal control method according to claim 1 , wherein the reinforcement learning control model is trained based on a learning method including:

acquiring current state information including at least one state variable from a simulation model;

inputting the current state information to the reinforcement learning control model;

calculating first control information based on the reinforcement learning control model;

inputting the calculated first control information to the simulation model; and

acquiring next state information and reward information from the simulation model.

5 . The optimal control method according to claim 4 , further comprising:

inputting the current state information and the first control information to a dynamic model that is distinct from the simulation model;

obtaining predicted state information from the dynamic model; and

updating the dynamic model by back-propagating an error between the predicted state information and the next state information returned by the simulation model.

6 . The optimal control method according to claim 1 , wherein the feedback control algorithm includes at least one of a proportional P control algorithm, a proportional-integral PI control algorithm, or a proportional-integral-differential PID control algorithm.

7 . The optimal control method according to claim 1 , wherein the first control information or the second control information includes at least one control data among compressor output control data, valve open/close control data, or heater output control data.

8 . The optimal control method according to claim 1 , wherein the calculating of optimal control information is performed by determining one of the first control information or the second control information as optimal control information according to a set determining condition.

9 . The optimal control method according to claim 8 , wherein the determining condition is a condition based on an error of the first control information and the second control information and a previously determined error threshold.

10 . The optimal control method according to claim 1 , wherein the calculating of optimal control information is performed based on a confidence score calculated by the reinforcement learning control model.

11 . The optimal control method according to claim 1 , wherein the calculating of optimal control information is performed based on a result of a weighted sum operation based on the first control information and the second control information.

12 . The optimal control method according to claim 11 , wherein a weight for the weighted sum operation includes:

a first weight which corresponds to the first control information and is previously determined; and

a second weight which corresponds to the second control information and is previously determined.

13 . The optimal control method according to claim 11 , wherein the weight for the weighted sum operation is determined according to the confidence score calculated by the reinforcement learning control model.

14 . The optimal control method according to claim 11 , wherein the weight for the weighted sum operation is set such that the higher the confidence score calculated by the reinforcement learning control model, the higher the ratio obtained by dividing a first weight corresponding to the first control information by a second weight corresponding to the second control information.

15 . The optimal control method according to claim 1 , wherein the reinforcement learning control model is consistently trained based on two or more state information acquired according to a time interval previously determined during a process of driving a vehicle.

16 . A computer program stored in a non-transitory computer readable storage medium, wherein when the computer program is executed by one or more processors, the computer program causes the one or more processors to perform the following operations to perform optimal control and the operations include:

an operation of acquiring state information including at least one state variable from one or more sensors mounted on a vehicle energy system;

an operation of calculating first control information by inputting the state information to a reinforcement learning control model;

an operation of calculating second control information from the state information based on a feedback control algorithm;

an operation of calculating optimal control information by performing a weighted sum of the first control information and the second control information, wherein a confidence score is calculated by the reinforcement learning control model, and a weight for the first control information is set to increase as the confidence score increases, and a weight for the second control information is set to decrease as the confidence score increases, such that as the confidence score increases, a contribution of the first control information relative to the second control information in the weighted sum increases; and

an operation of transmitting the optimal control information to one or more physical actuators to drive, in real time, at least one of a compressor, a valve, and a heater of the vehicle energy system.

17 . An apparatus for performing optimal control, comprising:

one or more processors;

a memory; and

a network unit, wherein the one or more processors are configured to:

acquire state information including at least one state variable from one or more sensors mounted on a vehicle energy system, the state information comprising real-time measurement values related to the operation of the vehicle energy system;

calculate first control information by inputting the state information to a trained reinforcement learning control model implemented as a neural network executed by the processor;

calculate second control information from the state information based on a feedback control algorithm executed by the processor, the feedback control algorithm comprising at least one of a proportional (P), proportional-integral (PI), and proportional-integral-differential (PID) control algorithm;

calculate an error between the first control information and the second control information;

compare the error to a previously determined error threshold;

select, as optimal control information, one of the first control information or the second control information according to whether the error exceeds the error threshold, such that if the error is less than or equal to the error threshold, the first control information is selected as the optimal control information, and if the error is greater than the error threshold, the second control information is selected as the optimal control information; and

transmit the optimal control information to one or more physical actuators to drive, in real time, at least one of a compressor, a valve, and a heater of the vehicle energy system, thereby manipulating the physical state of the vehicle energy system.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Feb 27, 2023
From: JOO, SUNGHO; LEE, JEONGHOON; KIM, JOONGJAE; LEE, JE YEOL; LEE, DONGMIN; KIM, MINSEOP
To: MAKINAROCKS CO., LTD.; HANON SYSTEMS
Reel/Frame 062813/0193 →
Priority Claims (2)
KR 10-2021-0078419 · Jun 17, 2021 · national
KR 10-2022-0051274 · Jun 17, 2021 · national
Continuity (1)
Related Publication 20240241486A1 · Jul 18, 2024
References Cited (22)
US 20030055798A1 · Hittle · 2003 [cited by examiner]
US 20050075993A1 · Jang et al. · 2005 [cited by applicant]
US 20190041811A1 · Drees · 2019 [cited by examiner]
US 20190187631A1 · Badgwell et al. · 2019 [cited by applicant]
US 20190309979A1 · Hurley et al. · 2019 [cited by applicant]
US 20200050920A1 · Idgunji et al. · 2020 [cited by applicant]
US 20200167437A1 · Mallya Kasaragod · 2020 [cited by examiner]
US 20200167611A1 · Yoon et al. · 2020 [cited by applicant]
US 20200249637A1 · Wee et al. · 2020 [cited by applicant]
US 20210026334A1 · Mazur · 2021 [cited by examiner]
US 20210132552A1 · Lawrence · 2021 [cited by examiner]
US 20210132587A1 · Lawrence · 2021 [cited by examiner]
US 20220204018A1 · Suplin · 2022 [cited by examiner]
KR 1020200062887A · 2020 [cited by applicant]
KR 1020200145079A · 2020 [cited by applicant]
KR 1020210006874A · 2021 [cited by applicant]
KR 1020210044177A · 2021 [cited by applicant]
KR 102247165B1 · 2021 [cited by applicant]
KR 1020210068378A · 2021 [cited by applicant]
Z. Cao, S. Xu, H. Peng, D. Yang and R. Zidek, “Confidence-Aware Reinforcement Learning for Self-Driving Cars,” in IEEE Transactions on Intelligent Transportation Systems, vol. 23, No. 7, pp. 7419-7430, Jul. 2022, date o… [cited by examiner]
P.E. An, S., et al. “A Reinforcement Learning Approach to On-Line Optimal Control”, Advanced Systems Research Group, Department of Aeronautics and Astronautics, 1994, (pp. 2465-2471). [cited by applicant]
Panagiotis Kofinas et al., “Online Tuning of a PID Controller with a Fuzzy Reinforcement Learning MAS for Flow Rate Control of a Desalination Unit”, MDPI, Dec. 10, 2018 (18 pgs). [cited by applicant]