IP Library Granted Patent US 11,886,153
Granted Patent B2
US 11,886,153 · App. 17/383,213 · Granted Jan 30, 2024

Building control system using reinforcement learning

Inventors: Sugumar Murugesan (Santa Clara, CA); Young M. Lee (Old Westbury, NY); Viswanath Ramamurti (San Leandro, CA)
Assignee: JOHNSON CONTROLS TYCO IP HOLDINGS LLP
G05B13/048G05B13/027G05B13/041G06N3/045G06N3/08H02J3/003
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,886,153
App. No.
17/383,213
Granted
Jan 30, 2024
Kind
B2
Abstract

A method of operating a building management system is disclosed. The method includes determining, by a processing circuit, policy rankings for a plurality of control policies based on building operation data of a first previous time period, selecting, by the processing circuit, a set of control policies from among the plurality of control policies based on the policy rankings of the set of control policies satisfying a ranking threshold, generating, by the processing circuit, a plurality of prediction models for the set of control policies, selecting, by the processing circuit, a first prediction model of the plurality of prediction models based on building operation data of a second previous time period, and responsive to selecting the first prediction model, operating, by the processing circuit, the building management system using the first prediction model.

Claims (40)

1. A building management system, comprising:

one or more memory devices configured to store instructions thereon that, when executed by one or more processors, cause the one or more processors to:

determine policy rankings for a plurality of control policies based on building operation data of a first previous time period;

select a set of control policies from among the plurality of control policies based on the policy rankings of the set of control policies satisfying a ranking threshold;

generate a plurality of prediction models for the set of control policies;

select a first prediction model of the plurality of prediction models based on building operation data of a second previous time period; and

responsive to selecting the first prediction model, operate the building management system using the first prediction model.

2. The building management system of claim 1 , wherein the plurality of prediction models includes reinforcement learning models, and wherein the plurality of control policies includes one or more state-spaces for the reinforcement learning models.

3. The building management system of claim 2 , wherein the one or more state-spaces includes one or more of outside air temperature, occupancy, zone-air temperature, or cloudiness.

4. The building management system of claim 2 , wherein the state-spaces include time-step forecasts for environmental variables.

5. The building management system of claim 2 , wherein the plurality of control policies includes one or more action spaces, and wherein the one or more action-spaces includes one or more of a potential zone-air temperature setpoints.

6. The building management system of claim 1 , wherein the policy rankings are determined based on a total amount of energy consumption over the first previous time period.

7. The building management system of claim 1 , wherein the policy rankings are determined based on a total amount of money spent on operating building equipment over the first previous time period.

8. The building management system of claim 1 , wherein the selecting the building operation data is based on one or both of simulated data or real-time building operation data.

9. A method for generating a prediction model for control of a building management system, comprising:

determining, by a processing circuit, policy rankings for a plurality of control policies based on building operation data of a first previous time period;

selecting, by the processing circuit, a set of control policies from among the plurality of control policies based on the policy rankings of the set of control policies satisfying a ranking threshold;

generating, by the processing circuit, a plurality of prediction models for the set of control policies;

selecting, by the processing circuit, a first prediction model of the plurality of prediction models based on building operation data of a second previous time period; and

responsive to selecting the first prediction model, operating, by the processing circuit, the building management system using the first prediction model.

10. The method of claim 9 , wherein the plurality of prediction models includes reinforcement learning models,

wherein the plurality of control policies includes one or more state-spaces for the reinforcement learning models, and

wherein the policy rankings is determined based on a combination of state-space complexity, action-space complexity, and/or reward performance.

11. The method of claim 10 , wherein a higher number of parameters used to determine the one or more state spaces corresponds to a greater state-space complexity, and wherein a higher number of actions corresponds to a greater action-space complexity.

12. The method of claim 10 , wherein a higher reward performance is based on one or more of energy consumption, total comfortability, or total number of errors.

13. The method of claim 9 , wherein the plurality of control policies include known baseline policies.

14. The method of claim 9 , wherein the building operation data includes one or more of current values, previous values, or forecasts for various points of a space of a building operatively connected to the processing circuit.

15. The method of claim 14 , wherein the building operation data includes total energy consumption or average comfortability.

16. A building management system, comprising:

building equipment operatively connected to one or more processors; and

one or more memory devices configured to store instructions thereon that, when executed by the one or more processors, cause the one or more processors to:

determine policy rankings for a plurality of control policies based on building operation data of a first previous time period;

select a set of control policies from among the plurality of control policies based on the policy rankings of the set of control policies satisfying a ranking thresh old;

generate a plurality of prediction models for the set of control policies;

select a first prediction model of the plurality of prediction models based on building operation data of a second previous time period; and

responsive to selecting the first prediction model, operate the building equipment using the first prediction model.

17. The building management system of claim 16 , wherein the plurality of prediction models includes reinforcement learning models, and wherein the plurality of control policies includes one or more state-spaces for the reinforcement learning models.

18. The building management system of claim 17 , wherein the one or more state-spaces includes one or more of outside air temperature, occupancy, zone-air temperature, or cloudiness.

19. The building management system of claim 17 , wherein the state-spaces include time-step forecasts for environmental variables.

20. The building management system of claim 17 , wherein the plurality of control policies includes one or more action spaces, and wherein the one or more action spaces includes one or more of a potential zone-air temperature setpoints.

Assignments (2)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Apr 9, 2024
From: JOHNSON CONTROLS TYCO IP HOLDINGS LLP
To: TYCO FIRE & SECURITY GMBH
Reel/Frame 067056/0552 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Aug 24, 2021
From: MURUGESAN, SUGUMAR; LEE, YOUNG M.; RAMAMURTI, VISWANATH
To: JOHNSON CONTROLS TYCO IP HOLDINGS LLP
Reel/Frame 057276/0563 →