IP Library › Granted Patent US 12,748,389
Granted Patent B2
US 12,748,389 · App. 17/979,787 · Granted Sep 29, 2026

Method for predicting benchmark value of unit equipment based on XGBoost algorithm and system thereof

Inventors: Yongkang Wang (Shanghai, CN); Gang Xu (Shanghai, CN); Ruijie Chen (Shanghai, CN); Chen Wang (Shanghai, CN); Qingping Li (Shanghai, CN); Bin Wu (Shanghai, CN); Yi Gong (Shanghai, CN)
Assignee: HUANENG SHANGHAI COMBINED CYCLE POWER CO. LTD.
G05B13/0265G05B13/042
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,748,389
App. No.
17/979,787
Granted
Sep 29, 2026
Kind
B2
Abstract

The invention relates to a method for predicting benchmark value of unit equipment based on XGBoost algorithm and a system thereof, wherein the method comprises the following steps: the historical operation data of unit equipment is obtained, the data is preprocessed, and a data set containing a plurality of samples is constructed, and each sample includes the benchmark value of a plurality of parameters of the equipment corresponding to a plurality of features; RF out-of-bag estimation is used for feature importance calculation to eliminate the features with low importance; the data is standardized to eliminate the dimensional effects among features; the data set is input to construct an XGBoost model, and Bayesian super parameter optimization is conducted to obtain the prediction model of benchmark values; and the real-time data of equipment operation is input, and the benchmark values of various equipment parameters are predicted by the prediction model of benchmark values. Compared with the prior art, the invention mines the correlation among data based on the XGBoost algorithm to predict a reasonable equipment benchmark value, and has the advantages of high generalization ability, high prediction accuracy and operation speed and great improvement of the automation ability of the unit.

Claims (231)

1 . A method for predicting benchmark value of unit equipment based on XGBoost algorithm is characterized by comprising the following steps:

S 1 . The historical operation data of unit equipment is obtained, the data is preprocessed, and a data set containing a plurality of samples is constructed, and each sample includes the benchmark value of a plurality of parameters of the equipment corresponding to a plurality of features;

S 2 . random forest (RF) out-of-bag estimation is used for feature importance calculation to eliminate the features with low importance;

S 3 . The data is standardized to eliminate the dimensional effects among features;

S 4 . The data set is input to construct an XGBoost model, and Bayesian super parameter optimization is conducted to obtain the prediction model of benchmark values;

S 5 . The real-time data of equipment operation is input, and the benchmark values of various equipment parameters are predicted by the prediction model of benchmark values,

S 6 . Adjust the operating parameters of the equipment to the benchmark values obtained through S 5 ;

wherein step S 2 is as follows:

for each feature of the sample, the random forest (RF) out-of-bag estimation is used to rank the importance of the features and select the features, wherein the average precision decline rate (MDA) is used as an indicator to calculate the importance of the feature, and wherein the formula is as follows:

MDA

=

1

n

⁢

∑

t

=

1

n

⁢

(

e

⁢

r

⁢

r

⁢

O

⁢

O

⁢

B

t

-

errOOB

t

′

)

,

wherein, n is the number of base classifiers constructed by random forests, errOOB t is the out-of-bag error of the t th base classifier, and errOOB′ t is the out-of-bag error of the t th base classifier after noise is added, wherein the more MDA decreases, the higher the importance of the feature;

in step S 3 , the data set contains N samples, each sample has L-type features, and Z-score standardization method is used to standardize each type of features of each sample, as follows:

x

nl

*

=

x

nl

-

μ

l

σ

l

wherein, x nl is the feature data of the type 1 features of the n th sample, and

x

nl

*

 is the feature data of the type 1 features of the n th sample after standardization, μ l is the mean value of the feature data of the type 1 features in the N th sample, and σ l is the standard deviation of the feature data of the type 1 features in the N th sample.

2 . The method for predicting benchmark value of unit equipment based on XGBoost algorithm according to claim 1 , wherein S 11 -S 14 is under S 1 , step S 1 is as follows:

S 11 . The historical operation data of the equipment is obtained from the plant level supervisory information system SIS of the unit;

S 12 . The data is checked for blank values and outliers, and the data with blank values and outliers are eliminated;

S 13 . Straightened line type data is filtered;

S 14 . Data features are dimensionally reduced by PCA to obtain a data set containing multiple samples, and each sample contains multiple features.

3 . The method for predicting benchmark value of unit equipment based on XGBoost algorithm according to claim 1 is characterized in that step S 4 comprises the following steps, wherein S 41 - 45 is under S 4 :

S 41 . The data set T containing N samples is input, T={(X 1 , Y 1 ), (X 2 , Y 2 ), (X 3 , Y 3 ), . . . , (X N , Y N )}, each sample has L-type features, X i =(x i1 , x i2 , . . . , x iL ), corresponding to the benchmark value of M parameters of the equipment, Y i =(y i1 , y i2 , . . . , y iM );

S 42 . The objective function of XGBoost model iteration is established:

O

⁡

(

t

)

=

-

1

2

⁢

∑

k

=

1

K

G

k

2

H

k

+

λ

+

γ

⁢

K

Wherein,

G

k

=

∑

i

∈

I

k

⁢

∂

Y

^

(

t

-

1

)

l

⁡

(

Y

i

,

Y

ˆ

i

(

t

-

1

)

)

;

H

k

=

∑

i

∈

I

k

⁢

∂

Y

^

(

t

-

1

)

2

l

⁡

(

Y

i

,

Y

ˆ

i

(

t

-

1

)

)

;

 λ is L2 regular penalty coefficient; γ is L 1 regular penalty coefficient; K is the total number of leaf nodes in the decision tree; Y i is the true value of the i th sample;

Y

^

i

(

t

-

1

)

 is the predicted value after the (t−1) th iteration of the i th sample; and the sample set on the leaf with index k is defined as as I k ;

S 43 . The adjustment range of XGBoost model super parameters is set, and Bayesian optimization algorithm is used to optimize XGBoost super parameters to obtain the optimal combination of super parameters;

S 44 . The optimal combination of super parameters is input into the XGBoost model, and the data set T is used to train according to the objective function 0 (t);

S 45 . The optimal combination of the super parameters is recorded if the prediction performance of the XGBoost model obtained through training meets the preset accuracy threshold, so as to obtain the prediction model of benchmark values; otherwise, step S 43 is executed to optimize the XGBoost super parameters again.

4 . The method for predicting benchmark value of unit equipment based on XGBoost algorithm according to claim 1 is characterized in that in step S 43 , the XGBoost model super parameters include: Learning rate with the parameter adjustment range of [0.1, 0.15]; Maximum depth of the tree with the parameter adjustment range of (5, 30); Penalty term of complexity with the parameter adjustment range of (0, 30); Randomly selected sample proportion with the parameter adjustment range of (0, 1); Random sampling ratio of features with the parameter adjustment range of (0.2, 0.6); L2 norm regular term of weight with the parameter adjustment range of (0, 10); Number of decision trees with the parameter adjustment range of (500, 1000); Minimum leaf node weight sum with the parameter adjustment range of (0, 10).

5 . The method for predicting benchmark value of unit equipment based on XGBoost algorithm according to claim 1 is characterized in that the prediction performance of XGBoost model in step S 45 includes average absolute percentage error and determination coefficient and the calculation formula is as follows:

eMAPE

=

∑

i

=

1

⁢

N

❘

"\[LeftBracketingBar]"

Y

^

i

-

YiYi

❘

"\[RightBracketingBar]"

⁢

N

R

⁢

2

=

1

-

∑

i

=

1

⁢

N

(

Y

^

i

-

Yi

)

⁢

2

⁢

∑

i

=

1

⁢

N

(

Y

^

i

-

Y_i

)

⁢

2

Wherein, e MAPE is the average absolute percentage error, R 2 is the determination coefficient, Yi is the benchmark value of the i th sample in the data set, Ŷ 1 is the benchmark value predicted by the XGBoost model according to the feature X of the i th sample, and Ŷ i is the average value of the benchmark values of the N th sample in the data set.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Nov 3, 2022
From: WANG, YONGKANG; XU, GANG; CHEN, RUIJIE; WANG, CHEN; LI, QINGPING; WU, BIN; GONG, YI
To: HUANENG SHANGHAI COMBINED CYCLE POWER CO, LTD.
Reel/Frame 061638/0730 →
Priority Claims (1)
CN 202111681654.1 · Dec 30, 2021 · national
Continuity (1)
Related Publication 20230213895A1 · Jul 6, 2023
References Cited (36)
US 11720962B2 · Kamkar · 2023 [cited by examiner]
US 12093817B2 · Kuo · 2024 [cited by examiner]
US 12182670B1 · Beauchesne · 2024 [cited by examiner]
US 20130044944A1 · Wang · 2013 [cited by examiner]
US 20170228645A1 · Wang · 2017 [cited by examiner]
US 20170286810A1 · Shigenaka · 2017 [cited by examiner]
US 20190086912A1 · Hsu · 2019 [cited by examiner]
US 20200380037A1 · Zhao · 2020 [cited by examiner]
US 20210011771A1 · Bonderson · 2021 [cited by examiner]
US 20210026314A1 · Phan · 2021 [cited by examiner]
US 20210117787A1 · Stal · 2021 [cited by examiner]
US 20210117869A1 · Plumbley · 2021 [cited by examiner]
US 20210180439A1 · Rawlinson · 2021 [cited by examiner]
US 20210182357A1 · Partee · 2021 [cited by examiner]
US 20210192370A1 · Toubiana · 2021 [cited by examiner]
US 20210334914A1 · Sadot · 2021 [cited by examiner]
US 20210350203A1 · Das · 2021 [cited by examiner]
US 20210358564A1 · Feinberg · 2021 [cited by examiner]
US 20220008976A1 · Sekimoto · 2022 [cited by examiner]
US 20220067249A1 · Steingrimsson · 2022 [cited by examiner]
US 20220092473A1 · Lee · 2022 [cited by examiner]
US 20220108195A1 · Kehler · 2022 [cited by examiner]
US 20220128983A1 · Zhang · 2022 [cited by examiner]
US 20220188521A1 · Mu · 2022 [cited by examiner]
US 20220261520A1 · Shigemori · 2022 [cited by examiner]
US 20220283575A1 · Mehta · 2022 [cited by examiner]
US 20220285938A1 · Mehta · 2022 [cited by examiner]
US 20220318641A1 · Carreira-Perpiñán · 2022 [cited by examiner]
US 20220335736A1 · Tizhoosh · 2022 [cited by examiner]
US 20220374827A1 · Lee · 2022 [cited by examiner]
US 20220397874A1 · Peng · 2022 [cited by examiner]
US 20230074606A1 · Jesus · 2023 [cited by examiner]
US 20230167997A1 · Bistany · 2023 [cited by examiner]
US 20230185652A1 · Hauser · 2023 [cited by examiner]
US 20240244410A1 · Fei · 2024 [cited by examiner]
US 20240320524A1 · Shen · 2024 [cited by examiner]