IP Library › Granted Patent US 12,602,568
Granted Patent B2
US 12,602,568 · App. 18/331,450 · Granted Apr 14, 2026

Born-again TSK fuzzy classifier based on knowledge distillation

Inventors: Yunliang Jiang (Huzhou, CN); Xiongtao Zhang (Huzhou, CN); Jungang Lou (Huzhou, CN); Qing Shen (Huzhou, CN); Jiangwei Weng (Qingdao, CN)
Assignee: Huzhou University
G06N3/043G06N3/08
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,602,568
App. No.
18/331,450
Filed
Jun 8, 2023
Granted
Apr 14, 2026
Kind
B2
Art Unit
2875
USPC
706/15
Abstract

This application provides a born-again TSK fuzzy classifier based on knowledge distillation. The born-again TSK fuzzy classifier based on knowledge distillation is denoted as CNNBaTSK, and a fuzzy rule of CNNBaTSK includes two parts: an antecedent part based on soft label information and a consequent part based on original data. A method for constructing the fuzzy rule of CNNBaTSK includes following steps: Step 1: taking, by the CNNBaTSK, the original data as input, obtaining a probability distribution of an output layer through a layer-by-layer neural expression, and introducing a distillation temperature to generate soft label information of DATASET; Step 2: partitioning the soft label information into five fixed fuzzy partitions to construct the fuzzy rule in a fuzzy part of the CNNBaTSK; Step 3: introducing the original data to calculate a consequent parameter, and optimizing the consequent parameter of CNNBaTSK using a non-iterative learning method.

Claims (1913)

1 . A born-again TSK fuzzy classifier based on knowledge distillation, wherein the born-again TSK fuzzy classifier based on knowledge distillation is denoted as CNNBaTSK, and a fuzzy rule of the CNNBaTSK comprises two parts: an antecedent part based on soft label information and a consequent part based on original data, and a method for constructing the fuzzy rule of the CNNBaTSK comprises following steps:

Step 1: taking, by the CNNBaTSK, the original data as input, obtaining a probability distribution of an output layer through a layer-by-layer neural expression, and introducing a distillation temperature to generate soft label information of DATASET;

Step 2: partitioning the soft label information into five fixed fuzzy partitions to construct the fuzzy rule in a fuzzy part of the CNNBaTSK;

Step 3: introducing the original data to calculate a consequent parameter, and optimizing the consequent parameter of CNNBaTSK using a non-iterative learning method.

2 . The born-again TSK fuzzy classifier based on knowledge distillation according to claim 1 , wherein in Step 2, each soft label information in different fuzzy rules have a center of {0, 0.25, 0.5, 0.75, and 1} respectively, and the soft label information is transformed into semantic interpretation to construct the fuzzy rule.

3 . The born-again TSK fuzzy classifier based on knowledge distillation according to claim 1 , wherein in Step 1, a specific method of generating the soft label information of DATASET comprises: firstly stacking multiple convolutional layers and pooling layers, in terms of layer by layer learning and generating a depth feature, the convolutional layers and the pooling layers used being disposed alternately, and then performing classification through several fully-connected layers and the output layer; the convolutional layer, the pooling layer, and the fully-connected layer are respectively denoted as a Conv layer, a Pool layer, and a FC layer;

assuming that a training dataset is X={x i , x i =(x 1 , x 2 , . . . , x n ) T , x i ∈R n , i=1, 2, . . . , N} and a label set is {tilde over (Y)}={Y i , Y i ∈{0, 1, . . . , K}, i=1, 2, . . . , N}, where n represents a sample dimension, N represents a sample quantity, and K represents a sample category; in a convolutional neural network, an output result (feature map) of a layer t is marked as Z t , where Z 0 represents the original data x i ;

in the Conv layer, a local connection method is employed to execute a convolution calculation, and then a bias be is added to the feature map, and then an activation function ƒ(⋅) is employed for a nonlinear transformation; and a calculation process for the Conv layer is as follows:

Z

t

=

j

⁡

(

W

i

⋆

Z

t

-

1

+

b

t

)

(

1

)

after the convolution calculation is completed, a nonlinear mapping of the PRelu activation function is performed, wherein a mathematical expression of the PRelu is expressed as follows:

f

⁡

(

a

)

=

{

a

,

if

⁢

a

>

0

κ

⁢

a

,

otherwise

(

2

)

wherein α represents an input variable and κ represents a slope coefficient;

in the Pool layer, a maximum value in a pooling window is selected as a result in a selected maximum pooling operation, and a pooling process can be expressed as follows:

Z

t

=

Pool

(

Z

t

-

1

)

(

3

)

after several convolution and maximum pooling operations, an extracted depth feature is input into a first FC layer; in the FC layer, all neurons between layers are connected, and the depth feature is further mapped to a new feature space to complete a classification task; calculation is performed through a weight W t and the bias b t , and a nonlinear transformation is performed using the activation function ƒ(⋅); and a calculation process is as follows:

Z

t

=

f

⁡

(

W

t

·

Z

t

-

1

+

b

t

)

(

4

)

an output layer of the convolutional neural network uses Softmax activation function, and an output thereof Z t =(z 1 , z 2 , . . . , z K ) T is transformed into a corresponding probability result E i =(e 1 , e 2 , . . . , e K ) T , wherein K represents a total number of category, and its calculation process is as follows:

e

K

=

exp

⁡

(

z

K

)

/

∑

k

=

1

K

exp

⁡

(

z

k

)

(

5

)

during the training process, a cross-entropy loss function is employed to measure a difference between the output of the convolutional neural network and a ground-truth label, and its calculation formula is as follows:

min

-

1

N

⁢

∑

i

=

1

N

∑

k

=

1

K

Y

i

,

k

⁢

log

⁡

(

e

k

(

E

i

)

)

(

6

)

the weight W t and bias b t of the convolutional neural network are iterated and optimized through an error backpropagation algorithm; a loss value of formula (6) is backpropagated from a last layer to a first layer, and parameter updating is performed according to the error of each layer; assuming that a derivative of cross-entropy loss for the weight W t is ΔW t , and a derivative of cross-entropy loss for the bias b t is Δb t , formula for parameter updating are expressed as follows:

W

t

l

=

W

t

l

-

θΔ

⁢

W

t

l

-

1

(

7

)

b

t

l

=

b

t

l

-

θΔ

⁢

b

t

l

-

1

(

8

)

wherein l represents a training iteration epoch and θ represents a learning rate;

in order to facilitate the distinction, a probability distribution of the output layer softened by the distillation temperature T is called the soft label information s i =(s 1 , s 2 , . . . , s K ) T , and its calculation method is as follows:

s

K

=

exp

⁢

(

z

K

/

T

)

/

∑

k

=

1

K

exp

⁢

(

z

k

/

T

)

(

9

)

in a binary-class classification task, assuming that the training dataset is expressed as X={x i , x i =(x 1 , x 2 , . . . , x n ) T , x i ∈R n , i=1, 2, . . . , N}, {tilde over (Y)}={Y i , Y i , ∈{0,1}, i=1, 2, . . . , N} is a ground-truth label vector of the binary-class classification task; the probability distribution of the output layer E i =(e 1 , e 2 ) T is transformed into the soft label information S={s i =(s 1 , s 2 ) T , s k ∈[0,1], k=1,2, i=1, 2, . . . , N} after introducing the distillation temperature T, wherein the sum of all probabilities in each sample is 1, that is, s 1 +s 2 =1.

4 . The born-again TSK fuzzy classifier based on knowledge distillation according to claim 3 , wherein in Step 2, the CNNBaTSK adopts following fuzzy rules:

In

⁢

rule

⁢

m

:

IF

⁢

s

1

⁢

is

*

^

s

2

⁢

is

*

THEN

⁢

y

m

=

p

0

m

+

p

1

m

⁢

x

1

+

p

2

m

⁢

x

2

+

…

+

p

n

m

⁢

x

n

.

m

=

1

,

2

,

…

,

M

(

10

)

wherein * denotes the membership of fuzzy partition, M is the total number of fuzzy rules and ∧ denotes a fuzzy conjunction operator, and the consequent part adopts a linear function of the input data x i ;

the antecedent part of the soft label information is taken into five fixed fuzzy partitions, wherein each fuzzy partition has a center of

c

d

m

∈

{

0

,

0.25

,

0.5

,

0.75

,

1

}

,

and the kernel width

σ

d

m

is set at a random positive value which is between 0 and 1; then, a Gaussian membership function as a fuzzy membership function is applied to estimate a membership of each rule, and its calculation method is as follows:

μ

m

(

s

i

,

d

)

=

exp

⁡

(

-

(

s

i

,

d

-

c

d

m

)

2

/

2

⁢

σ

d

m

)

(

11

)

μ

m

(

s

i

)

=

∏

d

=

1

2

μ

m

(

s

i

,

d

)

⁢

μ

m

(

s

i

)

=

∏

d

=

1

2

μ

m

(

s

i

,

d

)

(

12

)

a normalized membership function is formulated as follows:

μ

~

m

(

s

i

)

=

μ

m

(

s

i

)

/

∑

m

′

=

1

M

μ

m

′

(

s

i

)

(

13

)

therefore, the output {tilde over (y)} i of the CNNBaTSK can be expressed as:

y

~

i

=

∑

m

=

1

M

μ

~

m

(

s

i

)

⁢

y

m

(

14

)

making:

p

m

=

(

p

0

m

,

p

1

m

,

…

,

p

n

m

)

T

(

15

)

P

=

(

(

p

l

)

T

,

(

p

2

)

T

,

…

,

(

p

m

)

T

)

T

(

16

)

q

i

m

=

μ

~

m

(

s

i

)

⁢

(

l

,

x

i

)

T

(

17

)

Q

l

=

(

(

q

i

l

)

T

,

(

q

i

2

)

T

,

…

,

(

q

i

m

)

T

)

T

(

18

)

the output {tilde over (y)} i of the CNNBaTSK may be expressed as:

y

~

i

=

Q

i

T

⁢

P

.

(

19

)

5 . The born-again TSK fuzzy classifier based on knowledge distillation according to claim 4 , wherein in Step 3, a specific method for optimizing the consequent parameter of the CNNBaTSK is as follows:

an objective function is designed to balance the ground-truth label and the soft label information, and a following objective function for optimization is proposed:

min

⁢

1

2

⁢

∑

N

i

=

1

ς

i

2

+

α

2

⁢

P

2

2

+

β

2

⁢

∑

i

=

1

N

ξ

t

2

(

20

)

s

.

t

.

{

Q

i

T

⁢

P

-

Y

i

=

ς

i

Q

i

T

⁢

P

-

s

i

,

2

=

ξ

i

,

i

=

1

,

2

,

…

,

N

wherein P is the consequent parameter of the fuzzy rule, α and β are regularization factors, ζ i is a deviation between the prediction output and the ground-truth label, ξ i is a deviation between the prediction output and the soft label information; based on s 1 +s 2 =1 and the ground-truth label, index of the probability s 2 is matched with the ground-truth label;

a Lagrangian optimization formula of the objective function can be expressed as:

L

⁡

(

P

,

ς

,

ξ

,

τ

,

v

)

=

1

2

⁢

ς

2

2

+

α

2

⁢

P

2

2

+

β

2

⁢

ξ

2

2

-

τ

⁡

(

Q

~

⁢

P

-

Y

~

-

ς

)

-

v

⁡

(

Q

~

⁢

P

-

S

~

-

ξ

)

(

21

)

wherein ζ=(ζ 1 , ζ 2 , . . . , ζ N ) T , ξ=(ξ 1 , ξ 2 , . . . , ξ N ) T , τ=(τ 1 , τ 2 , . . . , τ N ) T and ν=(ν 1 , ν 2 , . . . , ν N ) T are Lagrangian multipliers with equality constraints, and an input matrix is denoted by {tilde over (Q)}=(Q 1 , Q 2 , . . . , Q N ) T ;

the consequent parameter P is calculated by setting a derivative gradient of the Lagrangian optimization formula with respect to (P, ζ, ξ, τ, ν) equal to be zero, resulting in following KKT optimality conditions:

{

∂

L

∂

P

⇒

α

⁢

P

T

=

τ

⁢

Q

~

+

v

⁢

Q

~

a

⁢

L

∂

ς

⇒

ς

T

+

τ

=

0

∂

L

∂

ξ

⇒

β

⁢

ξ

T

+

v

=

0

∂

L

∂

τ

⇒

Q

~

⁢

P

-

Y

~

-

ς

=

0

∂

L

∂

v

⇒

Q

~

⁢

P

-

S

~

-

ξ

=

0

(

22

)

yielding:

{

τ

=

-

(

Q

~

⁢

P

-

Y

~

)

T

v

=

-

β

⁡

(

Q

~

⁢

P

-

S

~

)

T

α

⁢

P

T

=

-

(

Q

~

⁢

P

-

Y

~

)

T

⁢

Q

~

-

β

⁡

(

Q

~

⁢

P

-

S

~

)

T

⁢

Q

~

(

23

)

therefore, an optimal analytical solution can be expressed as:

P

=

(

α

⁢

I

+

(

1

+

β

)

⁢

Q

˜

T

⁢

Q

˜

)

-

1

⁢

(

Q

˜

T

⁢

Y

˜

+

β

⁢

Q

˜

T

⁢

S

˜

)

(

24

)

wherein I denotes an identity matrix.

6 . The born-again TSK fuzzy classifier based on knowledge distillation according to claim 1 , wherein a learning algorithm of CNNBaTSK has inputs comprising: training dataset X={x i , x i =(x 1 , x 2 , . . . , x n ) T , x i ∈R n , i=1, 2 . . . , N} and {tilde over (Y)}={Y i , Y i ∈{0, 1, . . . , K}, i=1, 2, . . . , N} and a maximum iteration epoch {tilde over (l)}, initial learning rate θ, distillation temperature T, regularization factors α and β, testing sample x test of the convolutional neural network; outputs comprising CNNBaTSK after completion of training, and the output {tilde over (y)} test of the testing sample; the learning algorithm of the CNNBaTSK comprises two stages as follows:

S1: training stage:

S10: initialization: initializing W t and b t with random numbers, and proceeding step S11 to step S13 sequentially within a range of the maximum iteration epoch {tilde over (l)};

S11: inputting training samples X={x i , x i =x 1 , x 2 , . . . , x n ) T , x i ∈R n , i=1, 2, . . . , N}, and performing a convolutional and pooling operation on the training samples through the Conv layer and the Pool layer to extract depth features;

S12: calculating a probability that the sample belongs to each label at the FC layer;

S13: performing error backpropagation on 1D-CNN according to a cross-entropy loss function formula;

S14: generating a probability distribution E i =(e 1 , e 2 ) T of the output layer after the iteration is completed;

S15: introducing the distillation temperature T and converting it into the soft label information S={s i =(s 1 ,s 2 ) T , s k ∈[0,1], k=1,2, i=1,2, . . . , N};

S16: partitioning the antecedent part into five fixed fuzzy partitions, wherein the center of each fuzzy partition is

c

d

m

∈

{

0

,

0.25

,

0.5

,

0.75

,

1

}

,

 and the kernel width

σ

d

m

 is set at a random positive value which is between 0 and 1;

S17: computing a normalized membership {tilde over (μ)} m (s i ) of the training sample in different rules;

S18: computing and generating Q i , wherein the output of the CNNBaTSK can be expressed as

y

~

i

=

Q

i

T

⁢

P

;

S19: computing an analytical solution for P to be: P=(αI+(1+β){tilde over (Q)} T {tilde over (Q)}) −1 ({tilde over (Q)} T {tilde over (Y)}+β{tilde over (Q)} T {tilde over (S)});

S2: testing stage:

S21: inputting a testing sample x test ;

S22: generating the probability distribution E test =(e 1 , e 2 ) T of the output layer;

S23: computing the probability distribution s test =(s 1 , s 2 ) T of the output layer of the soft label information;

S24: computing a normalized membership {tilde over (μ)} m (s i ) of the testing sample in different rules;

S25: outputting {tilde over (y)} test of the testing sample according to

y

~

test

=

Q

test

T

⁢

P

.

7 . A method for constructing a fuzzy rule for a born-again TSK fuzzy classifier based on knowledge distillation, wherein the born-again TSK fuzzy classifier based on knowledge distillation is denoted as CNNBaTSK, and the fuzzy rule of the CNNBaTSK comprises two parts: an antecedent part based on soft label information and a consequent part based on original data, and the method for constructing the fuzzy rule of the CNNBaTSK comprises following steps:

Step 1: taking, by the CNNBaTSK, the original data as input, obtaining a probability distribution of an output layer through a layer-by-layer neural expression, and introducing a distillation temperature to generate soft label information of DATASET;

Step 2: partitioning the soft label information into five fixed fuzzy partitions to construct the fuzzy rule in a fuzzy part of the CNNBaTSK;

Step 3: introducing the original data to calculate a consequent parameter, and optimizing the consequent parameter of CNNBaTSK using a non-iterative learning method.

8 . The method for constructing the fuzzy rule according to claim 7 , wherein in Step 2, each soft label information in different fuzzy rules have a center of {0, 0.25, 0.5, 0.75, and 1} respectively, and the soft label information is transformed into semantic interpretation to construct the fuzzy rule.

9 . The method for constructing the fuzzy rule according to claim 7 , wherein in Step 1, a specific method of generating the soft label information of DATASET comprises: firstly stacking multiple convolutional layers and pooling layers, in terms of layer by layer learning and generating a depth feature, the convolutional layers and the pooling layers used being disposed alternately, and then performing classification through several fully-connected layers and the output layer; the convolutional layer, the pooling layer, and the fully-connected layer are respectively denoted as a Conv layer, a Pool layer, and a FC layer;

assuming that a training dataset is X={x i , x i =(x 1 , x 2 , . . . , x n ) T , x i ∈R n , i=1, 2, . . . , N} and a label set is {tilde over (Y)}={Y i , Y i ∈{0, 1, . . . , K}, i=1, 2, . . . , N}, where n represents a sample dimension, N represents a sample quantity, and K represents a sample category; in a convolutional neural network, an output result (feature map) of a layer t is marked as Z t , where Z 0 represents the original data x i ;

in the Conv layer, a local connection method is employed to execute a convolution calculation, and then a bias be is added to the feature map, and then an activation function ƒ(⋅) is employed for a nonlinear transformation; and a calculation process for the Conv layer is as follows:

Z

t

=

f

⁡

(

W

t

*

Z

t

-

1

+

b

t

)

(

1

)

after the convolution calculation is completed, a nonlinear mapping of the PRelu activation function is performed, wherein a mathematical expression of the PRelu is expressed as follows:

f

⁡

(

a

)

=

{

a

,

if

⁢

a

>

0

κ

⁢

a

,

otherwise

(

2

)

wherein α represents an input variable and κ represents a slope coefficient;

in the Pool layer, a maximum value in a pooling window is selected as a result in a selected maximum pooling operation, and a pooling process can be expressed as follows:

Z

t

=

Pool

(

Z

t

-

1

)

(

3

)

after several convolution and maximum pooling operations, an extracted depth feature is input into a first FC layer; in the FC layer, all neurons between layers are connected, and the depth feature is further mapped to a new feature space to complete a classification task;

calculation is performed through a weight W t and the bias b t , and a nonlinear transformation is performed using the activation function ƒ(⋅); and a calculation process is as follows:

Z

t

=

f

⁡

(

W

t

·

Z

t

-

1

+

b

t

)

(

4

)

an output layer of the convolutional neural network uses Softmax activation function, and an output thereof Z t =(z 1 , z 2 , . . . , z K ) T is transformed into a corresponding probability result E i ==(e 1 , e 2 , . . . , e K ) T , wherein K represents a total number of category, and its calculation process is

e

K

=

exp

⁡

(

z

K

)

/

∑

k

=

1

K

exp

⁡

(

z

k

)

(

5

)

during the training process, a cross-entropy loss function is employed to measure a difference between the output of the convolutional neural network and a ground-truth label, and its calculation formula is as follows:

min

-

1

N

⁢

∑

i

=

1

N

∑

k

=

1

K

Y

i

,

k

⁢

log

⁡

(

e

k

(

E

i

)

)

(

6

)

the weight W t and bias b t of the convolutional neural network are iterated and optimized through an error backpropagation algorithm; a loss value of formula (6) is backpropagated from a last layer to a first layer, and parameter updating is performed according to the error of each layer; assuming that assuming that a derivative of cross-entropy loss for the weight W t is ΔW t , and a derivative of cross-entropy loss for the bias b t is Δb t , formula for parameter updating are expressed as follows:

W

t

l

=

W

t

l

-

θΔ

⁢

W

t

l

-

1

(

7

)

b

t

l

=

b

t

l

-

θΔ

⁢

b

t

l

-

1

(

8

)

wherein l represents a training iteration epoch and θ represents a learning rate;

in order to facilitate the distinction, a probability distribution of the output layer softened by the distillation temperature T is called the soft label information s i =(s 1 , s 2 , . . . , s K ) T , and its calculation method is as follows:

s

K

=

exp

⁡

(

z

K

/

T

)

/

∑

k

=

1

K

exp

⁡

(

z

k

/

T

)

(

9

)

in a binary-class classification task, assuming that the training dataset is expressed as X={x i , x i =(x 1 , x 2 , . . . , x n ) T , x i ∈R n , i=1, 2, . . . , N}, {tilde over (Y)}={Y i , Y i ∈{0,1}, i=1, 2, . . . , N} is a ground-truth label vector of the binary-class classification task; the probability distribution of the output layer E i =(e 1 , e 2 ) T is transformed into the soft label information S={s i =(s 1 , s 2 ) T , s k ∈[0,1], k=1,2, i=1, 2, . . . , N} after introducing the distillation temperature T, wherein the sum of all probabilities in each sample is 1, that is, s 1 +s 2 =1.

10 . The method for constructing the fuzzy rule according to claim 9 , wherein the Step 2, the CNNBaTSK adopts following fuzzy rules:

In

⁢

rule

⁢

m

:

IF

⁢

s

1

⁢

is

*

^

s

2

⁢

is

*

THEN

⁢

y

m

=

p

0

m

+

p

1

m

⁢

x

1

+

p

2

m

⁢

x

2

+

…

+

p

n

m

⁢

x

n

.

m

=

1

,

2

,

…

,

M

(

10

)

wherein * denotes the membership of fuzzy partition, M is the total number of fuzzy rules and ∧ denotes a fuzzy conjunction operator, and the consequent part adopts a linear function of the input data x i ;

the antecedent part of the soft label information is taken into five fixed fuzzy partitions, wherein each fuzzy partition has a center of

c

d

m

∈

{

0

,

0.25

,

0.5

,

0.75

,

1

}

,

and the kernel width

σ

d

m

is set at a random positive value which is between 0 and 1; then, a Gaussian membership function as a fuzzy membership function is applied to estimate a membership of each rule, and its calculation method is as follows:

μ

m

(

s

i

,

d

)

=

exp

⁡

(

-

(

s

i

,

d

-

c

d

m

)

2

/

2

⁢

σ

d

m

)

(

11

)

μ

m

(

s

i

)

=

∏

2

d

=

1

μ

m

(

s

i

,

d

)

⁢

μ

m

(

s

i

)

=

∏

2

d

=

1

μ

m

(

s

i

,

d

)

(

12

)

a normalized membership function is formulated as follows:

μ

~

m

(

s

i

)

=

μ

m

⁢

〈

s

i

)

/

∑

m

′

=

1

M

μ

m

′

(

s

i

)

(

13

)

therefore, the output {tilde over (y)} i of the CNNBaTSK can be expressed as:

y

~

i

=

∑

m

=

1

M

μ

~

m

(

s

i

)

⁢

y

m

(

14

)

making:

p

m

=

(

p

0

m

,

p

1

m

,

…

,

p

n

m

)

T

(

15

)

P

=

(

(

p

l

)

T

,

(

p

2

)

T

,

…

,

(

p

m

)

T

)

T

(

16

)

q

i

m

=

μ

~

m

(

s

i

)

⁢

(

l

,

x

i

)

T

(

17

)

Q

l

=

(

(

q

i

l

)

T

,

(

q

i

2

)

T

,

…

,

(

q

i

m

)

T

)

T

(

18

)

the output {tilde over (y)} i of the CNNBaTSK may be expressed as:

y

~

i

=

Q

i

T

⁢

P

.

(

19

)

11 . The method for constructing the fuzzy rule according to claim 10 , wherein in Step 3, a specific method for optimizing the consequent parameter of the CNNBaTSK is as follows:

an objective function is designed to balance the ground-truth label and the soft label information, and a following objective function for optimization is proposed:

min

⁢

1

2

⁢

∑

N

i

=

1

ς

i

2

+

α

2

⁢

P

2

2

+

β

2

⁢

∑

i

=

1

N

ξ

i

2

(

20

)

s

.

t

.

{

Q

i

T

⁢

P

-

Y

i

=

ς

i

Q

i

T

⁢

P

-

s

i

,

2

=

ξ

i

,

i

=

1

,

2

,

…

,

N

wherein P is the consequent parameter of the fuzzy rule, α and β are regularization factors, ζ i is a deviation between the prediction output and the ground-truth label, ξ i is a deviation between the prediction output and the soft label information; based on s 1 +s 2 =1 and the ground-truth label, index of the probability s 2 is matched with the ground-truth label;

a Lagrangian optimization formula of the objective function can be expressed as:

L

⁡

(

P

,

ς

,

ξ

,

τ

,

v

)

=

1

2

⁢

ς

2

2

+

α

2

⁢

P

2

2

+

β

2

⁢

ξ

2

2

-

τ

⁡

(

Q

~

⁢

P

-

Y

~

-

ς

)

-

v

⁡

(

Q

~

⁢

P

-

S

~

-

ξ

)

(

21

)

wherein ζ=(ζ 1 , ζ 2 , . . . , ζ N ) T , ξ=(ξ 1 , ξ 2 , . . . , ξ N ) T , τ=(τ 1 , τ 2 , . . . , τ N ) T and ν=(ν 1 , ν 2 , . . . , ν N ) T are Lagrangian multipliers with equality constraints, and an input matrix is denoted by {tilde over (Q)}=(Q 1 , Q 2 , . . . , Q N ) T ;

the consequent parameter P is calculated by setting a derivative gradient of the Lagrangian optimization formula with respect to (P, ζ, ξ, τ, ν) equal to be zero, resulting in following KKT optimality conditions:

{

∂

L

∂

P

⇒

α

⁢

P

T

=

τ

⁢

Q

~

+

v

⁢

Q

~

∂

L

∂

ς

⇒

ς

T

+

τ

=

0

∂

L

∂

ξ

⇒

β

⁢

ξ

T

+

v

=

0

∂

L

∂

τ

⇒

Q

~

⁢

P

-

Y

~

-

ς

=

0

∂

L

∂

v

⇒

Q

~

⁢

P

-

S

~

-

ξ

=

0

(

22

)

yielding:

{

τ

=

-

(

Q

~

⁢

P

-

Y

~

)

T

v

=

-

β

⁡

(

Q

~

⁢

P

-

S

~

)

T

α

⁢

P

T

=

-

(

Q

~

⁢

P

-

Y

~

)

T

⁢

Q

~

-

β

⁡

(

Q

~

⁢

P

-

S

~

)

T

⁢

Q

~

(

23

)

therefore, an optimal analytical solution can be expressed as:

P

=

(

α

⁢

I

+

(

1

+

β

)

⁢

Q

~

T

⁢

Q

~

)

-

1

⁢

(

Q

~

T

⁢

Y

~

+

β

⁢

Q

~

T

⁢

S

~

)

(

24

)

wherein I denotes an identity matrix.

12 . The method for constructing the fuzzy rule according to claim 7 , wherein a learning algorithm of CNNBaTSK has inputs comprising: training dataset X={x i , x i =(x 1 , x 2 , . . . , x n ) T , x i ∈R n , i=1,2 . . . , N} and {tilde over (Y)}={Y i , Y i ∈{0, 1, . . . , K}, i=1,2, . . . , N}, and a maximum iteration epoch {tilde over (l)}, initial learning rate θ, distillation temperature T, regularization factors α and β, testing sample x test of the convolutional neural network; outputs comprising CNNBaTSK after completion of training, and the output {tilde over (y)} test of the testing sample; the learning algorithm of the CNNBaTSK comprises two stages as follows:

S1: training stage:

S10: initialization: initializing W t and b t with random numbers, and proceeding step S11 to step S13 sequentially within a range of the maximum iteration epoch {tilde over (l)};

S11: inputting training samples X={x i , x i =x 1 , x 2 , . . . , x n ) T , x i ∈R n , i=1, 2, . . . , N}, and performing a convolutional and pooling operation on the training samples through the Conv layer and the Pool layer to extract depth features;

S12: calculating a probability that the sample belongs to each label at the FC layer;

S13: performing error backpropagation on 1D-CNN according to a cross-entropy loss function formula;

S14: generating a probability distribution E i =(e 1 , e 2 ) T of the output layer after the iteration is completed;

S15: introducing the distillation temperature T and converting it into the soft label information S={s i =(s 1 ,s 2 ) T , s k ∈[0,1], k=1,2, i=1,2, . . . , N};

S16: partitioning the antecedent part into five fixed fuzzy partitions, wherein the center of each fuzzy partition is

c

d

m

∈

{

0

,

0.25

,

0.5

,

0.75

,

1

}

,

 and the kernel width

σ

d

m

 is set at a random positive value which is between 0 and 1;

S17: computing a normalized membership {acute over (μ)} m (s i ) of the training sample in different rules;

S18: computing and generating Q i , wherein the output of the CNNBaTSK can be expressed as

y

~

i

=

Q

i

T

⁢

P

;

S19: computing an analytical solution for P to be: P=(αI+(1+β){tilde over (Q)} T {tilde over (Q)}) −1 ({tilde over (Q)} T {tilde over (Y)}+β{tilde over (Q)} T {tilde over (S)});

S2: testing stage:

S21: inputting a testing sample x test ;

S22: generating the probability distribution E test =(e 1 , e 2 ) T of the output layer;

S23: computing the probability distribution s test =(s 1 , s 2 ) T of the output layer of the soft label information;

S24: computing a normalized membership {tilde over (ρ)} m (s i ) of the testing sample in different rules;

S25: outputting {tilde over (y)} test of the testing sample according to

y

~

test

=

Q

test

T

⁢

P

.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jun 30, 2023
From: JIANG, YUNLIANG; ZHANG, XIONGTAO; LOU, JUNGANG; SHEN, QING; WENG, JIANGWEI
To: HUZHOU UNIVERSITY
Reel/Frame 064131/0691 →
Priority Claims (1)
CN 202210651346.2 · Jun 9, 2022 · national
Continuity (1)
Related Publication 20230401424A1 · Dec 14, 2023
References Cited (7)
US 5245695A · Basehore · 1993 [cited by examiner]
US 6748369B2 · Khedkar · 2004 [cited by examiner]
US 11410029B2 · Fukuda · 2022 [cited by examiner]
Distilling a Deep Neural Network into a Takagi-Sugeno-Kang Fuzzy Inference System (Year: 2020). [cited by examiner]
The Scalable Fuzzy Inference-Based Ensemble Method for Sentiment Analysis (Year: 2022). [cited by examiner]
A CNN-Based Born-Again TSK Fuzzy Classifier Integrating Soft Label Information and Knowledge Distillation (Year: 2023). [cited by examiner]
Article; arXiv:2010.04974v1 [cs.AI] Oct. 10, 2020; Distilling a Deep Neural Network into a Takagi-Sugeno-Kang Fuzzy Inference System; Xiangming Gu, Tsinghua University, 30 Shuangqing Rd, Beijing, China; Xiang Cheng, Nat… [cited by applicant]