IP Library › Granted Patent US 11,755,916
Granted Patent B2
US 11,755,916 · App. 16/562,067 · Granted Sep 12, 2023

System and method for improving deep neural network performance

Inventors: Yanshuai Cao (Toronto, CA); Ruitong Huang (Toronto, CA); Junfeng Wen (Toronto, CA)
Assignee: ROYAL BANK OF CANADA
G06N3/084G06F18/24G06N3/047G06N3/08G06N7/01G06V10/764G06V10/7788G06V10/82
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,755,916
App. No.
16/562,067
Granted
Sep 12, 2023
Kind
B2
Abstract

An improved computer implemented method and corresponding systems and computer readable media for improving performance of a deep neural network are provided to mitigate effects related to catastrophic forgetting in neural network learning. In an embodiment, the method includes storing, in memory, logits of a set of samples from a previous set of tasks (D 1 ); and maintaining classification information from the previous set of tasks by utilizing the logits for matching during training on a new set of tasks (D 2 ).

Claims (333)

1. A computer implemented method for training performance of a deep neural network adapted to attain a model f T :X Δ C that maps data in an input space X to a C-dimensional probability simplex that reduces catastrophic forgetting on a first T data sets after training on T sequential tasks, D T representing a current available data set, and D t , t≤T representing additional data sets and the current available data set, the computer implemented method comprising:

storing, in non-transitory computer readable memory, logits of a set of samples from a previous set of tasks, D 1 , the storage establishing a memory cost m<<n 1 ;

maintaining classification information from the previous set of tasks by utilizing the logits for matching during training on a new set of tasks, D 2 , the logits selected to reduce a dependency on representation of D 1 ; and

training the deep neural network on D 2 , and applying a penalty on the deep neural network for prediction deviation, the penalty adapted to sample a memory x i (1) , i=1, . . . , m from D 1 and matching outputs for f 1 * when training f 2 ; wherein m is a memory size of the memory x i (1) , n 1 is a number of tasks in D 1 , and f 1 * is a model f trained on the first T data sets.

2. The method of claim 1 , wherein the penalty on the deep neural network for the prediction deviation is established according to a relation:

min

θ

1

n

2

⁢

∑

i

L

⁡

(

y

i

(

2

)

,

f

θ

(

x

i

(

2

)

)

)

+

1

m

⁢

∑

j

L

⁡

(

x

j

(

1

)

(

f

1

*

(

x

j

(

1

)

)

)

,

f

θ

(

x

j

(

1

)

)

)

where (x j (1) , y j (1) ) is a data pair of D 1 , (x i (2) , y i (2) ) is a data pair of D 2 , L is a Kullback-Leibler (KL) divergence, and f is parametrized by a vector Θθ∈ p ; wherein i and j are index notations.

3. The method of claim 2 , wherein L 2 regularization is applied to the logits, in accordance with a relation:

min

θ

⁢

1

n

2

⁢

∑

i

⁢

L

⁡

(

y

i

(

2

)

,

f

θ

⁡

(

x

i

(

2

)

)

)

+

1

m

⁢

∑

j

⁢

z

^

j

(

1

)

-

z

^

j

(

2

)

2

2

,

where {circumflex over (z)} j (1) ,{circumflex over (z)} j (2) are the logits produced by f 1 * and f θ respectively.

4. The method of claim 3 , comprising: applying a logits matching regularization R in accordance with a relation:

ℛ

⁡

(

,

)

=

1

K

⁢

∑

y

⁢

(

⁢

(

y

)

-

⁢

(

y

)

)

2

where:

(x,y) data pair,

ŷ predicted label

the output probability vector with logits

the output probability vector with logits

τ temperature hyperparameter

K number of class.

5. The method of claim 1 , wherein the performance improvement is a reduction of a forgetting behavior.

6. The method of claim 5 , wherein the reduction of the forgetting behavior includes while training on D 2 , the deep neural network is still capable of predicting on D 1 .

7. The method of claim 1 , wherein the non-transitory computer readable memory has a limited memory size having a float number memory size/task ratio selected from at least one of 10, 50, 100, 500, 1000, 1900, or 1994.

8. The method of claim 1 , wherein the deep neural network is configured for image recognition tasks, and wherein both the previous set of tasks and the new set of tasks are image classification tasks.

9. The method of claim 8 , wherein the previous set of tasks includes processing a permuted image data set, and wherein the new set of tasks includes processing the permuted image data set where pixels of each underlying image are linearly transformed.

10. The method of claim 8 , wherein the previous set of tasks includes processing a permuted image data set, and wherein the new set of tasks includes processing the permuted image data set where pixels of each underlying image are non-linearly transformed.

11. A computing device adapted for training performance of a deep neural network adapted to attain a model f T :X Δ C that maps data in an input space X to a C-dimensional probability simplex that reduces catastrophic forgetting on a first T data sets after training on T sequential tasks, D T representing a current available data set, and t , t≤T representing additional data sets and the current available data set, the computing device comprising a computer processor operating in conjunction with non-transitory computer memory, the computer processor configured to:

store, in the non-transitory computer readable memory, logits of a set of samples from a previous set of tasks, D 1 , the storage establishing a memory cost m<<n 1 ;

maintain classification information from the previous set of tasks by utilizing the logits for matching during training on a new set of tasks, D 2 , the logits selected to reduce a dependency on representation of D 1 ; and

train the deep neural network on D 2 , and apply a penalty on the deep neural network for prediction deviation, the penalty adapted to sample a memory x i (1) , i=1, . . . , m from D i and matching outputs for f 1 * when training f 2 ; wherein m is a memory size of the memory x i (1) , n 1 is a number of tasks in D 1 , and f 1 * is a model f trained on the first T data sets.

12. The device of claim 11 , wherein the penalty on the deep neural network for the prediction deviation is established according to a relation:

min

θ

⁢

1

n

2

⁢

∑

i

⁢

L

⁡

(

y

i

(

2

)

,

f

θ

⁡

(

x

i

(

2

)

)

)

+

1

m

⁢

∑

j

⁢

L

⁡

(

f

1

*

⁡

(

x

j

(

1

)

)

,

f

θ

⁡

(

x

j

(

1

)

)

)

,

where (x j (1) , y j (1) ) is a data pair of D 1 , (x i (2) ; y i (2) ) is a data pair of D 2 , Lis the Kullback-Leibler (KL) divergence, and f is parametrized by a vector Θθ∈ p ; wherein i and j are index notations.

13. The device of claim 12 , wherein L 2 regularization is applied to the logits, in accordance with a relation:

min

θ

⁢

1

n

2

⁢

∑

i

⁢

L

⁡

(

y

i

(

2

)

,

f

θ

⁡

(

x

i

(

2

)

)

)

+

1

m

⁢

∑

j

⁢

z

^

j

(

1

)

-

z

^

j

(

2

)

2

2

,

where {circumflex over (z)} j (1) ,{circumflex over (z)} j (2) are the logits produced by f 1 * and f θ respectively.

14. The device of claim 13 , wherein the computer processor is further configured to: apply logits matching regularization R in accordance with a relation:

ℛ

⁡

(

,

)

=

1

K

⁢

∑

y

⁢

(

⁢

(

y

)

-

⁢

(

y

)

)

2

where:

(x,y) data pair,

ŷ predicted label

the output probability vector with logits

the output probability vector with logits

τ temperature hyperparameter

K number of class.

15. The device of claim 11 , wherein a performance improvement is a reduction of a forgetting behavior.

16. The device of claim 15 , wherein the reduction of the forgetting behavior includes while training on D 2 , the deep neural network is still capable of predicting on D 1 .

17. The device of claim 11 , wherein the non-transitory computer readable memory has a limited memory size having a float number memory size/task ratio selected from at least one of 10, 50, 100, 500, 1000, 1900, or 1994.

18. The device of claim 11 , wherein the deep neural network is configured for image recognition tasks, and wherein both the previous set of tasks and the new set of tasks are image classification tasks.

19. The device of claim 18 , wherein the previous set of tasks includes processing a permuted image data set, and wherein the new set of tasks includes processing the permuted image data set where pixels of each underlying image are non-linearly transformed.

20. A non-transitory computer readable memory storing machine interpretable instructions, which when executed by a processor, cause the processor to execute a method for training performance of a deep neural network adapted to attain a model f T :X Δ C that maps data in an input space X to a C-dimensional probability simplex that reduces catastrophic forgetting on a first T data sets after training on T sequential tasks, D T representing a current available data set, and t , t≤T representing additional data sets and the current available data set, the method comprising:

storing, in non-transitory computer readable memory, logits of a set of samples from a previous set of tasks, D 1 , the storage establishing a memory cost m<<n 1 ;

maintaining classification information from the previous set of tasks by utilizing the logits for matching during training on a new set of tasks, D 2 , the logits selected to reduce a dependency on representation of D 1 ; and

training the deep neural network on D 2 , and applying a penalty on the deep neural network for prediction deviation, the penalty adapted to sample a memory x i (1) , i=1, . . . , m from D 1 and matching outputs for f 1 * when training f 2 ; wherein m is a memory size of the memory x i (1) , n 1 is a number of tasks in D 1 , and f 1 * is a model f trained on the first T data sets.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Apr 4, 2023
From: CAO, YANSHUAI; HUANG, RUITONG; WEN, JUNFENG
To: ROYAL BANK OF CANADA
Reel/Frame 063220/0808 →
Continuity (2)
Provisional Application 62727504 · Sep 5, 2018
Related Publication 20200074305A1 · Mar 5, 2020
Cited By (1)
US 12,572,775