IP Library › Granted Patent US 11,151,449
Granted Patent B2
US 11,151,449 · App. 15/878,933 · Granted Oct 19, 2021

Adaptation of a trained neural network

Inventors: Masayuki Suzuki (Tokyo, JP); Toru Nagano (Tokyo, JP)
Assignee: INTERNATIONAL BUSINESS MACHINES CORPORATION
G06N3/08G06N5/02
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,151,449
App. No.
15/878,933
Granted
Oct 19, 2021
Kind
B2
Abstract

A method, computer program product, and apparatus for adapting a trained neural network having one or more batch normalization layers are provided. The method includes adapting only the one or more batch normalization layers using adaptation data. The method also includes adapting the whole of the neural network having the one or more adapted batch normalization layers, using the adaptation data.

Claims (147)

1. A computer-implemented method for adapting a trained neural network having one or more batch normalization layers, the method comprising:

adapting only the one or more batch normalization layers using adaptation data; and

adapting the whole of the neural network having the one or more adapted batch normalization layers, using the adaptation data,

wherein each batch normalization layer has a mean, a variance, a scale parameter, and a shift parameter,

wherein the neural network having the one or more batch normalization layers was already trained using training data derived from source domain data, the adaptation data is derived from target domain data; and the scale parameter is adjusted according to the following equation:

γ

^

=

γ

⁢

σ

adaptation

2

+

ϵ

σ

base

2

+

ϵ

where {circumflex over (γ)} denotes an adjusted scale parameter; β denotes the scale parameter of the batch normalization layer before adapted, σ adaptation 2 denotes a final value of running variance recalculated using the adaptation data, σ base 2 denotes a final value of running variance of the batch normalization layer before adapted, and E denotes a predefined value.

2. The method according to claim 1 , wherein

the adapting only the one or more batch normalization layers comprises:

recalculating, for each batch normalization layer, the mean and the variance using the adaptation data; and

adjusting, for each batch normalization layer, the scale parameter and the shift parameter so that an output of the batch normalization layer with the recalculated mean and variance does not change.

3. The method according to claim 1 , wherein the neural network having the one or more batch normalization layers was already trained using training data derived from source domain data, the adaptation data is derived from target domain data; and the shift parameter is adjusted according to the following equation:

β

^

=

γ

⁢

μ

adaptation

-

μ

base

σ

base

2

+

ϵ

+

β

where {circumflex over (β)} denotes an adjusted shift parameter, β denotes the shift parameter of the batch normalization layer before adapted, γ denotes the scale parameter of the batch normalization layer before adapted, μ adaptation denotes a final value of running mean recalculated using the adaptation data, μ base denotes a final value of running mean of the batch normalization layer before adapted, σ base 2 denotes a final value of running variance of the batch normalization layer before adapted, and ϵ denotes a predefined value.

4. The method according to claim 1 , wherein the adapting only the one or more batch normalization layers comprises

freezing all parameters of the neural network other than parameters of each batch normalization layer so as to adapt only the parameters of the batch normalization layer using the adaptation data, wherein the parameters of the batch normalization layer are a mean, a variance, a scale parameter, and a shift parameter.

5. The method according to claim 1 , wherein the neural network is a convolutional neural network, a recurrent neural network, or a feed-forward neural network.

6. The method according to claim 1 , wherein the adapting only the one or more batch normalization layers and the adapting the whole of the neural network are done using a first adaptation data; and the method further comprising

adapting only the one or more adapted batch normalization layers using second adaptation data of a second target domain; and

adapting the whole of the adapted neural network having the one or more adapted batch normalization layers, using the second adaptation data.

7. The method according to claim 6 , wherein the following adaptations are iterated until parameters of each batch normalization layer converges: the adapting only the one or more adapted batch normalization layers using the first adaptation data, the adapting the whole of the adapted neural network using the first adaptation data, the adapting only the one or more adapted batch normalization layers using the second adaptation data, and the adapting the whole of the adapted neural network using the second adaptation data.

8. A computer system, comprising:

one or more processors; and

a memory storing a program which, when executed on the processor, performs an operation for adapting a trained neural network having one or more batch normalization layers, the operation comprising:

adapting only the one or more batch normalization layers using adaptation data; and

adapting the whole of the neural network having the one or more adapted batch normalization layers, using the adaptation data,

wherein each batch normalization layer has a mean, a variance, a scale parameter, and a shift parameter,

wherein the neural network having the one or more batch normalization layers was already trained using training data derived from source domain data, the adaptation data is derived from target domain data; and the scale parameter is adjusted according to the following equation:

γ

^

=

γ

⁢

σ

adaptation

2

+

ϵ

σ

base

2

+

ϵ

where {circumflex over (γ)} denotes an adjusted scale parameter, γ denotes the scale parameter of the batch normalization layer before adapted, σ adaptation 2 denotes a final value of running variance recalculated using the adaptation data, σ base 2 denotes a final value of running variance of the batch normalization layer before adapted, and ϵ denotes a predefined value.

9. The computer system according to claim 8 , wherein

the adapting only the one or more batch normalization layers comprises:

recalculating, for each batch normalization layer, the mean and the variance using the adaptation data; and

adjusting, for each batch normalization layer, the scale parameter and the shift parameter so that an output of the batch normalization layer with the recalculated mean and variance does not change.

10. The computer system according to claim 8 , wherein the neural network having the one or more batch normalization layers was already trained using training data derived from source domain data, the adaptation data is derived from target domain data; and the shift parameter is adjusted according to the following equation:

β

^

=

γ

⁢

μ

adaptation

-

μ

base

σ

base

2

+

ϵ

+

β

where {circumflex over (β)} denotes an adjusted shift parameter, β denotes the shift parameter of the batch normalization layer before adapted, γ denotes the scale parameter of the batch normalization layer before adapted, μ adaptation denotes a final value of running mean recalculated using the adaptation data, μ base denotes a final value of running mean of the batch normalization layer before adapted, σ base 2 denotes a final value of running variance of the batch normalization layer before adapted, and ϵ denotes a predefined value.

11. The computer system according to claim 8 , wherein the adapting only the one or more batch normalization layers comprises

freezing all parameters of the neural network other than parameters of each batch normalization layer so as to adapt only the parameters of the batch normalization layer using the adaptation data, wherein the parameters of the batch normalization layer are a mean, a variance, a scale parameter, and a shift parameter.

12. The computer system according to claim 8 , wherein the neural network is a convolutional neural network, a recurrent neural network, or a feed-forward neural network.

13. A computer program product for adapting a trained neural network having one or more batch normalization layers, the computer program product comprising a computer readable storage medium having program instructions embodied therewith, the program instructions executable by a computer to cause the computer to perform a method comprising:

adapting only the one or more batch normalization layers using adaptation data; and

adapting the whole of the neural network having the one or more adapted batch normalization layers, using the adaptation data,

wherein each batch normalization layer has a mean, a variance, a scale parameter, and a shift parameter

wherein the neural network having the one or more batch normalization layers was already trained using training data derived from source domain data, the adaptation data is derived from target domain data; and the scale parameter is adjusted according to the following equation:

γ

^

=

γ

⁢

σ

adaptation

2

+

ϵ

σ

base

2

+

ϵ

where {circumflex over (γ)} denotes an adjusted scale parameter, γ denotes the scale parameter of the batch normalization layer before adapted, σ adaptation 2 denotes a final value of running variance recalculated using the adaptation data, σ base 2 denotes a final value of running variance of the batch normalization layer before adapted, and ϵ denotes a predefined value.

14. The computer program product according to claim 13 , wherein

the adapting only the one or more batch normalization layers comprises:

recalculating, for each batch normalization layer, the mean and the variance using the adaptation data; and

adjusting, for each batch normalization layer, the scale parameter and the shift parameter so that an output of the batch normalization layer with the recalculated mean and variance does not change.

15. The computer program product according to claim 13 , wherein the neural network having the one or more batch normalization layers was already trained using training data derived from source domain data, the adaptation data is derived from target domain data; and the shift parameter is adjusted according to the following equation:

β

^

=

γ

⁢

μ

adaptation

-

μ

base

σ

base

2

+

ϵ

+

β

where {circumflex over (β)} denotes an adjusted shift parameter, β denotes the shift parameter of the batch normalization layer before adapted, γ denotes the scale parameter of the batch normalization layer before adapted, μ adaptation denotes a final value of running mean recalculated using the adaptation data, μ base denotes a final value of running mean of the batch normalization layer before adapted, σ base 2 denotes a final value of running variance of the batch normalization layer before adapted, and ϵ denotes a predefined value.

16. The computer program product according to claim 13 , wherein the adapting only the one or more batch normalization layers comprises

freezing all parameters of the neural network other than parameters of each batch normalization layer so as to adapt only the parameters of the batch normalization layer using the adaptation data, wherein the parameters of the batch normalization layer are a mean, a variance, a scale parameter, and a shift parameter.

17. The computer program product according to claim 13 , wherein the neural network is a convolutional neural network, a recurrent neural network, or a feed-forward neural network.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jan 24, 2018
From: SUZUKI, MASAYUKI; NAGANO, TORU
To: INTERNATIONAL BUSINESS MACHINES CORPORATION
Reel/Frame 044716/0889 →
Continuity (1)
Related Publication 20190228298A1 · Jul 25, 2019