IP Library › Granted Patent US 12,670,401
Granted Patent B2
US 12,670,401 · App. 18/207,949 · Granted Jun 30, 2026

Systems and methods for normalization in deep learning

Inventors: Hidenori Tanaka (Sunnyvale, CA); Ekdeep Singh Lubana (Sunnyvale, CA)
Assignee: NTT RESEARCH, INC.
G06N3/084G06N3/0464
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,670,401
App. No.
18/207,949
Filed
Jun 9, 2023
Granted
Jun 30, 2026
Kind
B2
Art Unit
2407
USPC
706/25
Abstract

A method of training a neural network may be provided. The method may include receiving a raw data set and normalizing the raw data set using a group normalization layer to generate a training data set for training a neural network, wherein the group normalization layer segments the raw data set into a plurality of groups of a predetermined size. The method may also include randomly initializing the neural network by randomly assigning values to a plurality of weights corresponding to a plurality of layers of the neural network. The method may further include training the neural network using the training data set.

Claims (40)

1 . A method of training a neural network, the method comprising:

receiving a raw data set;

normalizing the raw data set using a group normalization layer to generate a training data set for training a neural network, wherein the group normalization layer segments the raw data set into a plurality of groups of a predetermined size;

randomly initializing the neural network by randomly assigning values to a plurality of weights corresponding to a plurality of layers of the neural network; and

training the neural network using the training data set.

2 . The method of claim 1 , wherein the group normalization layer comprises an activation based layer.

3 . The method of claim 1 , wherein the group normalization layer comprises a parametric layer.

4 . The method of claim 1 , further comprising:

causing, by the group normalization layer, a linear growth of activation variance during a forward propagation within the neural network.

5 . The method of claim 1 , further comprising:

causing, by a non-linearity in a residual path of the neural network, a linear growth of activation variance during a forward propagation within the neural network.

6 . The method of claim 1 , further comprising:

preventing, by the group normalization layer, a rank collapse within at least one internal layer of the neural network.

7 . The method of claim 1 , further comprising:

avoiding, by the group normalization layer, gradient explosion during a backward propagation within the neural network.

8 . The method of claim 1 , further comprising:

determining the size based on training speed and training stability of the neural network.

9 . The method of claim 1 , wherein the neural network comprises a deep neural network.

10 . The method of claim 1 , wherein the neural network comprises a convolutional neural network.

11 . A system comprising:

a non-transitory storage medium storing computer program instructions; and

one or more processors configured to execute the computer program instructions to cause operations comprising:

receiving a raw data set;

normalizing the raw data set using a group normalization layer to generate a training data set for training a neural network, wherein the group normalization layer segments the raw data set into a plurality of groups of a predetermined size;

randomly initializing the neural network by randomly assigning values to a plurality of weights corresponding to a plurality of layers of the neural network; and

training the neural network using the training data set.

12 . The system of claim 11 , wherein the group normalization layer comprises an activation based layer.

13 . The system of claim 11 , wherein the group normalization layer comprises a parametric layer.

14 . The system of claim 11 , wherein the operations further comprise:

causing, by the group normalization layer, a linear growth of activation variance during a forward propagation within the neural network.

15 . The system of claim 11 , wherein the operations further comprise:

causing, by a non-linearity in a residual path of the neural network, a linear growth of activation variance during a forward propagation within the neural network.

16 . The system of claim 11 , wherein the operations further comprise:

preventing, by the group normalization layer, a rank collapse within at least one internal layer of the neural network.

17 . The system of claim 11 , wherein the operations further comprise:

avoiding, by the group normalization layer, gradient explosion during a backward propagation within the neural network.

18 . The system of claim 11 , wherein the operations further comprise:

determining the size based on training speed and training stability of the neural network.

19 . The system of claim 11 , wherein the neural network comprises a deep neural network.

20 . The system of claim 11 , wherein the neural network comprises a convolutional neural network.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jun 14, 2023
From: TANAKA, HIDENORI; LUBANA, EKDEEP SINGH
To: NTT RESEARCH, INC.
Reel/Frame 063949/0715 →
Continuity (2)
Provisional Application 63350820 · Jun 9, 2022
Related Publication 20230401448A1 · Dec 14, 2023
References Cited (54)
US 6442535B1 · Yifan · 2002 [cited by examiner]
US 9530042B1 · Saeed · 2016 [cited by examiner]
US 10354184B1 · Vitaladevuni · 2019 [cited by examiner]
US 11113578B1 · Brandt · 2021 [cited by examiner]
US 11232016B1 · Huynh · 2022 [cited by examiner]
US 11423303B1 · Jiao · 2022 [cited by examiner]
US 11494321B1 · Yu · 2022 [cited by examiner]
US 11941784B1 · Ikuta · 2024 [cited by examiner]
US 11961619B1 · LaBorde · 2024 [cited by examiner]
US 12020265B1 · Mao · 2024 [cited by examiner]
US 12175222B1 · Benfield · 2024 [cited by examiner]
US 12400106B1 · Diamant · 2025 [cited by examiner]
US 20040044503A1 · McConaghy · 2004 [cited by examiner]
US 20050147047A1 · Monk · 2005 [cited by examiner]
US 20050288812A1 · Cheng · 2005 [cited by examiner]
US 20160085430A1 · Moran · 2016 [cited by examiner]
US 20180088996A1 · Rossi · 2018 [cited by examiner]
US 20180137857A1 · Zhou · 2018 [cited by examiner]
US 20190164050A1 · Chen · 2019 [cited by examiner]
US 20190188567A1 · Yao · 2019 [cited by examiner]
US 20190251723A1 · Coppersmith, III · 2019 [cited by examiner]
US 20190311268A1 · Tilton · 2019 [cited by examiner]
US 20190340543A1 · Gerenstein · 2019 [cited by examiner]
US 20190362236A1 · Wang · 2019 [cited by examiner]
US 20200004815A1 · Weisberg · 2020 [cited by examiner]
US 20200065677A1 · Iriarte Lopez · 2020 [cited by examiner]
US 20200132547A1 · Yu · 2020 [cited by examiner]
US 20200175360A1 · Conti · 2020 [cited by examiner]
US 20200303078A1 · Mayhew · 2020 [cited by examiner]
US 20200311522A1 · Tzoufras · 2020 [cited by examiner]
US 20210142151A1 · Mixter · 2021 [cited by examiner]
US 20210158132A1 · Huynh · 2021 [cited by examiner]
US 20210168165A1 · Alsaeed · 2021 [cited by examiner]
US 20210174196A1 · Desmond · 2021 [cited by examiner]
US 20210256386A1 · Wieman · 2021 [cited by examiner]
US 20210357744A1 · Mittal · 2021 [cited by examiner]
US 20210397943A1 · Ribalta · 2021 [cited by examiner]
US 20220012860A1 · Zhang · 2022 [cited by examiner]
US 20220061746A1 · Lyman · 2022 [cited by examiner]
US 20220075659A1 · Mohapatra · 2022 [cited by examiner]
US 20220084732A1 · Liu · 2022 [cited by examiner]
US 20220114390A1 · Baughman · 2022 [cited by examiner]
US 20220147792A1 · Jiao · 2022 [cited by examiner]
US 20220189612A1 · Zhai · 2022 [cited by examiner]
US 20220198011A1 · Kumar · 2022 [cited by examiner]
US 20220237465A1 · Lewis · 2022 [cited by examiner]
US 20220327408A1 · Meng · 2022 [cited by examiner]
US 20220366532A1 · Shanmuga Vadivel · 2022 [cited by examiner]
US 20220408698A1 · Du · 2022 [cited by examiner]
US 20230098994A1 · Labatie · 2023 [cited by examiner]
US 20230106639A1 · Wagener · 2023 [cited by examiner]
US 20230156025A1 · Bracht · 2023 [cited by examiner]
US 20230215460A1 · Shiloh Perl · 2023 [cited by examiner]
US 20240331371A1 · Chen · 2024 [cited by examiner]