IP Library › Granted Patent US 12,705,484
Granted Patent B2
US 12,705,484 · App. 18/395,282 · Granted Aug 11, 2026

Image classification using batch normalization layers

Inventors: Sergey Ioffe (Mountain View, CA); Corinna Cortes (New York, NY)
Assignee: Google LLC
G06N3/08G06F18/2415G06N3/0464G06N3/084G06V10/70G06V10/82G06T2207/20081
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,705,484
App. No.
18/395,282
Filed
Dec 22, 2023
Granted
Aug 11, 2026
Kind
B2
Art Unit
2669
USPC
382/158
Abstract

Methods, systems, and apparatus, including computer programs encoded on computer storage media, for processing images or features of images using an image classification system that includes a batch normalization layer. One of the systems includes a convolutional neural network configured to receive an input comprising an image or image features of the image and to generate a network output that includes respective scores for each object category in a set of object categories, the score for each object category representing a likelihood that that the image contains an image of an object belonging to the category, and the convolutional neural network comprising: a plurality of neural network layers, the plurality of neural network layers comprising a first convolutional neural network layer and a second neural network layer; and a batch normalization layer between the first convolutional neural network layer and the second neural network layer.

Claims (94)

1 . A system comprising:

a user computer; and

a computer system comprising one or more computers and one or more storage devices storing instructions that when executed by the one or more computers cause the one or more computers to perform operations comprising:

receiving a network input comprising an image or image features of the image from the user computer;

processing the network input using a convolutional neural network configured to receive the network input and to generate a network output that characterizes the image, wherein:

the convolutional neural network includes a first neural network layer and a second neural network layer,

processing the network input using the convolutional neural network comprises processing a first layer input to the first neural network layer in accordance with trained values of a set of parameters of the first neural network layer to generate a first layer output having a plurality of components,

the trained values of the set of parameters of the first neural network layer are a result of training the neural network using a plurality of batches of training data,

each batch of training data comprises a respective plurality of training examples, and

the training of the neural network to determine the trained values of the set of parameters of the first neural network layer comprises, for each of the plurality of batches:

receiving a respective first layer output generated by the first neural network layer for each of the plurality of training examples in the batch;

computing a plurality of normalization statistics for the batch from the first layer outputs, comprising:

determining, for each of a plurality of subsets of the plurality of the components of the first layer outputs, a mean of the components of the first layer outputs for each of the plurality of training examples in the batch that are in the respective subset, and

determining, for each of the plurality of subsets of the plurality of the components of the first layer outputs, a standard deviation of the components of the first layer outputs for each of the plurality of training examples in the batch that are in the respective subset;

generating a respective batch normalization layer output for each training example in the batch, comprising:

for each first layer output and for each of the plurality of subsets, normalizing the components of the first layer output that are in the respective subset using the mean for the respective subset and the standard deviation for the respective subset; and

generating a respective batch normalization layer output for each of the training examples from the normalized layer outputs; and

providing the respective batch normalization layer outputs as inputs to the second neural network layer; and

providing the network output as output of the computer system.

2 . The system of claim 1 , wherein training the neural network to determine the trained values of the set of parameters of the first neural network layer further comprises, for each of the plurality of batches:

generating a respective network output for each training example, comprising processing the respective batch normalization layer outputs using the second neural network layer; and

updating the set of parameters of the first neural network layer using the respective network outputs using a backpropagation technique.

3 . The system of claim 2 , wherein updating the set of parameters of the first neural network layer using the respective network outputs using a backpropagation technique comprises:

backpropagating through the normalization statistics.

4 . The system of claim 1 , wherein the plurality of the components of the first layer output are indexed by dimension, and wherein computing a plurality of normalization statistics for the first layer outputs comprises:

computing, for each of the dimensions, a mean of the components of the first layer outputs in the dimension; and

computing, for each of the dimensions, a standard deviation of the components of the first layer outputs in the dimension.

5 . The system of claim 4 , wherein normalizing each of the plurality of the components of each first layer output comprises:

normalizing the component using the computed mean and computed standard deviation for the dimension corresponding to the component.

6 . The system of claim 4 , wherein generating the respective batch normalization layer output for each of the training examples from the normalized layer outputs comprises:

transforming, for each dimension, the component of the normalized layer output for the training example in the dimension in accordance with current values of a set of parameters for the dimension.

7 . The system of claim 1 , wherein the first neural network layer is a convolutional layer, wherein the plurality of the components of the first layer output are indexed by feature index and spatial location index, and wherein computing a plurality of normalization statistics for the first layer outputs comprises, for each of the feature indices:

computing a mean of the components of the first layer outputs that correspond to the feature index; and

computing a variance of the components of the first layer outputs that correspond to the feature index.

8 . The system of claim 7 , wherein normalizing each of the plurality of the components of each layer output comprises:

normalizing the component using the mean and the variance for the feature index corresponding to the component.

9 . The system of claim 7 , wherein generating the respective batch normalization layer output for each of the training examples from the normalized layer outputs comprises:

transforming each of the plurality of the components of the normalized layer output in accordance with current values of a set of parameters for the feature index corresponding to the component.

10 . One or more non-transitory computer-readable storage media storing instructions that when executed by one or more computers cause the one or more computers to perform operations comprising:

receiving a network input comprising an image or image features of the image from the user computer;

processing the network input using a convolutional neural network configured to receive the network input and to generate a network output that characterizes the image, wherein:

the convolutional neural network includes a first neural network layer and a second neural network layer,

processing the network input using the convolutional neural network comprises processing a first layer input to the first neural network layer in accordance with trained values of a set of parameters of the first neural network layer to generate a first layer output having a plurality of components,

the trained values of the set of parameters of the first neural network layer are a result of training the neural network using a plurality of batches of training data,

each batch of training data comprises a respective plurality of training examples, and

the training of the neural network to determine the trained values of the set of parameters of the first neural network layer comprises, for each of the plurality of batches:

receiving a respective first layer output generated by the first neural network layer for each of the plurality of training examples in the batch;

computing a plurality of normalization statistics for the batch from the first layer outputs, comprising:

determining, for each of a plurality of subsets of the plurality of the components of the first layer outputs, a mean of the components of the first layer outputs for each of the plurality of training examples in the batch that are in the respective subset, and

determining, for each of the plurality of subsets of the plurality of the components of the first layer outputs, a standard deviation of the components of the first layer outputs for each of the plurality of training examples in the batch that are in the respective subset;

generating a respective batch normalization layer output for each training example in the batch, comprising:

for each first layer output and for each of the plurality of subsets, normalizing the components of the first layer output that are in the respective subset using the mean for the respective subset and the standard deviation for the respective subset; and

generating a respective batch normalization layer output for each of the training examples from the normalized layer outputs; and

providing the respective batch normalization layer outputs as inputs to the second neural network layer; and

providing the network output as output of the one or more computers.

11 . The non-transitory computer-readable storage media of claim 10 , wherein training the neural network to determine the trained values of the set of parameters of the first neural network layer further comprises, for each of the plurality of batches:

generating a respective network output for each training example, comprising processing the respective batch normalization layer outputs using the second neural network layer; and

updating the set of parameters of the first neural network layer using the respective network outputs using a backpropagation technique.

12 . The non-transitory computer-readable storage media of claim 11 , wherein updating the set of parameters of the first neural network layer using the respective network outputs using a backpropagation technique comprises:

backpropagating through the normalization statistics.

13 . The non-transitory computer-readable storage media of claim 11 , wherein the plurality of the components of the first layer output are indexed by dimension, and wherein computing a plurality of normalization statistics for the first layer outputs comprises:

computing, for each of the dimensions, a mean of the components of the first layer outputs in the dimension; and

computing, for each of the dimensions, a standard deviation of the components of the first layer outputs in the dimension.

14 . The non-transitory computer-readable storage media of claim 13 , wherein normalizing each of the plurality of the components of each first layer output comprises:

normalizing the component using the computed mean and computed standard deviation for the dimension corresponding to the component.

15 . The non-transitory computer-readable storage media of claim 11 , wherein the first neural network layer is a convolutional layer, wherein the plurality of the components of the first layer output are indexed by feature index and spatial location index, and wherein computing a plurality of normalization statistics for the first layer outputs comprises, for each of the feature indices:

computing a mean of the components of the first layer outputs that correspond to the feature index; and

computing a variance of the components of the first layer outputs that correspond to the feature index.

16 . The non-transitory computer-readable storage media of claim 15 , wherein normalizing each of the plurality of the components of each layer output comprises:

normalizing the component using the mean and the variance for the feature index corresponding to the component.

17 . The non-transitory computer-readable storage media of claim 15 , wherein generating the respective batch normalization layer output for each of the training examples from the normalized layer outputs comprises:

transforming each of the plurality of the components of the normalized layer output in accordance with current values of a set of parameters for the feature index corresponding to the component.

18 . A method performed by one or more computers, the method comprising:

receiving a network input comprising an image or image features of the image from the user computer;

processing the network input using a convolutional neural network configured to receive the network input and to generate a network output that characterizes the image, wherein:

the convolutional neural network includes a first neural network layer and a second neural network layer,

processing the network input using the convolutional neural network comprises processing a first layer input to the first neural network layer in accordance with trained values of a set of parameters of the first neural network layer to generate a first layer output having a plurality of components,

the trained values of the set of parameters of the first neural network layer are a result of training the neural network using a plurality of batches of training data,

each batch of training data comprises a respective plurality of training examples, and

the training of the neural network to determine the trained values of the set of parameters of the first neural network layer comprises, for each of the plurality of batches:

receiving a respective first layer output generated by the first neural network layer for each of the plurality of training examples in the batch;

computing a plurality of normalization statistics for the batch from the first layer outputs, comprising:

determining, for each of a plurality of subsets of the plurality of the components of the first layer outputs, a mean of the components of the first layer outputs for each of the plurality of training examples in the batch that are in the respective subset, and

determining, for each of the plurality of subsets of the plurality of the components of the first layer outputs, a standard deviation of the components of the first layer outputs for each of the plurality of training examples in the batch that are in the respective subset;

generating a respective batch normalization layer output for each training example in the batch, comprising:

for each first layer output and for each of the plurality of subsets, normalizing the components of the first layer output that are in the respective subset using the mean for the respective subset and the standard deviation for the respective subset; and

generating a respective batch normalization layer output for each of the training examples from the normalized layer outputs; and

providing the respective batch normalization layer outputs as inputs to the second neural network layer; and

providing the network output as output of the one or more computers.

19 . The method of claim 18 , wherein training the neural network to determine the trained values of the set of parameters of the first neural network layer further comprises, for each of the plurality of batches:

generating a respective network output for each training example, comprising processing the respective batch normalization layer outputs using the second neural network layer; and

updating the set of parameters of the first neural network layer using the respective network outputs using a backpropagation technique.

20 . The method of claim 19 , wherein updating the set of parameters of the first neural network layer using the respective network outputs using a backpropagation technique comprises:

backpropagating through the normalization statistics.

Assignments (2)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded May 3, 2024
From: IOFFE, SERGEY; CORTES, CORINNA
To: GOOGLE INC.
Reel/Frame 067307/0733 →
ENTITY CONVERSION Recorded May 3, 2024
From: GOOGLE INC.
To: GOOGLE LLC
Reel/Frame 067309/0200 →
Continuity (6)
Continuation 17723007 · Apr 18, 2022
Continuation 16837959 · Apr 1, 2020
Continuation 16226483 · Dec 19, 2018
Continuation 15009647 · Jan 28, 2016
Provisional Application 62108984 · Jan 28, 2015
Related Publication 20240249138A1 · Jul 25, 2024
References Cited (90)
US 5479576A · Watanabe et al. · 1995 [cited by applicant]
US 5541590A · Nishio et al. · 1996 [cited by applicant]
US 5729662A · Rozmus · 1998 [cited by examiner]
US 5790758A · Streit · 1998 [cited by applicant]
US 5875284A · Watanabe et al. · 1999 [cited by applicant]
US 6134537A · Pao · 2000 [cited by applicant]
US 6539267B1 · Eryurek · 2003 [cited by applicant]
US 6650779B2 · Vachtesvanos et al. · 2003 [cited by applicant]
US 7016529B2 · Simard · 2006 [cited by examiner]
US 8331682B2 · Kaehler · 2012 [cited by examiner]
US 8370279B1 · Lin et al. · 2013 [cited by applicant]
US 9058517B1 · Collet et al. · 2015 [cited by applicant]
US 9251407B2 · Kaehler · 2016 [cited by examiner]
US 9373057B1 · Erhan · 2016 [cited by examiner]
US 10127475B1 · Corrado et al. · 2018 [cited by applicant]
US 20020054694A1 · Vachtsevanos et al. · 2002 [cited by applicant]
US 20030023382A1 · Nyland · 2003 [cited by applicant]
US 20030174881A1 · Simard et al. · 2003 [cited by applicant]
US 20030236662A1 · Goodman · 2003 [cited by applicant]
US 20040193559A1 · Hoya · 2004 [cited by applicant]
US 20050125369A1 · Buck et al. · 2005 [cited by applicant]
US 20050283450A1 · Matsugu · 2005 [cited by applicant]
US 20070183686A1 · Ioffe · 2007 [cited by examiner]
US 20080071710A1 · Serre · 2008 [cited by applicant]
US 20140089361A1 · Shibayama · 2014 [cited by applicant]
US 20140279774A1 · Wang · 2014 [cited by examiner]
US 20140365195A1 · Lahiri · 2014 [cited by applicant]
US 20160140425A1 · Kulkarni et al. · 2016 [cited by applicant]
CN 1470022 · 2004 [cited by applicant]
CN 1846218A · 2006 [cited by applicant]
CN 1945602 · 2007 [cited by applicant]
CN 102622418 · 2012 [cited by applicant]
CN 102955946 · 2013 [cited by applicant]
CN 103049792A · 2013 [cited by applicant]
CN 103810999 · 2014 [cited by applicant]
CN 103824055 · 2014 [cited by applicant]
EP 2345984 · 2011 [cited by applicant]
EP 3251059 · 2018 [cited by applicant]
JP 2013069132 · 2014 [cited by applicant]
RU 2424561 · 2009 [cited by applicant]
WO WO2005048185A1 · 2005 [cited by applicant]
WO WO2007027452 · 2007 [cited by applicant]
WO WO2016123409 · 2016 [cited by applicant]
Ciresan, D. C., Meier, U., Masci, J., Maria Gambardella, L., & Schmidhuber, J. (Jul. 2011). Flexible, high performance convolutional neural networks for image classification. In IJCAI proceedings-international joint con… [cited by examiner]
Hinton, G. E., Srivastava, N., Krizhevsky, A., Sutskever, I., & Salakhutdinov, R. R. (2012). Improving neural networks by preventing co-adaptation of feature detectors. arXiv preprint arXiv:1207.0580. (Year: 2012). [cited by examiner]
Chen et al., “Data Mining Method for Electric Power Information Network Differentiated Intrusion,” East China Electric Power, Dec. 2014, 42(12):2672-2675. [cited by applicant]
Chen et al., “Method for Neural Network Ensemble Based on Analytical Hierarchy Process,” Journal of University of Electronic Science and Technology of China, May 2008, 37(3):432-435. [cited by applicant]
Engelbrecht et al. “Automatic Scaling using gamma learning for Feedforward Neural Networks” IWANN 1995 pp. 374-381 [Published online 2005] [Retrieved Online Oct. 22, 2018] <URL:https://link.springer.conn/content/pdf/10.… [cited by applicant]
Extended European Search Report in European Appln. No. 21161358.3, mailed on Jul. 30, 2021, 10 pages. [cited by applicant]
Extended European Search Report issued in European Appln. No. 18207898.0, mailed on Apr. 3, 2019, 8 pages. [cited by applicant]
Gülçehre et al., “Knowledge Matters: Importance of Prior Information for Optimization,” CoRR, Submitted on Jul. 13, 2013, arXiv:1301.4083v6, pp. 1-37. [cited by applicant]
International Preliminary Report on Patentability issued in International Appln. No. PCT/US2016/015476, mailed on Aug. 1, 2017, 8 pages. [cited by applicant]
International Search Report and Written Opinion in International Appln. No. PCT/US2016/015476, mailed May 4, 2016, 13 pages. [cited by applicant]
Ioffe et al., “Batch Normalization: Accelerating Deep Network Training by Reducing Internal Covariate Shift,” CoRR, Submitted on Mar. 2, 2015, arXiv:1502.03167v3, pp. 1-11. [cited by applicant]
Jiang, “A literature survey on domain adaptation of statistical classifiers,” Mar. 2008 [retrieved on Jun. 6, 2016]. Retrieved from the Internet: URL<http://sifaka.cs.uiuc.edu/jiang4/domainadaptation/survey>, pp. 1-12. [cited by applicant]
Kavukcuoglu et al. “Learning Convolutional Feature Hierarchies for Visual Recognition” NIPS '10 Proceedings vol. 1 pp. 1090-1098 [Published Online 2010] [Retrieved online Oct. 22, 2018] <URL:https://papers.nips.cc/paper… [cited by applicant]
Krizhevsky et al. “ImageNet Classification with Deep Convolutional Neural Networks” NIPS '12 Proceedings vol. 1 p. 1097-1105 Published Online 2012] [retrieved online Oct. 22, 2018] <URL:https://papers.nips.cc/paper/4824… [cited by applicant]
LeCun et al. “Efficient BackProp” yann.lecun.conn [Published Online 2012] [Retrieved online Oct. 22, 2018] <URL:http://yann.lecun.conn/exdb/publis/pdf/lecun-98b.pdf> (Year: 2012). [cited by applicant]
LeCun et al., “Efficient BackProp,” Jan. 1, 1901, Correct System Design; [Lecture Notes in Computer Science; Lect.Notes Computer], Springer International Publishing, Cham, pp. 9-48, XP047292571. [cited by applicant]
Lin et al. “Network in Network” National University of Singapore [Published online Mar. 4, 2014] [Retrieved online Oct. 22, 2018] URL: https://arxiv.org/pdf/1312.4400.pdf> (Year: 2014). [cited by applicant]
Lyu and Simoncelli, “Nonlinear image representation using divisive normalization,” In Proc. Computer Vision and Pattern Recognition, IEEE Computer Society, pp. 1-8, Jun. 2008. [cited by applicant]
Notice of Allowance in Japanese Appln. No. 2018-232445, mailed on Apr. 27, 2020, 5 pages (with English translation). [cited by applicant]
Notice of Allowance in Singapore Appln. No. 11201706127R, mailed on Mar. 27, 2018, 7 pages. [cited by applicant]
Office Action in Australian Appln. No. 2019200309, mailed on Jun. 24, 2020, 3 pages. [cited by applicant]
Office Action in Brazilian Appln. No. 11201706306-3, mailed on Jul. 30, 2020, 8 pages (with English translation). [cited by applicant]
Office Action in Canadian Appln No. 2,975,251, mailed on May 30, 2019, 5 pages. [cited by applicant]
Office Action in Canadian Appln. No. 2975251, mailed on Jun. 8, 2018, 5 pages. [cited by applicant]
Office Action in Chinese Appln. No. 201680012517, mailed on Mar. 19, 2020, 11 pages (with English translation). [cited by applicant]
Office Action in Chinese Appln. No. 201680012517.X, mailed on Jul. 2, 2024, 14 pages (with English translation). [cited by applicant]
Office Action in Indian Appln. No. 201747026857, mailed on Aug. 17, 2020, 7 pages (with English translation). [cited by applicant]
Office Action in Korean Appln. No. 10-2019-7036115, mailed on Jul. 14, 2020, 6 pages (with English translation). [cited by applicant]
Office Action in Russian Appln. No. 2017130151, mailed on Jun. 27, 2018, 25 pages (with English translation). [cited by applicant]
Office Action issued in Australian Application No. 2016211333, mailed on Feb. 19, 2018, 2 pages. [cited by applicant]
Povey et al., “Parallel training of deep neural networks with natural gradient and parameter averaging,” CoRR, abs/1410.7455, pp. 1-28, Oct. 2014. [cited by applicant]
Raiko et al., “Deep learning made easier by linear transformations in perceptrons,” In International Conference on Artificial Intelligence and Statistics (AISTATS), pp. 924-932, 2012. [cited by applicant]
Sainath et al. Learning Filter Banks within a Deep Neural Network Framework 2013 IEEE workshop on Automatic Speech Recognition and Understanding [Published Online 2014] [Retrieved Online Oct. 22, 2018] <URL:https://ieee… [cited by applicant]
Szegedy et al. “Intriguing properties of neural networks” Cornell University Library [Published Online 2014] [Retrieved online Oct. 22, 2018] <URL: https://arxiv.org/pdf/1312.6199.pdf> (Year: 2014). [cited by applicant]
Wiesler et al., “A convergence analysis of log-linear training,” In Advances in Neural Information Processing Systems 24, pp. 657-665, Dec. 2011. [cited by applicant]
Wiesler et al., “Mean-normalized stochastic gradient for large-scale deep learning,” In IEEE International Conference on Acoustics, Speech, and Signal Processing, pp. 180-184, May 2014. [cited by applicant]
Office Action in Japanese Appln. No. 2024-131114, mailed on May 27, 2025, 6 pages (with English translation). [cited by applicant]
Extended European Search Report in European Appln. No. 24199485.4, mailed on Dec. 19, 2024, 9 pages. [cited by applicant]
Office Action in Chinese Appln. No. 201680012517.X, mailed on Dec. 1, 2024, 6 pages (with English translation). [cited by applicant]
Office Action in Australian Appln. No. 2023285952, mailed on Mar. 21, 2025, 3 pages. [cited by applicant]
Notice of Allowance in Japanese Appln. No. 2024-131114, mailed on Jan. 27, 2026, 6 pages (with machine translation). [cited by applicant]
Office Action in Chinese Appln. No. 202510128788.2, mailed on Feb. 10, 2026, 8 pages (with English translation). [cited by applicant]
Office Action in Indian Appln. No. 202348087215, mailed on Mar. 13, 2026, 7 pages. [cited by applicant]
Office Action in Indian Appln. No. 202348087290, mailed on Mar. 13, 2026, 7 pages. [cited by applicant]
Office Action in Indian Appln. No. 202348087540, mailed on Mar. 13, 2026, 7 pages. [cited by applicant]
Office Action in Indian Appln. No. 202348087675, mailed on Mar. 13, 2026, 7 pages. [cited by applicant]
Office Action in Chinese Appln. No. 202510129770.4, mailed on Mar. 7, 2026, 16 pages (with English translation). [cited by applicant]