IP Library › Granted Patent US 10,373,049
Granted Patent B2
US 10,373,049 · App. 15/385,642 · Granted Aug 6, 2019

Generating an output for a neural network output layer

Inventor: Reginald Clifford Young (Palo Alto, CA)
Assignee: Google LLC
G06N3/0454G06N3/04G06N3/063G06N7/005
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 10,373,049
App. No.
15/385,642
Filed
Dec 20, 2016
Granted
Aug 6, 2019
Kind
B2
Art Unit
2122
USPC
706/27
Abstract

Systems, methods, and apparatus, including computer programs encoded on a computer storage medium for processing a network input through a neural network having one or more initial neural network layers followed by a softmax output layer. In one aspect, the methods include obtaining a layer output generated by the one or more initial neural network layers and processing the layer output through the softmax output layer to generate a neural network output. Processing the layer output through the softmax output layer includes determining, for each possible output value, a number of occurrences in the layer output values; for each possible output value occurring in the layer output values, determining a respective exponentiation measure; determining a normalization factor for the layer output by combining the exponentiation measures in accordance with the number of occurrences of the possible output values; and determining, for each of layer output values, a softmax probability value.

Claims (63)

1. A method of processing a network input through a neural network having one or more initial neural network layers followed by a softmax output layer to generate a neural network output for the network input, wherein the initial neural network layers are implemented by a processing system that performs computations specified by the one or more initial neural network layers using quantized arithmetic such that output values generated by the processing system can take only values from a predetermined finite set of values, the method comprising:

precomputing a respective exponentiation measure for each possible output value of a predetermined finite set of possible output values that can be included in layer outputs generated by processing network inputs through the one or more initial neural network layers using the processing system;

storing the respective precomputed exponentiation measures;

obtaining a first layer output generated by the processing system by processing a first network input through the one or more initial neural network layers,

the first layer output having a plurality of first layer output values, and

each first layer output value being a respective one of the predetermined finite set of possible output values; and

processing the layer output through the softmax output layer to generate a first neural network output for the network input, comprising:

determining, for each possible output value in the predetermined finite set of possible output values, a count of occurrences of the possible output value among the plurality of first layer output values that are in the first layer output;

for each possible output value that occurs at least once in the in the plurality of layer output values, accessing the stored respective precomputed exponentiation measure for the possible output value, instead of re-computing the respective exponentiation measure;

determining a normalization factor for the layer output by combining the stored precomputed exponentiation measures for the possible output values that occur in the first layer output at least once in accordance with the counts of occurrences of the possible output values; and

determining, for each of the plurality of first layer output values, a softmax probability value from the respective precomputed exponentiation measure for the first layer output value and the normalization factor.

2. The method of claim 1 , wherein obtaining the first layer output comprises:

receiving a plurality of initial layer output values from the processing system, the plurality of initial layer output values being unmapped output values of the one or more initial neural network layers that are each a respective one of the predetermined finite set of values;

obtaining mapping data defining a mapping from the predetermined finite set of values to the predetermined finite set of possible layer output values; and

determining, for each initial layer output value, a corresponding first layer output value based on the mapping data.

3. The method of claim 2 , wherein the mapping data specifies a scaling factor for scaling each of the plurality of initial layer output values to generate the layer output values.

4. The method of claim 1 , wherein each of the finite set of possible output values map to a respective value of an integer data type.

5. The method of claim 1 , further comprising:

generating the first network input by converting one or more floating point values to fixed point values.

6. The method of claim 1 ,

wherein determining the respective precomputed exponentiation measure for each possible output value comprises exponentiating Euler's number by a multiplication of the possible output value.

7. The method of claim 1 , wherein each softmax probability value is determined by dividing each respective precomputed exponentiation measure by the normalization factor.

8. The method of claim 1 , wherein each of the finite set of possible output values is an output of a mapping function.

9. The method of claim 1 , wherein each of the finite set of possible output values is an output of a compression function.

10. A system comprising one or more computers and one or more storage devices storing instructions that are operable, when executed by the one or more computers, to cause the one or more computers to perform operations for processing a network input through a neural network having one or more initial neural network layers followed by a softmax output layer to generate a neural network output for the network input, wherein the initial neural network layers are implemented by a processing system that performs computations specified by the one or more initial neural network layers using quantized arithmetic such that output values generated by the processing system can take only values from a predetermined finite set of values, the operations comprising:

precomputing a respective exponentiation measure for each possible output value of a predetermined finite set of possible output values that can be included in layer outputs generated by processing network inputs through the one or more initial neural network layers using the processing system;

storing the respective precomputed exponentiation measures;

obtaining a first layer output generated by the processing system processing a first network input through the one or more initial neural network layers,

the first layer output having a plurality of first layer output values, and

each first layer output value being a respective one of the predetermined finite set of possible output values; and

processing the layer output through the softmax output layer to generate a first neural network output for the network input, comprising:

determining, for each possible output value in the predetermined finite set of possible output values, a count of occurrences of the possible output value among the plurality of first layer output values that are in the first layer output;

for each possible output value that occurs at least once in the in the plurality of layer output values, accessing the stored precomputed exponentiation measure for the possible output value, instead of re-computing the respective exponentiation measure;

determining a normalization factor for the layer output by combining the stored precomputed exponentiation measures for the possible output values that occur in the first layer output at least once in accordance with the counts of occurrences of the possible output values; and

determining, for each of the plurality of first layer output values, a softmax probability value from the respective precomputed exponentiation measure for the first layer output value and the normalization factor.

11. The system of claim 10 , wherein obtaining the first layer output comprises:

receiving a plurality of initial layer output values from the processing system, the plurality of initial layer output values being unmapped output values of the one or more initial neural network layers that are each a respective one of the predetermined finite set of values;

obtaining mapping data defining a mapping from the predetermined finite set of values to the predetermined finite set of possible layer output values; and

determining, for each initial layer output value, a corresponding first layer output value based on the mapping data.

12. The system of claim 11 , wherein the mapping data specifies a scaling factor for scaling each of the plurality of initial layer output values to generate the layer output values.

13. The system of claim 10 , wherein each of the finite set of possible output values map to a respective value of an integer data type.

14. The system of claim 10 , further comprising:

generating the first network input by converting one or more floating point values to fixed point values.

15. The system of claim 10 ,

wherein determining the respective precomputed exponentiation measure for each possible output value comprises exponentiating Euler's number by a multiplication of the possible output value.

16. A non-transitory computer storage medium encoded with instructions that, when executed by one or more computers, cause the one or more computers to perform operations for processing a network input through a neural network having one or more initial neural network layers followed by a softmax output layer to generate a neural network output for the network input, wherein the initial neural network layers are implemented by a processing system that performs computations specified by the one or more initial neural network layers using quantized arithmetic such that output values generated by the processing system can take only values from a predetermined finite set of values, the operations comprising:

precomputing a respective exponentiation measure for each possible output value of a predetermined finite set of possible output values that can be included in layer outputs generated by processing network inputs through the one or more initial neural network layers using the processing system;

storing the respective precomputed exponentiation measures;

obtaining a first layer output generated by the processing system processing a first network input through the one or more initial neural network layers,

the first layer output having a plurality of first layer output values, and

each first layer output value being a respective one of the predetermined finite set of possible output values; and

processing the layer output through the softmax output layer to generate a first neural network output for the network input, comprising:

determining, for each possible output value in the predetermined finite set of possible output values, a count of occurrences of the possible output value among the plurality of first layer output values that are in the first layer output;

for each possible output value that occurs at least once in the in the plurality of layer output values, accessing the stored precomputed exponentiation measure for the possible output value, instead of re-computing the respective exponentiation measure;

determining a normalization factor for the layer output by combining the stored precomputed exponentiation measures for the possible output values that occur in the first layer output at least once in accordance with the counts of occurrences of the possible output values; and

determining, for each of the plurality of first layer output values, a softmax probability value from the respective precomputed exponentiation measure for the first layer output value and the normalization factor.

17. The non-transitory computer storage medium of claim 16 , wherein obtaining the first layer output comprises:

receiving a plurality of initial layer output values from the processing system, the plurality of initial layer output values being unmapped output values of the one or more initial neural network layers that are each a respective one of the predetermined finite set of values;

obtaining mapping data defining a mapping from the predetermined finite set of values to the predetermined finite set of possible layer output values; and

determining, for each initial layer output value, a corresponding first layer output value based on the mapping data.

18. The non-transitory computer storage medium of claim 17 , wherein the mapping data specifies a scaling factor for scaling each of the plurality of initial layer output values to generate the layer output values.

19. The non-transitory computer storage medium of claim 16 , wherein each of the finite set of possible output values map to a respective value of an integer data type.

20. The non-transitory computer storage medium of claim 16 , wherein determining the respective precomputed exponentiation measure for each possible output value comprises exponentiating Euler's number by a multiplication of the possible output value.

Assignments (2)
CHANGE OF NAME Recorded Oct 20, 2017
From: GOOGLE INC.
To: GOOGLE LLC
Reel/Frame 044567/0001 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Dec 20, 2016
From: YOUNG, REGINALD CLIFFORD
To: GOOGLE INC.
Reel/Frame 041038/0441 →
Continuity (1)
Related Publication 20180174022A1 · Jun 21, 2018
Cited By (1)
US 12,327,179