IP Library › Granted Patent US 11,790,236
Granted Patent B2
US 11,790,236 · App. 16/809,096 · Granted Oct 17, 2023

Minimum deep learning with gating multiplier

Inventor: Gil Shamir (Sewickley, PA)
Assignee: GOOGLE LLC
G06N3/084G06N3/04
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,790,236
App. No.
16/809,096
Granted
Oct 17, 2023
Kind
B2
Abstract

Systems and methods according to the present disclosure can employ a computer-implemented method for inference using a machine-learned model. The method can be implemented by a computing system having one or more computing devices. The method can include obtaining data descriptive of a neural network including one or more network units and one or more gating paths, wherein each of the gating path(s) includes one or more gating units. The method can include obtaining data descriptive of one or more input features. The method can include determining one or more network unit outputs from the network unit(s) based at least in part on the input feature(s). The method can include determining one or more gating values from the gating path(s). The method can include determining one or more gated network unit outputs based at least in part on a combination of the network unit output(s) and the gating value(s).

Claims (33)

1. A computing system for performing gating-based regularization of a neural network, the computing system comprising:

one or more processors; and

one or more non-transitory computer-readable media that collectively store:

a neural network, the neural network comprising:

a gated network unit, the gated network unit comprising one or more network parameters; and

a gating path associated with the gated network unit, wherein the gating path comprises one or more gating units, wherein each of the one or more gating units comprises one or more gating parameters, wherein the gating path is configured to produce a gating value that represents an overall magnitude of an effect that the gated network unit has on predictions of the neural network without an adjustment by the gating unit;

wherein a gated output of the gated network unit comprises an intermediate output of the gated network unit multiplied by the gating value; and

instructions that, when executed by the one or more processors, cause the computing system to perform operations to train the neural network based on one or more training examples, wherein the operations comprise, for each of the one or more training examples:

determining a gradient of a loss function with respect to at least one of the one or more network parameters and one or more gating parameters; and

updating a respective value of at least one of the one or more network parameters and the one or more gating parameters based on the gradient of the loss function.

2. The computing system of claim 1 , wherein the one or more gating units comprises at least one of a benefit path configured to produce a benefit score associated with the gated network unit, a scaling function configured to produce the gating value based at least in part on the benefit score, or a clipping function configured to clip the gating value based at least in part on a clipping threshold.

3. The computing system of claim 2 , wherein the benefit path comprises at least one of a benefit bias unit or a benefit link weight, wherein the one or more gating parameters comprises a benefit score bias, wherein the benefit bias unit is configured to store the benefit score bias, and wherein the benefit path produces the benefit score based at least in part on the benefit score bias, wherein the benefit path comprises at least one of a stateful benefit bias unit or a stateful benefit link weight.

4. The computing system of claim 2 , wherein the neural network comprises one or more network layers, each of the one or more network layers comprising one or more network units, wherein a first network layer of the one or more network layers comprises the gated network unit, and wherein the benefit path comprises a weighted sum of inputs, wherein the weighted sum of inputs comprises outputs of the one or more network units.

5. The computing system of claim 2 , wherein the benefit path comprises one or more benefit path layers, each of the benefit path layers comprising one or more benefit units.

6. The computing system of claim 5 , wherein at least one of the one or more benefit path layers comprises a bottleneck layer, wherein a dimensionality of the bottleneck layer is less than a dimensionality of a preceding layer of the one or more benefit path layers.

7. The computing system of claim 1 , wherein the neural network comprises a second gated network unit, wherein a second gated output of the second gated network unit comprises a second intermediate output of the second gated network unit multiplied by the gating value.

8. The computing system of claim 1 , wherein at least one of the one or more gated network units and one or more gating units are stateful gating units that are respective to a state of the at least one gated network unit, wherein the at least one gated network unit having the state comprises at least one of an embedding unit, a link weight, or a bias unit.

9. The computing system of claim 1 , wherein the one or more network units comprises an embedding vector, and wherein the gating value is associated with each embedding component in the embedding vector.

10. The computing system of claim 2 , wherein the scaling function comprises one of a sigmoid function, a half-sigmoid function, or shifted piecewise smooth activation function.

11. The computing system of claim 2 , wherein the scaling function comprises a self-gating scaling function.

12. The computing system of claim 1 , wherein the one or more gating parameters are learned during training of the neural network.

13. The computing system of claim 1 , wherein the one or more gating units are employed as an activation function for the gated network unit.

14. A computer-implemented method for performing inference using a machine-learned model, the computer-implemented method comprising:

obtaining, by a computing system comprising one or more computing devices, data descriptive of a neural network comprising:

one or more network units; and

one or more gating paths, each of the one or more gating paths associated with each of the one or more network units, wherein each of the one or more gating paths comprises one or more gating units;

obtaining, by the computing system, data descriptive of one or more input features;

determining, by the computing system, one or more network unit outputs from the one or more network units based at least in part on the one or more input features;

determining, by the computing system, one or more gating values from the one or more gating paths, wherein the one or more gating values represents an overall magnitude of an effect that the one or more network units has on predictions of the neural network without an adjustment by the one or more gating units; and

determining one or more gated network unit outputs based at least in part on a combination of the one or more network unit outputs and the one or more gating values.

15. The computer-implemented method of claim 14 , wherein the one or more gating units comprises at least one of a benefit path configured to produce a benefit score associated with the gated network unit, a scaling function configured to produce the gating value based at least in part on the benefit score, or a clipping function configured to clip the gating value based at least in part on a clipping threshold.

16. The computer-implemented method of claim 15 , wherein the benefit path comprises at least one of a benefit bias unit or a benefit link weight, wherein the one or more gating parameters comprises a benefit score bias, wherein the benefit bias unit is configured to store the benefit score bias, and wherein the benefit path produces the benefit score based at least in part on the benefit score bias, wherein the benefit path comprises at least one of a stateful benefit bias unit or a stateful benefit link weight.

17. The computer-implemented method of claim 15 , wherein the neural network comprises one or more network layers, each of the one or more network layers comprising one or more network units, wherein a first network layer of the one or more network layers comprises the gated network unit, and wherein the benefit path comprises a weighted sum of inputs, wherein the weighted sum of inputs comprises outputs of the one or more network units.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Apr 27, 2020
From: SHAMIR, GIL
To: GOOGLE LLC
Reel/Frame 052506/0539 →
Continuity (1)
Related Publication 20210279591A1 · Sep 9, 2021