IP Library Granted Patent US 12,112,254
Granted Patent B1
US 12,112,254 · App. 16/780,843 · Granted Oct 8, 2024

Optimizing loss function during training of network

Inventors: Steven L. Teig (Menlo Park, CA); Eric A. Sather (Palo Alto, CA)
Assignee: PERCEIVE CORPORATION
G06N3/047G06N3/048G06N3/084
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,112,254
App. No.
16/780,843
Granted
Oct 8, 2024
Kind
B1
Abstract

Some embodiments provide a method for training a machine-trained (MT) network. The method uses a set of training inputs to train parameters of the MT network according to an initial loss function. The method uses a set of validation inputs to compute an error measure for the MT network as trained by the first set of training inputs. The method modifies the loss function for subsequent training of the MT network based on the computed error measure. The method uses the set of training inputs to train the parameters of the MT network according to the modified loss function.

Claims (32)

1. A method for training a machine-trained (MT) network, the method comprising:

using a set of training inputs to train parameters of the MT network according to an initial loss function that is a first combination of a plurality of possible loss functions defined by initial values for a plurality of coefficients;

using a set of validation inputs to compute an error measure for the MT network as trained by the first set of training inputs;

modifying the loss function for subsequent training of the MT network based on the error measure computed using the set of validation inputs to generate a modified loss function that is a second combination of the plurality of possible loss functions defined by modified values for the plurality of coefficients, wherein the plurality of coefficients are continuously differentiable with respect to a description length score that accounts for (i) an amount of information required to modify the loss function and (ii) improvements to predictiveness of the MT network based on modifications to the loss function; and

using the set of training inputs to train the parameters of the MT network according to the loss function as modified based on the error measure computed using the set of validation inputs.

2. The method of claim 1 , wherein the initial loss function comprises the plurality of possible loss functions multiplied by the initial values for the plurality of coefficients and the loss function as modified based on the error measure computed using the set of validation inputs comprises the plurality of possible loss functions multiplied by the modified values for the plurality of coefficients.

3. The method of claim 1 , wherein a first one of the possible loss functions is a logarithmic function and a second one of the possible loss functions is a polynomial function.

4. The method of claim 1 , wherein the plurality of possible loss functions comprises a set of basis functions.

5. The method of claim 1 , wherein using the set of training inputs to train the parameters of the MT network according to the loss function as modified based on the error measure computed using the set of validation inputs comprises using the set of training inputs along with a subset of the set of validation inputs.

6. The method of claim 1 further comprising modifying at least one hyperparameter, separate from the loss function, that defines how the MT network is trained based on the error measure computed using the set of validation inputs.

7. The method of claim 6 , wherein:

the at least one hyperparameter comprises a plurality of hyperparameters separate from the loss function that are also continuously differentiable with respect to the description length score; and

the description length score accounts for improvements to predictiveness of the MT network based on modifications to the plurality of hyperparameters.

8. The method of claim 1 , wherein the MT network is a neural network.

9. The method of claim 1 , wherein the initial loss function is a first type of loss function defined by the initial values for all but a first one of the coefficients being set to zero while the modified loss function is a second type of loss function defined by the modified values for all but a second one of the coefficients being set to zero.

10. The method of claim 1 , wherein the combinations of the plurality of possible loss functions defined by values for the plurality of coefficients enables construction of any differentiable function as a loss function.

11. A non-transitory machine-readable medium storing a program which when executed by at least one processing unit trains a machine-trained (MT) network, the program comprising sets of instructions for:

using a set of training inputs to train parameters of the MT network according to an initial loss function that is a first combination of a plurality of possible loss functions defined by initial values for a plurality of coefficients;

using a set of validation inputs to compute an error measure for the MT network as trained by the first set of training inputs;

modifying the loss function for subsequent training of the MT network based on the error measure computed using the set of validation inputs to generate a modified loss function that is a second combination of the plurality of possible loss functions defined by modified values for the plurality of coefficients, wherein the plurality of coefficients are continuously differentiable with respect to a description length score that accounts for (i) an amount of information required to modify the loss function and (ii) improvements to predictiveness of the MT network based on modifications to the loss function, wherein the modification of the loss function minimizes the description length score by trading off modifications to the loss function with improvements to predictiveness of the MT network; and

using the set of training inputs to train the parameters of the MT network according to the loss function as modified based on the error measure computed using the set of validation inputs.

12. The non-transitory machine-readable medium of claim 11 , wherein the initial loss function comprises the plurality of possible loss functions multiplied by the initial values for the plurality coefficients and the loss function as modified based on the error measure computed using the set of validation inputs comprises the plurality of possible loss functions multiplied by the modified values for the plurality of coefficients.

13. The non-transitory machine-readable medium of claim 11 , wherein a first one of the possible loss functions is a logarithmic function and a second one of the possible loss functions is a polynomial function.

14. The non-transitory machine-readable medium of claim 11 , wherein the plurality of possible loss functions comprises a set of basis functions.

15. The non-transitory machine-readable medium of claim 11 , wherein the set of instructions for using the set of training inputs to train the parameters of the MT network according to the loss function as modified based on the error measure computed using the set of validation inputs comprises a set of instructions for using the set of training inputs along with a subset of the set of validation inputs.

16. The non-transitory machine-readable medium of claim 11 wherein the program further comprises a set of instructions for modifying at least one hyperparameter, separate from the loss function, that defines how the MT network is trained based on the error measure computed using the set of validation inputs.

17. The non-transitory machine-readable medium of claim 16 , wherein:

the at least one hyperparameter comprises a plurality of hyperparameters separate from the loss function that are also continuously differentiable with respect to the description length score; and

the description length score accounts for improvements to predictiveness of the MT network based on modifications to the plurality of hyperparameters.

18. The non-transitory machine-readable medium of claim 11 , wherein the MT network is a neural network.

19. The non-transitory machine-readable medium of claim 11 , wherein the initial loss function is a first type of loss function defined by the initial values for all but a first one of the coefficients being set to zero while the modified loss function is a second type of loss function defined by the modified values for all but a second one of the coefficients being set to zero.

20. The non-transitory machine-readable medium of claim 11 , wherein the combinations of the plurality of possible loss functions defined by values for the plurality of coefficients enables construction of any differentiable function as a loss function.

Assignments (3)
BILL OF SALE Recorded Oct 31, 2024
From: AMAZON.COM SERVICES LLC
To: AMAZON TECHNOLOGIES, INC.
Reel/Frame 069288/0490 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Oct 31, 2024
From: PERCEIVE CORPORATION
To: AMAZON.COM SERVICES LLC
Reel/Frame 069288/0731 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Feb 5, 2020
From: TEIG, STEVEN L.; SATHER, ERIC A.
To: PERCEIVE CORPORATION
Reel/Frame 051729/0666 →
Continuity (3)
Continuation In Part 16453622 · Jun 26, 2019
Provisional Application 62913707 · Oct 10, 2019
Provisional Application 62838629 · Apr 25, 2019
Cited By (2)
US 12,619,879 US 12,718,097