IP Library Granted Patent US 12,217,179
Granted Patent B2
US 12,217,179 · App. 18/372,415 · Granted Feb 4, 2025

Intelligent regularization of neural network architectures

Inventors: Zoubin Ghahramani (Cambridge, GB); Douglas Bemis (San Francisco, CA); Theofanis Karaletsos (San Francisco, CA)
Assignee: Uber Technologies, Inc.
G06N3/08G06N3/045G06N3/082G06N3/0985G06N7/01
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,217,179
App. No.
18/372,415
Granted
Feb 4, 2025
Kind
B2
Abstract

A trained computer model includes a direct network and an indirect network. The indirect network generates expected weights or an expected weight distribution for the nodes and layers of the direct network. These expected characteristics may be used to regularize training of the direct network weights and encourage the direct network weights towards those expected, or predicted by the indirect network. Alternatively, the expected weight distribution may be used to probabilistically predict the output of the direct network according to the likelihood of different weights or weight sets provided by the expected weight distribution. The output may be generated by sampling weight sets from the distribution and evaluating the sampled weight sets.

Claims (46)

1. A non-transitory computer-readable medium storing instructions that, when executed by a computing system, cause the computing system to perform operations comprising:

receiving a set of direct inputs;

providing the set of direct inputs to a direct network having a set of weights, wherein the set of weights were trained by a process including:

setting initial values for the set of weights;

processing training input using the initial values to generate training output;

obtaining distributions of a set of expected weights for the direct network, the distributions of set of expected weights generated by an indirect network using a set of indirect parameters, wherein at least some of the set of indirect parameters for a node describe a location of the node within the direct network by specifying a layer and an index of the node;

identifying an error between an expected output and the training output generated from the direct network; and

updating the set of weights based on the error and the distributions of the set of expected weights; and

generating, by the direct network, a direct output from the direct network, using the set of weights.

2. The non-transitory computer-readable medium of claim 1 , wherein the indirect network applied the set of indirect parameters to a set of indirect control inputs corresponding to a characteristic conditioning a generation of the set of expected weights from the set of indirect parameters.

3. The non-transitory computer-readable medium of claim 2 , wherein the characteristic includes one or more of: a location on an image, or a type of input of the direct network.

4. The non-transitory computer-readable medium of claim 2 , wherein the indirect network is configured to also generate another set of expected weights for another neural network having a different characteristic.

5. The non-transitory computer-readable medium of claim 1 , wherein the operations further comprise updating the set of indirect parameters using a derivative of the error with respect to the set of indirect parameters.

6. The non-transitory computer-readable medium of claim 1 , wherein the set of weights is updated based on an unregulated set of weights and the set of expected weights.

7. The non-transitory computer-readable medium of claim 6 , wherein the set of weights is updated by a linear combination of the unregulated set of weights and the set of expected weights.

8. The non-transitory computer-readable medium of claim 1 , wherein the indirect network is a parametric model.

9. A method comprising:

receiving a set of direct inputs;

providing the set of direct inputs to a direct network having a set of weights, wherein the set of weights were trained by a process including:

setting initial values for the set of weights;

processing training input using the initial values to generate training output;

obtaining distributions of a set of expected weights for the direct network, the distributions of the set of expected weights generated by an indirect network using a set of indirect parameters, wherein at least some of the set of indirect parameters for a node describe a location of the node within the direct network by specifying a layer and an index of the node;

identifying an error between an expected output and the training output generated from the direct network; and

updating the set of weights based on the error and the distributions of the set of expected weights; and

generating, by the direct network, a direct output from the direct network, using the set of weights.

10. The method of claim 9 , wherein the indirect network applied the set of indirect parameters to a set of indirect control inputs corresponding to a characteristic conditioning a generation of the set of expected weights from the set of indirect parameters.

11. The method of claim 10 , wherein the characteristic includes one or more of: a location on an image, or a type of input of the direct network.

12. The method of claim 10 , wherein the indirect network is configured to also generate another set of expected weights for another neural network having a different characteristic.

13. The method of claim 9 , further comprising, updating the set of indirect parameters using a derivative of the error with respect to the set of indirect parameters.

14. The method of claim 9 , wherein the set of weights is updated based on an unregulated set of weights and the set of expected weights.

15. The method of claim 14 , wherein the set of weights is updated by a linear combination of the unregulated set of weights and the set of expected weights.

16. The method of claim 9 , wherein the indirect network is a parametric model.

17. A computing system comprising:

one or more computer processors; and

a non-transitory computer-readable medium storing instructions that, when executed by the computing system, cause the computing system to perform operations comprising:

receiving a set of direct inputs;

providing the set of direct inputs to a direct network having a set of weights, wherein the set of weights were trained by a process including:

setting initial values for the set of weights;

processing training input using the initial values to generate training output;

obtaining distributions of a set of expected weights for the direct network, the distributions of the set of expected weights generated by an indirect network using a set of indirect parameters, wherein at least some of the set of indirect parameters for a node describe a location of the node within the direct network by specifying a layer and an index of the node;

identifying an error between an expected output and the training output generated from the direct network; and

updating the set of weights based on the error and the distributions of the set of expected weights; and

generating, by the direct network, a direct output from the direct network, using the set of weights.

18. The computing system of claim 17 , wherein the indirect network applied the set of indirect parameters to a set of indirect control inputs corresponding to a characteristic conditioning a generation of the set of expected weights from the set of indirect parameters.

19. The computing system of claim 18 , wherein the characteristic includes one or more of: a location on an image, or a type of input of the direct network.

20. The computing system of claim 18 , wherein the indirect network is configured to also generate another set of expected weights for another neural network having a different characteristic.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Dec 14, 2023
From: GHAHRAMANI, ZOUBIN; BEMIS, DOUGLAS; KARALETSOS, THEOFANIS
To: UBER TECHNOLOGIES, INC.
Reel/Frame 065871/0623 →
Continuity (5)
Continuation 17513517 · Oct 28, 2021
Continuation 15789898 · Oct 20, 2017
Provisional Application 62451818 · Jan 30, 2017
Provisional Application 62410393 · Oct 20, 2016
Related Publication 20240013049A1 · Jan 11, 2024
References Cited (2)
US 20200184337A1 · Baker · 2020 [cited by applicant]
United States Office Action, U.S. Appl. No. 15/789,898, filed Mar. 8, 2021, seven pages. [cited by applicant]