IP Library Granted Patent US 12,020,167
Granted Patent B2
US 12,020,167 · App. 17/051,982 · Granted Jun 25, 2024

Gradient adversarial training of neural networks

Inventors: Ayan Tuhinendu Sinha (San Francisco, CA); Andrew Rabinovich (San Francisco, CA); Zhao Chen (Mountain View, CA); Vijay Badrinarayanan (Mountain View, CA)
Assignee: Magic Leap, Inc.
G06N3/088G06N3/045G06N3/084
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,020,167
App. No.
17/051,982
Filed
Oct 30, 2020
Granted
Jun 25, 2024
Kind
B2
Art Unit
2127
USPC
706/25
Abstract

Systems and methods for gradient adversarial training of a neural network are disclosed. In one aspect of gradient adversarial training, an auxiliary neural network can be trained to classify a gradient tensor that is evaluated during backpropagation in a main neural network that provides a desired task output. The main neural network can serve as an adversary to the auxiliary network in addition to a standard task-based training procedure. The auxiliary neural network can pass an adversarial gradient signal back to the main neural network, which can use this signal to regularize the weight tensors in the main neural network. Gradient adversarial training of the neural network can provide improved gradient tensors in the main network. Gradient adversarial techniques can be used to train multitask networks, knowledge distillation networks, and adversarial defense networks.

Claims (48)

1. A system comprising:

non-transitory memory configured to store:

executable instructions;

a main neural network configured to determine an output associated with a task; and

an auxiliary neural network configured to train a gradient tensor that is used to calculate a plurality of weights in the main neural network; and

a hardware processor in communication with the non-transitory memory, the hardware processor programmed by the executable instructions to:

receive, by the main neural network, training data associated with the task to be performed by the main neural network;

evaluate the gradient tensor during backpropagation in the main neural network;

receive, by the auxiliary neural network, the gradient tensor;

train the auxiliary neural network using an auxiliary loss function;

provide to the main neural network, by the auxiliary neural network, an adversarial gradient signal;

obtain a trained main neural network by updating the weights in the main neural network based at least in part on the gradient tensor and the adversarial gradient signal;

output the trained main neural network;

receive image data from a camera corresponding to the task, wherein the task includes at least one of face recognition, visual search, gesture identification or recognition, semantic segmentation, object detection, room layout estimation, cuboid detection, lighting detection, simultaneous localization and mapping, relocalization of an object or an avatar, speech recognition, or natural language processing;

analyze the image data using the trained main neural network; and

perform the task based on analyzing the image data using the trained main neural network.

2. The system of claim 1 , wherein the training data comprises images, and the task comprises a computer vision task.

3. The system of claim 1 , wherein the main neural network comprises a multitask network, a knowledge distillation network, or an adversarial defense network.

4. The system of claim 1 , wherein to evaluate the gradient tensor, the hardware processor is programmed by the executable instructions to evaluate a gradient of a main loss function with respect to a weight tensor in each layer of the main neural network.

5. The system of claim 1 , wherein to provide to the main neural network the adversarial gradient signal, the hardware processor is programmed to utilize gradient reversal.

6. The system of claim 1 , wherein to provide to the main neural network the adversarial gradient signal, the hardware processor is programmed to determine a signal for a layer in the main neural network that is based at least partly on weights in preceding layers of the main neural network.

7. The system of claim 1 , wherein to update weights in the main neural network, the hardware processor is programmed to regularize the weights in the main neural network based at least in part on the adversarial gradient signal.

8. The system of claim 1 , wherein the main neural network comprises a multitask network, the task comprises a plurality of tasks, and the multitask network comprises:

a shared encoder;

a plurality of task-specific decoders associated with a respective task from the plurality of tasks; and

a plurality of gradient-alignment layers (GALs), each GAL of the plurality of GALs located after the shared encoder and before at least one of the task-specific decoders.

9. The system of claim 8 , wherein the hardware processor is programmed to train each GAL of the plurality of GALs using a reversed gradient signal from the auxiliary neural network.

10. The system of claim 8 , wherein the hardware processor is programmed to train each GAL of the plurality of GALs to make statistical distributions of a plurality of gradient tensors for the plurality of tasks indistinguishable.

11. The system of claim 8 , wherein the plurality of GALs are dropped during forward inference in the multitask network.

12. The system of claim 1 , wherein the main neural network comprises a knowledge distillation network that comprises a student network and a teacher network, and the auxiliary loss function comprises a binary classifier trainable to discriminate between (1) a gradient tensor of the student network and (2) a gradient tensor of the teacher network.

13. The system of claim 1 , wherein the main neural network is configured to analyze images, and the hardware processor is programmed to utilize a modified cross-entropy loss function in which a cross-entropy loss of the main neural network is modified based at least in part on a soft-max function evaluated on an output activation from the auxiliary neural network.

14. The system of claim 1 , wherein the main neural network is configured to analyze images, and the hardware processor is programmed to utilize a modified cross-entropy loss function configured to add weight to negative classes whose gradient tensors are similar to gradient tensors of a primary class.

15. A method for training a neural network comprising a main neural network configured to determine an output associated with a task and an auxiliary neural network configured to train a gradient tensor that is used to calculate a plurality of weights in the main neural network, the method comprising:

receiving, by the main neural network, training data associated with the task to be performed by the main neural network;

evaluating a gradient tensor during backpropagation in the main neural network;

receiving, by the auxiliary neural network, the gradient tensor;

training the auxiliary neural network using an auxiliary loss function;

providing to the main neural network, by the auxiliary neural network, an adversarial gradient signal;

obtaining a trained main neural network by updating the weights in the main neural network based at least in part on the gradient tensor and the adversarial gradient signal;

outputting the trained main neural network corresponding to the task, wherein the task includes at least one of face recognition, visual search, gesture identification or recognition, semantic segmentation, object detection, room layout estimation, cuboid detection, lighting detection, simultaneous localization and mapping, relocalization of an object or an avatar, speech recognition, or natural language processing;

receiving image data from a camera;

analyzing the image data using the trained main neural network; and

performing the task based on the analyzing the image data using the trained main neural network.

16. The method of claim 15 , wherein the training data comprises images, and the task comprises a computer vision task.

17. The method of claim 15 , wherein the main neural network comprises a multitask network, a knowledge distillation network, or an adversarial defense network.

18. The method of claim 15 , wherein evaluating the gradient tensor comprises evaluating a gradient of a main loss function with respect to a weight tensor in each layer of the main neural network.

19. The method of claim 15 , wherein providing the adversarial gradient signal to the main neural network comprises utilizing gradient reversal.

20. The method of claim 15 , wherein providing the adversarial gradient signal comprises determining a signal for a layer in the main neural network that is based at least partly on weights in preceding layers of the main neural network.

Assignments (3)
SECURITY INTEREST Recorded Oct 28, 2025
From: MAGIC LEAP, INC.; MENTOR ACQUISITION ONE, LLC; MOLECULAR IMPRINTS, INC.
To: CITIBANK, N.A., AS COLLATERAL AGENT
Reel/Frame 073388/0027 →
SECURITY INTEREST Recorded Oct 24, 2025
From: MAGIC LEAP, INC.; MENTOR ACQUISITION ONE, LLC; MOLECULAR IMPRINTS, INC.
To: CITIBANK, N.A., AS COLLATERAL AGENT
Reel/Frame 073255/0581 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Dec 17, 2020
From: SINHA, AYAN TUHINENDU; RABINOVICH, ANDREW; CHEN, ZHAO; BADRINARAYANAN, VIJAY
To: MAGIC LEAP, INC.
Reel/Frame 054687/0018 →
Cited By (1)
US 12,688,429