IP Library › Granted Patent US 11,062,179
Granted Patent B2
US 11,062,179 · App. 16/179,469 · Granted Jul 13, 2021

Method and device for generative adversarial network training

Inventors: Avishek Bose (Toronto, CA); Yanshuai Cao (Toronto, CA)
Assignee: ROYAL BANK OF CANADA
G06K9/6257G06F16/532G06K9/6262G06K9/6267G06K9/726G06N3/0472G06N3/088
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,062,179
App. No.
16/179,469
Granted
Jul 13, 2021
Kind
B2
Abstract

An electronic device for neural network training includes at least one processor and one or more memories configured to provide or train: a generative adversarial network (GAN) using a generator and a discriminator for: receiving a plurality of training cases; and training the generative adversarial network, based on the plurality of training cases, to classify the training cases; wherein the generator generates hard negative examples for the discriminator.

Claims (46)

1. An electronic device for neural network training comprising:

one or more processors;

a non-transitory computer-readable medium storing one or more programs and data representative of a generative adversarial network (GAN) having a plurality of nodes and weights, the GAN including a generator and a discriminator, wherein the one or more programs are configured to be executed by the one or more processors, the one or more programs including instructions for:

receiving data representative of a plurality of training samples;

generating, by the generator, a plurality of hard negative samples based on the plurality of training samples, wherein the hard negative samples are generated from: samples which are initially misclassified by the GAN, or samples which are classified differently than statistically-similar samples, wherein the generator is based on a fixed noise distribution and a conditional distribution defining a categorical distribution over possible training samples; and

generating an output, by the discriminator, based on the plurality of hard negative samples from the generator.

2. The electronic device of claim 1 , wherein at least one of the generator and the discriminator includes a neural network.

3. The electronic device of claim 2 , wherein the generator is defined for an image or caption retrieval.

4. The electronic device of claim 3 , wherein the generator for the image retrieval is based on:

λ p noise ( i )+(1−λ) g θ ( i|c )

wherein p noise is a fixed noise distribution, g θ is a conditional distribution with learnable parameters θ and λ is a hyperparameter, g θ (i|c) and defines a categorical distribution over all possible images from a set of images in the plurality of training cases.

5. The electronic device of claim 3 , wherein the generator for the caption retrieval is based on:

λ p noise ( c )+(1−λ) g θ ( c|i )

wherein p noise is a fixed noise distribution, g θ is a conditional distribution with learnable parameters θ and λ is a hyperparameter, and g θ (i|c) defines a categorical distribution over all possible images from a set of images including correctly labelled images.

6. The electronic device of claim 2 , wherein an output of the discriminator is used as a feedback to the generator.

7. The electronic device of claim 3 , wherein the a score for an image-caption pair is used as a reward for the generator.

8. The electronic device of claim 7 , wherein the GAN is configured to: receive data representative of an image, and generate an output to identify, from a plurality of text strings, a corresponding text string based on the data representative of the image.

9. The electronic device of claim 7 , wherein the GAN is configured to: receive data representative of a text string, and generate an output to identify, from a plurality of images, a corresponding image based on the data representative of the text string.

10. A computer-implemented method comprising:

receiving, by a generative adversarial network (GAN) having a plurality of nodes and weights, the GAN including a generator and a discriminator, data representative of a plurality of training samples;

generating, by the generator, a plurality of hard negative samples based on the plurality of training samples, wherein the hard negative samples are generated from: samples which are initially misclassified by the GAN, or samples which are classified differently than statistically-similar samples, wherein the generator is based on a fixed noise distribution and a conditional distribution defining a categorical distribution over possible training samples; and

generating an output, by the discriminator, based on the plurality of hard negative samples from the generator.

11. The method of claim 10 , wherein at least one of the generator and the discriminator includes a neural network.

12. The method of claim 11 , wherein the generator for the image retrieval is based on:

λ p noise ( i )+(1−λ) g θ ( i|c )

wherein p noise is a fixed noise distribution, g θ is a conditional distribution with learnable parameters θ and λ is a hyperparameter, and g θ (i|c) defines a categorical distribution over all possible images from a set of images in the plurality of training cases.

13. The method of claim 11 , wherein the generator for the caption retrieval is based on:

λ p noise ( c )+(1−λ) g θ ( c|i )

wherein p noise is a fixed noise distribution, g θ is a conditional distribution with learnable parameters θ and λ is a hyperparameter, g θ (i|c) and defines a categorical distribution over all possible images from a set of images including correctly labelled images.

14. The method of claim 10 , wherein an output of the discriminator is used as a feedback to the generator.

15. The method of claim 11 , wherein the a score for an image-caption pair is used as a reward for the generator.

16. The method of claim 15 , wherein the GAN is configured to: receive data representative of an image, and generate an output to identify, from a plurality of text strings, a corresponding text string based on the data representative of the image.

17. The method of claim 15 , wherein the GAN is configured to: receive data representative of a text string, and generate an output to identify, from a plurality of images, a corresponding image based on the data representative of the text string.

18. An electronic device comprising:

one or more processors; and

a memory storing one or more program configured to be executed by the one or more processors, the one or more programs including instructions for:

receiving a text string or an image;

processing the text string or the image using a generative adversarial network including a generator and a discriminator; and

choosing an matched image based on the processed text string from a plurality of images, or choosing a matched text string based on the processed image from a plurality of text strings;

wherein during a training phase, the generator is configured to generate a plurality of hard negatives samples for the discriminator, wherein the hard negative samples are generated from: samples which are initially misclassified by the GAN, or samples which are classified differently than statistically-similar samples, wherein the generator is based on a fixed noise distribution and a conditional distribution defining a categorical distribution over possible training samples.

19. The method of claim 18 , wherein the generator is based on:

λ p noise ( i )+(1−λ) g θ ( i|c )

wherein p noise is a fixed noise distribution, g θ is a conditional distribution with learnable parameters θ and λ is a hyperparameter, and g θ (i|c) defines a categorical distribution over all possible images from a set of images in a plurality of training cases.

20. The method of claim 18 , wherein the generator for is based on:

λ p noise ( c )+(1−λ) g θ ( c|i )

wherein p noise is a fixed noise distribution, g θ is a conditional distribution with learnable parameters θ and λ is a hyperparameter, g θ (i|c) and defines a categorical distribution over all possible images from a set of images including correctly labelled images.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Feb 18, 2020
From: BOSE, AVISHEK; CAO, YANSHUAI
To: ROYAL BANK OF CANADA
Reel/Frame 051841/0562 →
Continuity (2)
Provisional Application 62580821 · Nov 2, 2017
Related Publication 20190130221A1 · May 2, 2019
Cited By (1)
US 12,632,541