IP Library Granted Patent US 12,437,528
Granted Patent B2
US 12,437,528 · App. 18/199,896 · Granted Oct 7, 2025

Training a speaker neural network using one or more listener neural networks

Inventors: Aaditya K. Singh (McLean, VA); Fengning Ding (London, GB); Felix George Hill (London, GB); Andrew Kyle Lampinen (Palo Alto, CA)
Assignee: DeepMind Technologies Limited
G06V10/82G06V20/635
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,437,528
App. No.
18/199,896
Granted
Oct 7, 2025
Kind
B2
Abstract

Methods, systems, and apparatus, including computer programs encoded on computer storage media, for training a speaker neural network using one or more listener neural networks.

Claims (61)

1. A method performed by one or more computers, the method comprising:

obtaining a set of one or more training images;

for each training image in the set:

processing the training image using a speaker neural network to generate a text caption for the training image;

generating a plurality of listener inputs, each listener input comprising (i) the text caption for the training image and (ii) a respective set of images, wherein the respective set of images includes a corresponding version of the training image and a corresponding set of one or more distractor images that are each different from the training image;

for each listener input, processing the listener input using a respective listener neural network from a set of one or more listener neural networks to generate a respective match score for each image in the respective set of images in the listener input; and

generating a reward for the training image based at least in part on the respective match scores for the training image generated by processing each of the listener inputs; and

training the speaker neural network using the rewards for the training images.

2. The method of claim 1 , wherein the speaker neural network comprises:

a vision encoder neural network configured to process the training image to generate an encoded representation of the training image;

an adapter neural network configured to process the encoded representation to generate an adapted representation of the training image; and

a language neural network configured to process the adapted representation to generate the text caption of the training image.

3. The method of claim 2 , wherein the adapter neural network is configured to apply one or more QKV-attention layers over the encoded representation to generate the adapted representation.

4. The method of claim 2 , wherein training the speaker neural network using the rewards for the training images comprises:

updating the adapter neural network while holding the vision encoder neural network and the language neural network fixed.

5. The method of claim 4 , further comprising, prior to the training:

pre-training the speaker neural network on an image captioning data set through supervised learning to pre-train the adapter neural network while holding the vision encoder neural network and the language neural network fixed.

6. The method of claim 4 , wherein the language neural network has been pre-trained on one or more language modeling tasks.

7. The method of claim 4 , wherein the vision encoder neural network has been pre-trained on one or more representation learning tasks.

8. The method of claim 1 , wherein the corresponding set of distractor images is the same set of distractor images for all of the listener inputs.

9. The method of claim 1 , wherein generating a reward for the training image based at least in part on the respective match scores for the training image generated by processing each of the listener inputs comprises:

generating a respective accuracy score for each listener input based on whether the training image was assigned a highest match score of any image in the respective set of images in the listener input; and

combining the respective accuracy scores to generate the reward.

10. The method of claim 1 , wherein generating a reward for the training image based at least in part on the respective match scores for the training image generated by processing each of the listener inputs comprises:

determining whether only a designated subset of the listener inputs assigned a highest match score of any image in the respective set of images in the listener input to the training image; and

assigning a maximum reward for the training image only in response to determining that only the designated subset of the listener inputs assigned a highest match score of any image in the respective set of images in the listener input to the training image.

11. The method of claim 1 , wherein the set of one or more listener neural networks includes only one listener neural network.

12. The method of claim 1 , wherein the set of one or more listener neural networks includes a plurality of different listener neural networks, wherein each listener input is processed by a different one of the different listener neural networks, and wherein the corresponding version of the training image in each of the training inputs is the same for all of the training inputs.

13. The method of claim 12 , wherein one of the different listener neural networks has been trained on training data that reflects preferences of a particular user.

14. The method of claim 12 , further comprising:

after the training, determining, based on test text captions generated by the speaker neural network by processing test images, a level of bias exhibited by one or more of the plurality of different listener neural networks.

15. The method of claim 1 , wherein, for a first listener input of the plurality of listener inputs, the corresponding version of the training image is the training image.

16. The method of claim 1 , wherein, for at least one of the listener inputs of the plurality of listener inputs, the corresponding version of the training image is the training image after a corresponding image transformation has been applied to the training image.

17. The method of claim 16 , wherein the corresponding image transformation comprises one or more of:

a transformation that changes one or more color values in the training image;

a transformation that crops a portion of the training image;

a transformation that rotates at least a portion of the training image; or

a transformation that blurs at least a portion of the training image.

18. The method of claim 1 , wherein training the speaker neural network using the rewards for the training images comprises:

training the speaker neural network using the rewards through reinforcement learning.

19. The method of claim 18 , wherein training the speaker neural network using the rewards through reinforcement learning comprises:

training the speaker neural network using a policy gradient reinforcement learning technique.

20. The method of claim 19 , wherein the policy gradient reinforcement learning technique is REINFORCE.

21. A system comprising:

one or more computers; and

one or more storage devices storing instructions that, when executed by one or more computers, cause the one or more computers to perform operations comprising:

obtaining a set of one or more training images;

for each training image in the set:

processing the training image using a speaker neural network to generate a text caption for the training image;

generating a plurality of listener inputs, each listener input comprising (i) the text caption for the training image and (ii) a respective set of images, wherein the respective set of images includes a corresponding version of the training image and a corresponding set of one or more distractor images that are each different from the training image;

for each listener input, processing the listener input using a respective listener neural network from a set of one or more listener neural networks to generate a respective match score for each image in the respective set of images in the listener input; and

generating a reward for the training image based at least in part on the respective match scores for the training image generated by processing each of the listener inputs; and

training the speaker neural network using the rewards for the training images.

22. One or more non-transitory computer-readable storage media storing instructions that when executed by one or more computers cause the one or more computers to perform operations comprising:

obtaining a set of one or more training images;

for each training image in the set:

processing the training image using a speaker neural network to generate a text caption for the training image;

generating a plurality of listener inputs, each listener input comprising (i) the text caption for the training image and (ii) a respective set of images, wherein the respective set of images includes a corresponding version of the training image and a corresponding set of one or more distractor images that are each different from the training image;

for each listener input, processing the listener input using a respective listener neural network from a set of one or more listener neural networks to generate a respective match score for each image in the respective set of images in the listener input; and

generating a reward for the training image based at least in part on the respective match scores for the training image generated by processing each of the listener inputs; and

training the speaker neural network using the rewards for the training images.

Assignments (2)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jun 6, 2025
From: DEEPMIND TECHNOLOGIES LIMITED
To: GDM HOLDING LLC
Reel/Frame 071498/0210 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Feb 20, 2024
From: SINGH, AADITYA K.; DING, FENGNING; HILL, FELIX GEORGE; LAMPINEN, ANDREW KYLE
To: DEEPMIND TECHNOLOGIES LIMITED
Reel/Frame 066496/0557 →