IP Library Granted Patent US 12,437,528
Granted Patent B2
US 12,437,528 · App. 18/199,896 · Granted Oct 7, 2025

Training a speaker neural network using one or more listener neural networks

Inventors: Aaditya K. Singh (McLean, VA); Fengning Ding (London, GB); Felix George Hill (London, GB); Andrew Kyle Lampinen (Palo Alto, CA)
Assignee: DeepMind Technologies Limited
G06V10/82G06V20/635
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,437,528
App. No.
18/199,896
Granted
Oct 7, 2025
Kind
B2
Abstract

Methods, systems, and apparatus, including computer programs encoded on computer storage media, for training a speaker neural network using one or more listener neural networks.

Claims (61)

1. A method performed by one or more computers, the method comprising:

obtaining a set of one or more training images;

for each training image in the set:

processing the training image using a speaker neural network to generate a text caption for the training image;

generating a plurality of listener inputs, each listener input comprising (i) the text caption for the training image and (ii) a respective set of images, wherein the respective set of images includes a corresponding version of the training image and a corresponding set of one or more distractor images that are each different from the training image;

for each listener input, processing the listener input using a respective listener neural network from a set of one or more listener neural networks to generate a respective match score for each image in the respective set of images in the listener input; and

generating a reward for the training image based at least in part on the respective match scores for the training image generated by processing each of the listener inputs; and

training the speaker neural network using the rewards for the training images.

2. The method of claim 1 , wherein the speaker neural network comprises:

a vision encoder neural network configured to process the training image to generate an encoded representation of the training image;

an adapter neural network configured to process the encoded representation to generate an adapted representation of the training image; and

a language neural network configured to process the adapted representation to generate the text caption of the training image.

3. The method of claim 2 , wherein the adapter neural network is configured to apply one or more QKV-attention layers over the encoded representation to generate the adapted representation.

4. The method of claim 2 , wherein training the speaker neural network using the rewards for the training images comprises:

updating the adapter neural network while holding the vision encoder neural network and the language neural network fixed.

5. The method of claim 4 , further comprising, prior to the training:

pre-training the speaker neural network on an image captioning data set through supervised learning to pre-train the adapter neural network while holding the vision encoder neural network and the language neural network fixed.

6. The method of claim 4 , wherein the language neural network has been pre-trained on one or more language modeling tasks.

7. The method of claim 4 , wherein the vision encoder neural network has been pre-trained on one or more representation learning tasks.

8. The method of claim 1 , wherein the corresponding set of distractor images is the same set of distractor images for all of the listener inputs.

9. The method of claim 1 , wherein generating a reward for the training image based at least in part on the respective match scores for the training image generated by processing each of the listener inputs comprises:

generating a respective accuracy score for each listener input based on whether the training image was assigned a highest match score of any image in the respective set of images in the listener input; and

combining the respective accuracy scores to generate the reward.

10. The method of claim 1 , wherein generating a reward for the training image based at least in part on the respective match scores for the training image generated by processing each of the listener inputs comprises:

determining whether only a designated subset of the listener inputs assigned a highest match score of any image in the respective set of images in the listener input to the training image; and

assigning a maximum reward for the training image only in response to determining that only the designated subset of the listener inputs assigned a highest match score of any image in the respective set of images in the listener input to the training image.

11. The method of claim 1 , wherein the set of one or more listener neural networks includes only one listener neural network.

12. The method of claim 1 , wherein the set of one or more listener neural networks includes a plurality of different listener neural networks, wherein each listener input is processed by a different one of the different listener neural networks, and wherein the corresponding version of the training image in each of the training inputs is the same for all of the training inputs.

13. The method of claim 12 , wherein one of the different listener neural networks has been trained on training data that reflects preferences of a particular user.

14. The method of claim 12 , further comprising:

after the training, determining, based on test text captions generated by the speaker neural network by processing test images, a level of bias exhibited by one or more of the plurality of different listener neural networks.

15. The method of claim 1 , wherein, for a first listener input of the plurality of listener inputs, the corresponding version of the training image is the training image.

16. The method of claim 1 , wherein, for at least one of the listener inputs of the plurality of listener inputs, the corresponding version of the training image is the training image after a corresponding image transformation has been applied to the training image.

17. The method of claim 16 , wherein the corresponding image transformation comprises one or more of:

a transformation that changes one or more color values in the training image;

a transformation that crops a portion of the training image;

a transformation that rotates at least a portion of the training image; or

a transformation that blurs at least a portion of the training image.

18. The method of claim 1 , wherein training the speaker neural network using the rewards for the training images comprises:

training the speaker neural network using the rewards through reinforcement learning.

19. The method of claim 18 , wherein training the speaker neural network using the rewards through reinforcement learning comprises:

training the speaker neural network using a policy gradient reinforcement learning technique.

20. The method of claim 19 , wherein the policy gradient reinforcement learning technique is REINFORCE.

21. A system comprising:

one or more computers; and

one or more storage devices storing instructions that, when executed by one or more computers, cause the one or more computers to perform operations comprising:

obtaining a set of one or more training images;

for each training image in the set:

processing the training image using a speaker neural network to generate a text caption for the training image;

generating a plurality of listener inputs, each listener input comprising (i) the text caption for the training image and (ii) a respective set of images, wherein the respective set of images includes a corresponding version of the training image and a corresponding set of one or more distractor images that are each different from the training image;

for each listener input, processing the listener input using a respective listener neural network from a set of one or more listener neural networks to generate a respective match score for each image in the respective set of images in the listener input; and

generating a reward for the training image based at least in part on the respective match scores for the training image generated by processing each of the listener inputs; and

training the speaker neural network using the rewards for the training images.

22. One or more non-transitory computer-readable storage media storing instructions that when executed by one or more computers cause the one or more computers to perform operations comprising:

obtaining a set of one or more training images;

for each training image in the set:

processing the training image using a speaker neural network to generate a text caption for the training image;

generating a plurality of listener inputs, each listener input comprising (i) the text caption for the training image and (ii) a respective set of images, wherein the respective set of images includes a corresponding version of the training image and a corresponding set of one or more distractor images that are each different from the training image;

for each listener input, processing the listener input using a respective listener neural network from a set of one or more listener neural networks to generate a respective match score for each image in the respective set of images in the listener input; and

generating a reward for the training image based at least in part on the respective match scores for the training image generated by processing each of the listener inputs; and

training the speaker neural network using the rewards for the training images.

Assignments (2)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jun 6, 2025
From: DEEPMIND TECHNOLOGIES LIMITED
To: GDM HOLDING LLC
Reel/Frame 071498/0210 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Feb 20, 2024
From: SINGH, AADITYA K.; DING, FENGNING; HILL, FELIX GEORGE; LAMPINEN, ANDREW KYLE
To: DEEPMIND TECHNOLOGIES LIMITED
Reel/Frame 066496/0557 →
Continuity (2)
Provisional Application 63343960 · May 19, 2022
Related Publication 20230401835A1 · Dec 14, 2023
References Cited (60)
US 10460033B2 · Cohen · 2019 [cited by examiner]
US 10467274B1 · Ren · 2019 [cited by examiner]
US 10496884B1 · Nguyen · 2019 [cited by examiner]
US 12167100B2 · Lin · 2024 [cited by examiner]
US 12210825B2 · Cho · 2025 [cited by examiner]
US 20180268297A1 · Okazaki · 2018 [cited by examiner]
US 20210326660A1 · Krishnan · 2021 [cited by examiner]
US 20230019211A1 · Wang · 2023 [cited by examiner]
US 20230401877A1 · Srinivasa · 2023 [cited by examiner]
EP 3889836A1 · 2021 [cited by examiner]
WO WO2020190112A1 · 2020 [cited by examiner]
Yun et al, “Gated Object-Attribute Matching Network for Detailed Image Caption” (pp. 1-11). (Year: 2020). [cited by examiner]
He et al., “A Modularized Architecture of Multi-Branch Convolutional Neural Network for Image Captioning” (pp. 1-15) (Year: 2019). [cited by examiner]
Agarwal et al., “Evaluating CLIP: towards characterization of broader capabilities and downstream implications,” CoRR, Aug. 5, 2021, arxiv.org/abs/2108.02818, 5 pages. [cited by applicant]
Alayrac et al., “Flamingo: a visual language model for few-shot learning,” CoRR, Apr. 29, 2022, arxiv.org/abs/2204.14198, 54 pages. [cited by applicant]
Baumli et al., “Relative variational intrinsic control,” CoRR, Dec. 14, 2020, arxiv.org/abs/2012.07827, 9 pages. [cited by applicant]
Brock et al., “High-performance large scale image recognition without normalization,” CoRR, Feb. 11, 2021, arxiv.org/abs/2102.06171, 22 pages. [cited by applicant]
Brown et al., “Language models are few-shot learners,” CoRR, Jul. 22, 2020, 75 pages. [cited by applicant]
Burns et al., “Women also Snowboard: Overcoming Bias in Captioning Models,” CoRR, Mar. 26, 2018, arXiv:1803.09797, 22 pages. [cited by applicant]
Clark et al., “Referring as a collaborative process,” Cognition, 22, 1986, 387(1):1-39. [cited by applicant]
Cogswell et al., “Dialog without dialog data: Learning visual dialog agents from VQA data,” CoRR, Jul. 24, 2020, Jul. 24, 2020, 19 pages. [cited by applicant]
Elhagry et al., “A thorough review on recent deep learning methodologies for image captioning,” CoRR, Jul. 28, 2021, arxiv.org/abs/2107.13114, 6 pages. [cited by applicant]
Github.com [online], “Haiku: SONNET for JAX,” Sep. 2023, retrieved on Nov. 22, 2023, retrieved from URL<https://github.com/google/jax/>, 11 pages. [cited by applicant]
Github.com [online], “JAX: Autograd and XLA,” Nov. 2018, retrieved on Nov. 22, 2023, retrieved from URL<https://github.com/google/jax/>, 11 pages. [cited by applicant]
Hawkins et al., “From partners to populations: A hierarchical bayesian account of coordination and convention,” Psychological Review, 2022, 130(4):977-1016. [cited by applicant]
Hawkins et al., “The emergence of social norms and conventions,” Trends in cognitive sciences, Feb. 2019, 23(2):158-169. [cited by applicant]
Hoffmann et al., “Training compute-optimal large language models,” CoRR, Mar. 29, 2022, arxiv.org/abs/2203.15556, 36 pages. [cited by applicant]
Holtzman et al., “The curious case of neural text degeneration,” CoRR, Apr. 22, 2019, arxiv.org/abs/1904.09751, 16 pages. [cited by applicant]
Hu et al., “Lora: Low-rank adaptation of large language models,” CoRR, Jun. 17, 2021, arxiv.org/abs/2106.09685, 26 pages. [cited by applicant]
Jaderberg et al., Human-level performance in 3D multiplayer games with population-based reinforcement learning, Science, May 31, 2019, 364(6443):859-865. [cited by applicant]
Jaegle et al., “Perceiver IO: A general architecture for structured inputs & outputs,” CoRR, Jul. 30, 2021, arxiv.org/abs/2107.14795, 29 pages. [cited by applicant]
Jia et al., “Scaling up visual and vision-language representation learning with noisy text supervision,” CoRR, Feb. 11, 2021, arxiv.org/abs/2102.05918, 14 pages. [cited by applicant]
Keys, “Cubic convolution interpolation for digital image processing,” IEEE Transactions on Acoustics, Speech, and Signal Processing, Dec. 1981, 29(6):1153-1160. [cited by applicant]
Kingma et al., “Adam: A method for stochastic optimization,” CoRR, Dec. 22, 2014, arxiv.org/abs/1412.6980, 15 pages. [cited by applicant]
Kudo et al., “Sentencepiece: A simple and language independent subword tokenizer and detokenizer for neural text processing,” CoRR, Aug. 19, 2018, arxiv.org/abs/1808.06226, 6 pages. [cited by applicant]
Kunda et al., “Creative captioning: An AI grand challenge based on the dixit board game,” CoRR, Sep. 30, 2020, arxiv.org/abs/2010.00048, 9 pages. [cited by applicant]
Lazaridou et al., “Multi-agent communication meets natural language: Synergies between functional and structural language learning,” CoRR, May 14, 2020, arxiv.org/abs/2005.07064, 11 pages. [cited by applicant]
Lee et al., “Countering language drift via visual grounding,” CoRR, Sep. 10, 2019, arXiv:1909.04499, 11 pages. [cited by applicant]
Li et al., “Prefix-Tuning: Optimizing continuous prompts for generation,” CoRR, Jan. 1, 2021, arxiv.org/abs/2101.00190, 15 pages. [cited by applicant]
Lin et al., “Microsoft COCO: common objects in context,” CoRR, May 1, 2014, arxiv.org/abs/1405.0312, 15 pages. [cited by applicant]
Meade et al., “An empirical survey of the effectiveness of debiasing techniques for pre-trained language models,” CoRR, Oct. 16, 2021, arxiv.org/abs/2110.08527, 21 pages. [cited by applicant]
Mokady et al., “Clipcap: CLIP prefix for image captioning,” CoRR, Nov. 18, 2021, arxiv.org/abs/2111.09734, 10 pages. [cited by applicant]
Ouyang et al., “Training language models to follow instructions with human feedback,” CoRR, Mar. 4, 2022, arXiv:2203.02155, 68 pages. [cited by applicant]
Perez et al., “Red teaming language models with language models,” CoRR, Feb. 7, 2022, arxiv.org/abs/2202.03286, 31 pages. [cited by applicant]
Purwins et al., “Deep learning for audio signal processing,” CoRR, Apr. 30, 2019, arxiv.org/abs/1905.00078, 15 pages. [cited by applicant]
Radford et al., “Learning transferable visual models from natural language supervision,” CoRR, 2021, arxiv.org/abs/2103.00020, 48 pages. [cited by applicant]
Rajbhandari et al., “ZeRO; Memory optimizations toward training trillion parameter models,” CoRR, Oct. 4, 2019, arxiv.org/abs/1910.02054, 24 pages. [cited by applicant]
Schick et al., “Self-Diagnosis and Self-Debiasing: A Proposal for Reducing Corpus-Based Bias in NLP,” Transactions of the Association for Computational Linguistics, Dec. 17, 2021, 9:1408-1424. [cited by applicant]
Sharma et al., “Conceptual captions: A cleaned, hypernymed, image alt-text dataset for automatic image captioning,” Proceedings of the 56th Annual Meeting of the Association for Computational Linguistics (vol. 1: Long P… [cited by applicant]
Stiennon et al., “Learning to summarize with human feedback,” Advances in Neural Information Processing Systems, 2020, 33:3008-3021. [cited by applicant]
Tsimpoukelli et al., “Multimodal few-shot learning with frozen language models,” CoRR, Jun. 25, 2021, arxiv.org/abs/2106.13884, 19 pages. [cited by applicant]
Vaswani et al., “Attention is all you need,” CoRR, Dec. 6, 2017, arxiv.org/abs/1706.03762, 15 pages. [cited by applicant]
Vélez et al., “The rare preference effect: Statistical information influences social affiliation judgments,” Cognition, Nov. 2019, 192:103994. [cited by applicant]
Wang et al., “Learning to represent student knowledge on programming exercises using deep learning,” International Educational Data Mining Society, Jun. 25-28, 2017, pp. 324-329. [cited by applicant]
Weidinger et al., “Ethical and social risks of harm from language models,” CoRR, Dec. 8, 2021, arxiv.org/abs/2112.04359, 64 pages. [cited by applicant]
Williams et al., “A learning algorithm for continually running fully recurrent neural networks,” Neural Computation, Jun. 1, 1989, 1(2):270-280. [cited by applicant]
Williams et al., “Simple statistical gradient-following algorithms for connectionist reinforcement learning,” Mach. Learn., May 1992, 8(3-4):229-256. [cited by applicant]
Xiao et al., “Fashion-MNIST: a novel image dataset for benchmarking machine learning algorithms,” CoRR, Aug. 25, 2017, arxiv.org/abs/1708.07747, 6 pages. [cited by applicant]
Yoon et al., “Talking with tact: Polite language as a balance between kindness and informativity,” Proceedings of the 38th Annual Conference of the Cognitive Science Society, 2016, pp. 2771-2776. [cited by applicant]
Ziegler et al., “Fine-tuning language models from human preferences,” CoRR, Sep. 18, 2019, arxiv.org/abs/1909.08593, 26 pages. [cited by applicant]