IP Library Granted Patent US 12,242,947
Granted Patent B2
US 12,242,947 · App. 16/759,561 · Granted Mar 4, 2025

Machine learning systems with memory based parameter adaptation for learning fast and slower

Inventors: Pablo Sprechmann (London, GB); Siddhant Jayakumar (London, GB); Jack William Rae (London, GB); Alexander Pritzel (London, GB); Adrià Puigdomènech Badia (London, GB); Oriol Vinyals (London, GB); Razvan Pascanu (London, GB); Charles Blundell (London, GB)
Assignee: DeepMind Technologies Limited
G06N3/045G06N3/02G06N3/04G06N3/044G06N3/084
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,242,947
App. No.
16/759,561
Granted
Mar 4, 2025
Kind
B2
Abstract

There is described herein a computer-implemented method of processing an input data item. The method comprises processing the input data item using a parametric model to generate output data, wherein the parametric model comprises a first sub-model and a second sub-model. The processing comprises processing, by the first sub-model, the input data to generate a query data item, retrieving, from a memory storing data point-value pairs, at least one data point-value pair based upon the query data item and modifying weights of the second sub-model based upon the retrieved at least one data point-value pair. The output data is then generated based upon the modified second sub-model.

Claims (38)

1. A computer-implemented method of processing an input data item, comprising:

processing the input data item using a parametric model to generate output data, wherein the parametric model comprises a first sub-model and a second sub-model, the processing comprising:

processing, by the first sub-model, the input data to generate a query data item;

retrieving, from an external memory storing data point-value pairs, at least one data point-value pair based upon the query data item;

modifying weights of the second sub-model using the at least one data-point value pair that was retrieved based upon the query data item without modifying weights of the first sub-model; and

processing the query data item using the second sub-model in accordance with the modified weights to generate the output data.

2. The method as claimed in claim 1 wherein the parametric model comprises a neural network, and wherein the first and second sub-models comprise first and second sub-networks of the neural network.

3. The method according to claim 2 , wherein modifying the weights of the second sub-network using the at least one data-point value pair that was retrieved based upon the query data item comprises generating a plurality of weights for the second sub-network.

4. The method according to claim 3 , wherein the plurality of weights are generated based upon a relationship between a data point and a value of the data point-value pairs.

5. The method according to claim 2 , wherein modifying the weights of the second sub-network using the at least one data-point value pair that was retrieved based upon the query data item comprises minimizing a loss function, wherein the loss function is based upon a relationship between the at least one data point-value pair and the second sub-network.

6. The method according to claim 5 , wherein the loss function is further based upon the query data item.

7. The method according to claim 6 , wherein the loss function is weighted based upon a relationship between the query data item and the data-point value pair.

8. The method according to claim 1 , wherein the query data item comprises a hidden state of the first sub-model.

9. The method according to claim 1 , wherein the external memory comprises an episodic memory.

10. The method according to claim 2 , wherein the first sub-network comprises a first plurality of first neural network layers and the second sub-network comprises one or more second neural network layers, wherein the number of first neural network layers is greater than the number of second neural network layers.

11. The method according to claim 2 , wherein the data point-value pair comprises:

a data point of the data point-value pair associated with a hidden state of the first sub-network; and

a value of the data point-value pair associated with an output of the second sub-network for the data point-value pair.

12. The method according to claim 1 , further comprising resetting the second sub-model after the output data is generated using the second sub-model in accordance with the modified weights.

13. The method according to claim 1 , wherein the input data item is processed as part of a reinforcement learning system.

14. The method of claim 1 , wherein the input data item is a data item associated with data of a category selected from the group consisting of: image data, video data, motion data, speech data, audio data, an electronic document, data representing a state of an environment and data representing an action.

15. A system comprising one or more computers and one or more storage devices storing instructions that when executed by the one or more computers cause the one or more computers to perform operations for processing an input data item, the operations comprising:

processing the input data item using a parametric model to generate output data, wherein the parametric model comprises a first sub-model and a second sub-model, the processing comprising:

processing, by the first sub-model, the input data to generate a query data item;

retrieving, from an external memory storing data point-value pairs, at least one data point-value pair based upon the query data item;

modifying weights of the second sub-model using the at least one data-point value pair that was retrieved based upon the query data item without modifying weights of the first sub-model; and

processing the query data item using the second sub-model in accordance with the modified weights to generate the output data.

16. One or more non-transitory computer storage media storing instructions that when executed by one or more computers cause the one or more computers to perform operations for processing an input data item, the operations comprising:

processing the input data item using a parametric model to generate output data, wherein the parametric model comprises a first sub-model and a second sub-model, the processing comprising:

processing, by the first sub-model, the input data to generate a query data item;

retrieving, from an external memory storing data point-value pairs, at least one data point-value pair based upon the query data item;

modifying weights of the second sub-model using the at least one data-point value pair that was retrieved based upon the query data item without modifying weights of the first sub-model; and

processing the query data item using the second sub-model in accordance with the modified weights to generate the output data.

17. The system according to claim 15 wherein the parametric model comprises a neural network, and wherein the first and second sub-models comprise first and second sub-networks of the neural network.

18. The system according to claim 17 , wherein modifying the weights of the second sub-network using the at least one data-point value pair that was retrieved based upon the query data item comprises generating a plurality of weights for the second sub-network.

19. The system according to claim 18 , wherein the plurality of weights are generated based upon a relationship between a data point and a value of the data point-value pairs.

20. The system according to claim 17 , wherein modifying the second sub-network using the at least one data-point value pair that was retrieved based upon the query data item comprises minimizing a loss function, wherein the loss function is based upon a relationship between the at least one data point-value pair and the second sub-network.

21. The method according to claim 1 , further comprising adding the output data to the external memory.

Assignments (2)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jun 6, 2025
From: DEEPMIND TECHNOLOGIES LIMITED
To: GDM HOLDING LLC
Reel/Frame 071498/0210 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jun 9, 2020
From: SPRECHMANN, PABLO; JAYAKUMAR, SIDDHANT; RAE, JACK WILLIAM; PRITZEL, ALEXANDER; BADIA, ADRIÀ PUIGDOMÈNECH; VINYALS, ORIOL; PASCANU, RAZVAN; BLUNDELL, CHARLES
To: DEEPMIND TECHNOLOGIES LIMITED
Reel/Frame 052884/0180 →
Continuity (2)
Provisional Application 62578319 · Oct 27, 2017
Related Publication 20200285940A1 · Sep 10, 2020
References Cited (59)
US 20030033347A1 · Bolle · 2003 [cited by examiner]
US 20140358546A1 · Fernandez · 2014 [cited by examiner]
US 20170024645A1 · Socher · 2017 [cited by examiner]
US 20170228637A1 · Santoro · 2017 [cited by examiner]
Graves et al., “Hybrid Computing Using a Neural Network with Dynamic External Memory” Oct. 12, 2016, doi: 10.1038/nature20101, pp. 471-476. (Year: 2016). [cited by examiner]
Gu et al., “Search Engine Guided Non-Parametric Neural Machine Translation” May 20, 2017, arXiv:1705.07267v1. (Year: 2017). [cited by examiner]
Kumar et al., “Ask Me Anything: Dynamic memory Network for Natural Language Processing” Mar. 5, 2016, arXiv: 1506.07285v5. (Year: 2016). [cited by examiner]
Turchi et al., “Continuous Learning from Humans Post Edits for Neural Machine Translation” Jun. 1, 2017, doi: 10.1515/pralin-2017-0023, (Year: 2017). [cited by examiner]
Aharoni et al., “Gradual Learning of Deep Recurrent Neural Networks,” CoRR, Aug. 2017, https://arxiv.org/abs/1708.08863v1, 7 pages. [cited by applicant]
Anselmi et al., “Unsupervised learning of invariant representations with low sample complexity: the magic of sensory cortex or a new framework for machine learning?” CoRR, Mar. 2014, https://arxiv.org/abs/1311.4158v5, 2… [cited by applicant]
Ba et al., “Using Fast Weights to Attend to the Recent Past,” Advances in Neural Information Processing Systems 29 (NIPS 2016), 2016, 9 pages. [cited by applicant]
Bahdanau et al., “Neural Machine Translation by Jointly Learning to Align and Translate,” CoRR, Sep. 2014, ahttps://arxiv.org/abs/1409.0473v1, 15 pages. [cited by applicant]
Blundell et al., “Model-Free Episodic Control,” CoRR, Jun. 2016, https://arxiv.org/abs/1606.04460, Jun. 2016, 12 pages. [cited by applicant]
Farajian et al., “Multi-Domain Neural Machine Translation through Unsupervised Adaptation,” Proceedings of the Conference on Machine Translation (WMT), Sep. 2017, 127-137. [cited by applicant]
Finn et al., “Model-Agnostic Meta-Learning for Fast Adaptation of Deep Networks,” CoRR, Mar. 2017, https://arxiv.org/abs/1703.03400v3, 13 pages. [cited by applicant]
Fortunato et al., “Bayesian Recurrent Neural Networks,” CoRR, Apr. 2017, https://arxiv.org/abs/1704.02798v2, 11 pages. [cited by applicant]
French, “Catastrophic forgetting in connectionist networks,” Trends in Cognitive Sciences, Apr. 1999, 3(4):128-135. [cited by applicant]
Furlanello et al., “Active Long Term Memory Networks,” CoRR, Jun. 2016, https://arxiv.org/abs/1606.02355, 10 pages. [cited by applicant]
Goodfellow et al., “Maxout Networks,” CoRR, Feb. 2013, https://arxiv.org/abs/1302.4389v4, 9 pages. [cited by applicant]
Grave et al., “Improving Neural Language Models with a Continuous Cache,” CoRR, Dec. 2016, https://arxiv.org/abs/1612.04426, 9 pages. [cited by applicant]
Gu et al., “Search Engine Guided Non-Parametric Neural Machine Translation,” CoRR, May 2017, https://arxiv.org/abs/1705.07267v1, 11 pages. [cited by applicant]
He et al., “Deep Residual Learning for Image Recognition,” 2016 IEEE Conference on Computer Vision and Pattern Recognition (CVPR), Jun. 2016, pp. 770-778. [cited by applicant]
Hinton et al., “Distilling the Knowledge in a Neural Network,” CoRR, Mar. 2015, https://arxiv.org/abs/1503.02531, 9 pages. [cited by applicant]
Hochreiter et al., “Long Short-Term Memory,” Neural Computation, Nov. 1997, 9(8): 1735-1780. [cited by applicant]
Kaiser et al., “Learning to Remember Rare Events,” CoRR, Mar. 2017, https://arxiv.org/abs/1703.03129, 10 pages. [cited by applicant]
Kingma et al., “Adam: A Method for Stochastic Optimization,” CoRR, Dec. 2014, https://arxiv.org/abs/1412.6980v1, 9 pages. [cited by applicant]
Kirkpatrick et al., “Overcoming catastrophic forgetting in neural networks,” Proceedings of the National Academy of Sciences, Mar. 2017, 114(13):3521-3526. [cited by applicant]
Krause et al., “Dynamic Evaluation of Neural Sequence Models,” CoRR, Sep. 2017, https://arxiv.org/abs/1709.07432v2, 10 pages. [cited by applicant]
Krizhevsky et al., “ImageNet Classification with Deep Convolutional Neural Networks,” Advances in Neural Information Processing Systems 25 (NIPS 2012), 2012, 9 pages. [cited by applicant]
Kumaran et al., “What Learning Systems do Intelligent Agents Need? Complementary Learning Systems Theory Updated,” Trends in Cognitive Sciences, Jul. 2016, 20(7):512-534. [cited by applicant]
LeCun et al., “Gradient-Based Learning Applied to Document Recognition,” Proceedings of the IEEE, Nov. 1998, 86(11):2278-2324. [cited by applicant]
Leibo et al., “Approximate Hubel-Wiesel Modules and the Data Structures of Neural Computation,” CoRR, Dec. 2015, arxiv.org/abs/1512.08457, 13 pages. [cited by applicant]
Li et al., “Learning Without Forgetting,” IEEE Transactions on Pattern Analysis and Machine Intelligence, Nov. 2017, pp. 2935-2947. [cited by applicant]
Li et al., “One Sentence One Model for Neural Machine Translation,” CoRR, Sep. 2016, arxiv.org/abs/1609.06490, 7 pages. [cited by applicant]
Lopez-Paz et al., “Gradient Episodic Memory for Continual Learning,” 31st Conference on Neural Information Processing Systems, Jun. 2017, 10 pages. [cited by applicant]
Marcus et al., “Building a Large Annotated Corpus of English: The Penn Treebank,” Computational Linguistics, Oct. 1993, 19(2):313-330. [cited by applicant]
McClelland et al., “Why There Are Complementary Learning Systems in the Hippocampus and Neocortex: Insights From the Successes and Failures of Connectionist Models of Learning and Memory,” Psychological Review, 1995, 10… [cited by applicant]
McCloskey et al., “Catastrophic Interference in Connectionist Networks: The Sequential Learning Problem,” Psychology of Learning and Motivation, 1989, 24:109-165. [cited by applicant]
Melis et al., “On the State of the Art of Evaluation in Neural Language Models,” CoRR, Jul. 2017, arxiv.org/abs/1707.05589v1, 7 pages. [cited by applicant]
Merity et al., “Pointer Sentinel Mixture Models,” CoRR, Sep. 2016, arxiv.org/abs/1609.07843, 13 pages. [cited by applicant]
Merity et al., “Regularizing and Optimizing LSTM Language Models,” CoRR, Aug. 2017, arxiv.org/abs/1708.02182, 10 pages. [cited by applicant]
Metz et al., “Unrolled Generative Adversarial Networks,” CoRR, Nov. 2016, arxiv.org/abs/1611.02163v1, Nov. 2016, 19 pages. [cited by applicant]
Mnih et al., “Human-level control through deep reinforcement learning,” Nature, Feb. 2015, 518(7540):529-533. [cited by applicant]
Munkhdalai et al., “Meta Networks,” CoRR, Mar. 2017, arxiv.org/abs/1703.00837v2, 11 pages. [cited by applicant]
PCT International Preliminary Report on Patentability in International Appln. No. PCT/EP2018/079559, mailed May 7, 2020, 10 pages. [cited by applicant]
PCT International Search Report and Written Opinion in International Appln. No. PCT/EP2018/079559, mailed Feb. 4, 2019, 16 pages. [cited by applicant]
Pritzel et al., “Neural Episodic Control,” CoRR, Mar. 2017, arxiv.org/abs/1703.01988, 12 pages. [cited by applicant]
Ravi et al., “Optimization as a Model for Few-Shot Learning,” ICLR, Nov. 2016, 11 pages. [cited by applicant]
Russakovsky et al., “ImageNet Large Scale Visual Recognition Challenge,” International Journal of Computer Vision, Apr. 2015, 115(3):211-252. [cited by applicant]
Santoro et al., “One-shot Learning with Memory-Augmented Neural Networks,” CoRR, May 2016, arxiv.org/abs/1605.06065, 13 pages. [cited by applicant]
Silver et al., “Mastering the game of Go without human knowledge,” Nature, Oct. 2017, 550(7676):354-359. [cited by applicant]
Snell et al., “Prototypical Networks for Few-shot Learning,” 31st Conference on Neural Information Processing Systems, Jun. 2017, 11 pages. [cited by applicant]
Sprechmann et al., “Memory-based Parameter Adaptation,” CoRR, Feb. 2018, arxiv.org/abs/1802.10542, 16 pages. [cited by applicant]
Turchi et al., “Continuous Learning from Human Post-Edits for Neural Machine Translation,” The Prague Bulletin of Mathematical Linguistics, Jun. 2017, 108(1):233-244. [cited by applicant]
Van den Oord et al., “WaveNet: A Generative Model for Raw Audio,” CoRR, Sep. 2016, arxiv.org/abs/1609.03499v2, 15 pages. [cited by applicant]
Vinyals et al., “Matching Networks for One Shot Learning,” Advances in Neural Information Processing Systems 29, 2016, 9 pages. [cited by applicant]
Weston et al., “Memory Networks,” CoRR, Oct. 2014, arxiv.org/abs/1410.3916v1, 15 pages. [cited by applicant]
Wu et al., “Google's Neural Machine Translation System: Bridging the Gap between Human and Machine Translation,” CoRR, Sep. 2016, arxiv.org/abs/1609.08144v2, 23 pages. [cited by applicant]
Zhang et al., “Character-level Convolutional Networks for Text Classification,” Advances in Neural Information Processing Systems 28, 2015, 9 pages. [cited by applicant]