IP Library Granted Patent US 12,373,706
Granted Patent B2
US 12,373,706 · App. 17/066,457 · Granted Jul 29, 2025

Systems and methods for unsupervised continual learning

Inventors: Soheil Kolouri (Agoura Hills, CA); Mohammad Rostami (Los Angeles, CA); Praveen K. Pilly (West Hills, CA)
Assignee: HRL LABORATORIES, LLC
G06N5/022G06N20/00
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,373,706
App. No.
17/066,457
Granted
Jul 29, 2025
Kind
B2
Abstract

Described is a system for continual adaptation of a machine learning model implemented in an autonomous platform. The system adapts knowledge previously learned by the machine learning model for performance in a new domain. The system receives a consecutive sequence of new domains comprising new task data. The new task data and past learned tasks are forced to share a data distribution in an embedding space, resulting in a shared generative data distribution. The shared generative data distribution is used to generate a set of pseudo-data points for the past learned tasks. Each new domain is learned using both the set of pseudo-data points and the new task data. The machine learning model is updated using both the set of pseudo-data points and the new task data.

Claims (39)

1. A system for continual adaptation of a machine learning model implemented in an autonomous platform, the system comprising:

one or more processors and one or more associated memories, each associated memory being a non-transitory computer-readable medium having executable instructions encoded thereon such that when executed, the one or more processors perform an operation of:

adapting a set of knowledge previously learned by a machine learning model for performance in a new domain, wherein adapting the set of knowledge comprises:

receiving a consecutive sequence of new domains, where each new domain comprises new task data;

forcing, through dimensionality reduction, the new task data and a plurality of past learned tasks to share a same parametric data distribution in a task invariant embedding space as clusters of consolidated classes, resulting in a shared generative data distribution;

using the shared generative data distribution, generating a set of pseudo-data points for the past learned tasks;

learning each new domain using both the set of pseudo-data points and the new task data such that each new domain learned has an empirical data distribution in the task invariant embedding space that matches the same parametric data distribution;

updating the machine learning model using both the set of pseudo-data points and the new task data; and

causing one or more mechanical components of the autonomous platform to actuate and, in doing so, causing the autonomous platform to perform a physical driving operation based on the new task data.

2. The system as set forth in claim 1 , wherein the one or more processors further perform an operation of using knowledge of data distributions obtained from the plurality of past learned tasks to match a data distribution of new task data in the embedding space.

3. The system as set forth in claim 1 , wherein the embedding space is invariant with respect to any learned task, such that new task data does not interfere with remembering any past learned task.

4. The system as set forth in claim 1 , wherein the one or more processors further perform an operation of using a Sliced Wasserstein Distance metric to force the new task data and the plurality of past learned tasks to share the data distribution in the embedding space.

5. The system as set forth in claim 1 , wherein the shared generative data distribution is a multi-modal distribution modeled as a Gaussian mixture model.

6. A computer implemented method for continual adaptation of a machine learning model implemented in an autonomous platform, the method comprising an act of:

causing one or more processors to execute instructions encoded on one or more associated memories, each associated memory being a non-transitory computer-readable medium, such that upon execution, the one or more processors perform operations of:

adapting a set of knowledge previously learned by a machine learning model for performance in a new domain, wherein adapting the set of knowledge comprises:

receiving a consecutive sequence of new domains, where each new domain comprises new task data;

forcing, through dimensionality reduction, the new task data and a plurality of past learned tasks to share a same parametric data distribution in a task invariant embedding space as clusters of consolidated classes, resulting in a shared generative data distribution;

using the shared generative data distribution, generating a set of pseudo-data points for the past learned tasks;

learning each new domain using both the set of pseudo-data points and the new task data such that each new domain learned has an empirical data distribution in the task invariant embedding space that matches the same parametric data distribution;

updating the machine learning model using both the set of pseudo-data points and the new task data; and

causing one or more mechanical components of the autonomous platform to actuate and, in doing so, causing the autonomous platform to perform a physical driving operation based on the new task data.

7. The method as set forth in claim 6 , wherein the one or more processors further perform an operation of using knowledge of data distributions obtained from the plurality of past learned tasks to match a data distribution of new task data in the embedding space.

8. The method as set forth in claim 6 , wherein the embedding space is invariant with respect to any learned task, such that new task data does not interfere with remembering any past learned task.

9. The method as set forth in claim 6 , wherein the one or more processors further perform an operation of using a Sliced Wasserstein Distance metric to force the new task data and the plurality of past learned tasks to share the data distribution in the embedding space.

10. The method as set forth in claim 6 , wherein the shared generative data distribution is a multi-modal distribution modeled as a Gaussian mixture model.

11. A computer program product for continual adaptation of a machine learning model implemented in an autonomous platform, the computer program product comprising:

computer-readable instructions stored on a non-transitory computer-readable medium that are executable by a computer having one or more processors for causing the processor to perform operations of:

adapting a set of knowledge previously learned by a machine learning model for performance in a new domain, wherein adapting the set of knowledge comprises:

receiving a consecutive sequence of new domains, where each new domain comprises new task data;

forcing, through dimensionality reduction, the new task data and a plurality of past learned tasks to share a same parametric data distribution in a task invariant embedding space as clusters of consolidated classes, resulting in a shared generative data distribution;

using the shared generative data distribution, generating a set of pseudo-data points for the past learned tasks;

learning each new domain using both the set of pseudo-data points and the new task data such that each new domain learned has an empirical data distribution in the task invariant embedding space that matches the same parametric data distribution;

updating the machine learning model using both the set of pseudo-data points and the new task data; and

causing one or more mechanical components of the autonomous platform to actuate and, in doing so, causing the autonomous platform to perform a physical driving operation based on the new task data.

12. The computer program product as set forth in claim 11 , wherein the one or more processors further perform an operation of using knowledge of data distributions obtained from the plurality of past learned tasks to match a data distribution of new task data in the embedding space.

13. The computer program product as set forth in claim 11 , wherein the embedding space is invariant with respect to any learned task, such that new task data does not interfere with remembering any past learned task.

14. The computer program product as set forth in claim 11 , wherein the one or more processors further perform an operation of using a Sliced Wasserstein Distance metric to force the new task data and the plurality of past learned tasks to share the data distribution in the embedding space.

15. The computer program product as set forth in claim 11 , wherein the shared generative data distribution is a multi-modal distribution modeled as a Gaussian mixture model.

Assignments (2)
CONFIRMATORY LICENSE Recorded Jul 27, 2022
From: HRL LABORATORIES, LLC
To: GOVERNMENT OF THE UNITED STATES AS REPRESENTED BY THE SECRETARY OF THE AIR FORCE
Reel/Frame 060979/0589 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Oct 8, 2020
From: KOLOURI, SOHEIL; ROSTAMI, MOHAMMAD; PILLY, PRAVEEN K.
To: HRL LABORATORIES, LLC
Reel/Frame 054014/0037 →
Continuity (2)
Provisional Application 62953063 · Dec 23, 2019
Related Publication 20210192363A1 · Jun 24, 2021
References Cited (45)
US 9875440B1 · Commons · 2018 [cited by examiner]
US 20190130257A1 · Meyerson · 2019 [cited by examiner]
US 20200090045A1 · Baker · 2020 [cited by examiner]
US 20210110205A1 · Karimi · 2021 [cited by examiner]
Shin et al. “Continual Learning with Deep Generative Replay”, 2017 https://proceedings.neurips.cc/paper/2017/file/0efbe98067c6c73dba1250d2beaa81f9-Paper.pdf (Year: 2017). [cited by examiner]
Wu et al. “Sliced Wasserstein Generative Models”, 2018 https://arxiv.org/pdf/1706.02631v3.pdf (Year: 2018). [cited by examiner]
Kolouri et al. “Sliced Wasserstein Distance for Learning Gaussian Mixture Models”, 2018 https://ieeexplore.ieee.org/abstract/document/8578459 (Year: 2018). [cited by examiner]
Rao et al. “Continual Unsupervised Representation Learning”, 2019 https://proceedings.neurips.cc/paper/2019/file/861578d797aeb0634f77aff3f488cca2-Paper.pdf (Year: 2019). [cited by examiner]
“Continual Learning with Deep Generative Replay,” Shin et al(Year: 2017). [cited by examiner]
“Sliced Wasserstein Generative Models,” (Year: 2017). [cited by examiner]
“Iterative Machine Learning: A step towards Model Accuracy,” Packt, Bodarjee (Year: 2017). [cited by examiner]
Notification of the International Preliminary Report on Patentability Chapter I for PCT/US2020/054872; date of mailing Jul. 7, 2022. [cited by applicant]
The International Preliminary Report on Patentability Chapter I for PCT/US2020/054872; date of mailing Jul. 7, 2022. [cited by applicant]
The International Search Report of the International Searching Authority for PCT/US2020/054872; date of mailing Feb. 8, 2021. [cited by applicant]
The Written Opinion of the International Searching Authority for PCT/US2020/054872; date of mailing Feb. 8, 2021. [cited by applicant]
Rostami, M., et al., “Generative Continual Concept Learning,” arxiv.org, Cornell University Library, NY, 2019, pp. 1-8. [cited by applicant]
Rostami, M., et al., “Complementary Learning for Overcoming Catastrophic Forgetting Using Experience Replay,” Proceedings of the Twenty-Eight International Joint Conference on Artificial Intelligence, 2019, pp. 3339-334… [cited by applicant]
Lesort, T., et al., “Continual learning for robotics: Definition, framework, learning strategies, opportunities and challenges,” Information Fusion, Elsevier, US, vol. 58, 2019, pp. 52-68. [cited by applicant]
Yoo, J., et al., “Domain Adaptation Using Adversarial Learning for Autonomous Navigation,” arxiv.org, Cornell University Library, NY, 2017, pp. 1-20. [cited by applicant]
Nicolas Bonneel, Julien Rabin, Gabriel Peyr'e, and Hanspeter Pfister. Sliced and Radon Wasserstein Barycenters of Measures. Journal of Mathematical Imaging and Vision, 51(1): pp. 22-45, 2015. [cited by applicant]
Nicolas Bonnotte. Unidimensional and evolution methods for optimal transportation. PhD thesis, Chapter 5.1, pp. 119-125, Paris 11, 2013. [cited by applicant]
Nicolas Courty, Remi Flamary, Devis Tuia, and Alain Rakotomamonjy. Optimal Transport for Domain Adaptation. IEEE Transactions on Pattern Analysis and Machine Intelligence, 39(9): pp. 1853-1865, 2017. [cited by applicant]
Amir Globerson and Sam T Roweis. Metric Learning by Collapsing Classes. In Advances in Neural Information Processing Systems, pp. 451-458, 2006. [cited by applicant]
James Kirkpatrick, Razvan Pascanu, Neil Rabinowitz, Joel Veness, Guillaume Desjardins, Andrei A Rusu, Kieran Milan, John Quan, Tiago Ramalho, Agnieszka Grabska-Barwinska, et al. Overcoming Catastrophic Forgetting in Neu… [cited by applicant]
Alex Krizhevsky, Ilya Sutskever, and Geoffrey E Hinton. ImageNet Classification with Deep Convolutional Neural Networks. In Advances in Neural Information Processing Systems, pp. 1097-1105, 2012. [cited by applicant]
Brenden M Lake, Ruslan Salakhutdinov, and Joshua B Tenenbaum. Human-Level Concept Learning through Probabilistic Program Induction. Science, 350(6266): pp. 1332-1338, 2015. [cited by applicant]
Yann LeCun, Bernhard E Boser, John S Denker, Donnie Henderson, Richard E Howard, Wayne E Hubbard, and Lawrence D Jackel. Handwritten Digit Recognition with a Back-Propagation Network. In Advances in Neural Information P… [cited by applicant]
Marieke Longcamp, Marie-Therese Zerbato-Poudou, and Jean-Luc Velay. The Influence of Writing Practice on Letter Recognition in Preschool Children: A Comparison Between Handwriting and Typing. Acta Psychologica, 119(1): … [cited by applicant]
James L McClelland, Bruce L McNaughton, and Randall C O'Reilly. Why There are Complementary Learning Systems In the Hippocampus and Neocortex: Insights from the Successes and Failures of Connectionist Models of Learning… [cited by applicant]
James L McClelland and Timothy T Rogers. The Parallel Distributed Processing Approach to Semantic Cognition. Nature Reviews Neuroscience, 4(4): pp. 310-322, 2003. [cited by applicant]
James L McClelland, David E Rumelhart, PDP Research Group, et al. Parallel Distributed Processing. Explorations in the Microstructure of Cognition, 2: pp. 216-271, 1986. [cited by applicant]
Michael McCloskey and Neal J Cohen. Catastrophic Interference in Connectionist Networks: The Sequential Learning Problem. In Psychology of Learning and Motivation, vol. 24, pp. 109-165. Elsevier, 1989. [cited by applicant]
Leland McInnes, John Healy, and James Melville. UMap: Uniform Manifold Approximation and Projection for Dimension Reduction. arXiv preprint arXiv:1802.03426, pp. 1-63, 2018. [cited by applicant]
Saeid Motiian, Quinn Jones, Seyed Iranmanesh, and Gianfranco Doretto. Few-shot adversarial domain adaptation. In Advances in Neural Information Processing Systems, pp. 6670-6680, 2017. [cited by applicant]
German I Parisi, Ronald Kemker, Jose L Part, Christopher Kanan, and Stefan Wermter. Continual Lifelong Learning with Neural Networks: A Review. Neural Networks, pp. 54-71, 2019. [cited by applicant]
Julien Rabin and Gabriel Peyre. Wasserstein Regularization of Imaging Problem. In 2011 18th IEEE International Conference on Image Processing, pp. 1541-1544, IEEE, 2011. [cited by applicant]
Anthony Robins. Catastrophic Forgetting, Rehearsal and Pseudorehearsal. Connection Science, 7(2): pp. 123-146, 1995. [cited by applicant]
Mohammad Rostami, Soheil Kolouri, and Praveen Pilly. Complementary Learning for Overcoming Catastrophic Forgetting Using Experience Replay. In IJCAI, pp. 3339-3345, 2019. [cited by applicant]
Paul Ruvolo and Eric Eaton. Ella: An Efficient Lifelong Learning Algorithm. In International Conference on Machine Learning, pp. 507-515, 2013. [cited by applicant]
Andrew M Saxe, James L McClelland, and Surya Ganguli. A Mathematical Theory of Semantic Development in Deep Neural Networks. Proceedings of the National Academy of Sciences, pp. 11537-11546, 2019. [cited by applicant]
Shai Shalev-Shwartz and Shai Ben-David. Understanding Machine Learning: From Theory to Algorithms, Chapter 2.2, pp. 35-36, Cambridge University Press, 2014. [cited by applicant]
Hanul Shin, Jung Kwon Lee, Jaehong Kim, and Jiwon Kim. Continual Learning with Deep Generative Replay. In Advances in Neural Information Processing Systems, pp. 2990-2999, 2017. [cited by applicant]
Jake Snell, Kevin Swersky, and Richard Zemel. Prototypical Networks for Few-Shot Learning. In Advances in Neural Information Processing Systems, pp. 4077-4087, 2017. [cited by applicant]
Mark A. Kramer. Nonlinear Principal Component Analysis Using Autoassociative Neural Networks. AlChE Journal, 37(2), pp. 233-243. [cited by applicant]
Generative Continual Concept Learning. 33rd Conference on Neural Information Processing Systems (NeurIPS 2019), pp. 1-10, Vancouver, Canada. [cited by applicant]