IP Library › Granted Patent US 12,223,697
Granted Patent B1
US 12,223,697 · App. 17/901,703 · Granted Feb 11, 2025

System and method for unsupervised concept extraction from reinforcement learning agents

Inventors: Praveen K. Pilly (West Hills, CA); Nicholas A. Ketz (Topanga, CA)
Assignee: HRL LABORATORIES, LLC
G06V10/776G06V10/761G06V10/762G06V10/774G06V20/70
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,223,697
App. No.
17/901,703
Granted
Feb 11, 2025
Kind
B1
Abstract

Described is a method for improved performance of agent-based machine learning. The method includes training a reinforcement learning (RL) agent on an image processing task. A dataset of states and corresponding actions is then extracted from the RL agent. A measure of attention is applied to an input space of the RL agent. During action selection by the RL agent, image patches of the input space are extracted based on the applied measure of attention. Portions of a set of inputs are clustered based on similarity to the image patches, generating a set of clusters having cluster centers. Non-semantic concept labels are provided as distances to the cluster centers for each state in the dataset.

Claims (47)

1. A system for performance of agent-based machine learning with unsupervised concept extraction from reinforcement learning agents, the system comprising:

one or more processors and a non-transitory computer-readable medium having executable instructions encoded thereon such that when executed, the one or more processors perform operations of:

training a reinforcement learning (RL) agent on an image processing task;

extracting a dataset of a plurality of states and corresponding actions from the RL agent, the plurality of states corresponding to a plurality of image frames;

perturbing each image frame by blurring each pixel in the image frame and focusing the blur around each pixel using a Gaussian filter centered on the pixel;

processing each image frame with the retrained RL agent to obtain an output;

comparing the output to the image frame prior to perturbation to determine a saliency score for each pixel in each image frame;

extracting, during action selection by the RL agent, a plurality of image patches of the input space based on the saliency score;

clustering, using an unsupervised clustering algorithm, portions of a first set of inputs based on similarity to the plurality of image patches, thereby generating a first set of clusters having cluster centers; and

providing non-semantic concept labels as distances to the cluster centers for each state in the dataset.

2. The system as set forth in claim 1 , wherein the one or more processors further perform an operation of using the non-semantic concept labels, retraining the RL agent on the image processing task.

3. The system as set forth in claim 1 , wherein the one or more processors further perform an operation of using the non-semantic concept labels, training another system that relies on the same input space.

4. The system as set forth in claim 1 , wherein the one or more processors perform operations of:

determining a saliency map for the image frame based on the saliency scores of the pixels in the image frame;

using the saliency map, extracting contiguous salient portions of the image frame; and

filtering the image frame using the extracted contiguous salient portions to highlight portions of the image frame for the RL agent to use for action selection.

5. A method for performance of agent-based machine learning with unsupervised concept extraction from reinforcement learning agents, the method comprising acts of:

training a reinforcement learning (RL) agent on an image processing task;

extracting a dataset of a plurality of states and corresponding actions from the RL agent, the plurality of states corresponding to a plurality of image frames;

perturbing each image frame by blurring each pixel in the image frame and focusing the blur around each pixel using a Gaussian filter centered on the pixel;

processing each image frame with the retrained RL agent to obtain an output;

comparing the output to the image frame prior to perturbation to determine a saliency score for each pixel in each image frame;

extracting, during action selection by the RL agent, a plurality of image patches of the input space based on the saliency score;

clustering, using an unsupervised clustering algorithm, portions of a first set of inputs based on similarity to the plurality of image patches, thereby generating a first set of clusters having cluster centers; and

providing non-semantic concept labels as distances to the cluster centers for each state in the dataset.

6. The method as set forth in claim 5 , further comprising an act of using the non-semantic concept labels, retraining the RL agent on the image processing.

7. The method as set forth in claim 5 , further comprising an act of using the non-semantic concept labels, training another system that relies on the same input space.

8. The method as set forth in claim 5 , further comprising acts of:

determining a saliency map for the image frame based on the saliency scores of the pixels in the image frame;

using the saliency map, extracting contiguous salient portions of the image frame; and

filtering the image frame using the extracted contiguous salient portions to highlight portions of the image frame for the RL agent to use for action selection.

9. A computer program product for performance of agent-based machine learning with unsupervised concept extraction from reinforcement learning agents, the computer program product comprising:

a non-transitory computer-readable medium having executable instructions encoded thereon, such that upon execution of the instructions by one or more processors, the one or more processors perform operations of:

training a reinforcement learning (RL) agent on an image processing task;

extracting a dataset of a plurality of states and corresponding actions from the RL agent, the plurality of states corresponding to a plurality of image frames;

perturbing each image frame by blurring each pixel in the image frame and focusing the blur around each pixel using a Gaussian filter centered on the pixel;

processing each image frame with the retrained RL agent to obtain an output;

comparing the output to the image frame prior to perturbation to determine a saliency score for each pixel in each image frame;

extracting, during action selection by the RL agent, a plurality of image patches of the input space based on the saliency score;

clustering, using an unsupervised clustering algorithm, portions of a first set of inputs based on similarity to the plurality of image patches, thereby generating a first set of clusters having cluster centers; and

providing non-semantic concept labels as distances to the cluster centers for each state in the dataset.

10. The computer program product as set forth in claim 9 , further comprising instructions for causing the one or more processors to perform an operation of using the non-semantic concept labels, retraining the RL agent on the image processing task.

11. The computer program product as set forth in claim 9 , further comprising instructions for causing the one or more processors to perform an operation of using the non-semantic concept labels, training another system that relies on the same input space.

12. The computer program product as set forth in claim 9 , further comprising instructions for causing the one or more processors to perform operations of:

determining a saliency map for the image frame based on the saliency scores of the pixels in the image frame;

using the saliency map, extracting contiguous salient portions of the image frame; and

filtering the image frame using the extracted contiguous salient portions to highlight portions of the image frame for the RL agent to use for action selection.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Sep 1, 2022
From: PILLY, PRAVEEN K.; KETZ, NICHOLAS A.
To: HRL LABORATORIES, LLC
Reel/Frame 060971/0172 →
Continuity (5)
Continuation In Part 17590726 · Feb 1, 2022
Continuation In Part 16900145 · Jun 12, 2020
Provisional Application 63290031 · Dec 15, 2021
Provisional Application 63146314 · Feb 5, 2021
Provisional Application 62906269 · Sep 26, 2019
References Cited (60)
US 6564198B1 · Narayan · 2003 [cited by examiner]
US 9785145B2 · Gordon et al. · 2017 [cited by applicant]
US 11436484B2 · Farabet · 2022 [cited by examiner]
US 20110081073A1 · Skipper · 2011 [cited by examiner]
US 20180118219A1 · Hiei et al. · 2018 [cited by applicant]
US 20200111011A1 · Viswanathan · 2020 [cited by examiner]
US 20200151562A1 · Pietquin · 2020 [cited by examiner]
US 20200293064A1 · Wu · 2020 [cited by examiner]
US 20210053570A1 · Akella · 2021 [cited by examiner]
US 20210061298A1 · Balachandran · 2021 [cited by examiner]
US 20210101619A1 · Weast · 2021 [cited by examiner]
US 20210271968A1 · Ganin · 2021 [cited by examiner]
US 20210284173A1 · Ohnishi · 2021 [cited by examiner]
CN 106740114A · 2017 [cited by examiner]
CN 110222406A · 2019 [cited by examiner]
DE 102014224765A1 · 2016 [cited by examiner]
EP 2256667B1 · 2012 [cited by examiner]
FR 2977056A1 · 2012 [cited by examiner]
JP 2015182526A · 2015 [cited by examiner]
JP 2018062321A · 2018 [cited by examiner]
JP 6520066B2 · 2019 [cited by examiner]
KR 20110119989A · 2011 [cited by examiner]
WO WO2016092796A1 · 2016 [cited by examiner]
WO WO2019002465A1 · 2019 [cited by applicant]
S. Greydanus, A. Koul, J. Dodge, and A. Fern, “Visualizing and Understanding Atari Agents,” ArXiv171100138 Cs, Sep. 2018, Accessed: Jun. 23, 2020. [Online]. Available: http://arxiv.org/abs/1711.00138, pp. 1-10. [cited by applicant]
S. Kolouri, C. E. Martin, and H. Hoffmann, “Explaining distributed neural activations via unsupervised learning,” in Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition Workshops, 2017, pp. 20-… [cited by applicant]
Office Action 1 for U.S. Appl. No. 16/900,145, Date mailed: Dec. 29, 2021. [cited by applicant]
Response to Office Action 1 for U.S. Appl. No. 16/900,145, Date mailed: Mar. 29, 2022. [cited by applicant]
Notice of Allowance for U.S. Appl. No. 16/900,145, Date mailed: Apr. 29, 2022. [cited by applicant]
Dutordoir, V., Salimbeni, H., Deisenroth, M., & Hensman, J. (2018). Gaussian Process Conditional Density Estimation. ArXiv:1810.12750, pp. 1-19. [cited by applicant]
Fawcett, Tom (2006). “An Introduction to ROC Analysis”. Pattern Recognition Letters. 27 (8): pp. 861-874. [cited by applicant]
Ketz, N., Kolouri, S., & Pilly, P. (2019). Using World Models for Pseudo-Rehearsal in Continual Learning. ArXiv:1903.02647, pp. 1-16. [cited by applicant]
Kansky K, et al., (2017). “Schema networks: Zero-Shot Transfer with a Generative Causal Model of Intuitive Physics.” In Proceedings of the 34th International Conference on Machine Learning. vol. 70: pp. 1809-1818. [cited by applicant]
Kolouri, Soheil, Charles E. Martin, and Heiko Hoffmann. (2017). “Explaining Distributed Neural Activations via Unsupervised Learning.” In CVPR Workshop on Explainable Computer Vision and Job Candidate Screening Competit… [cited by applicant]
Merrild, J., Rasmussen, M. A., & Risi, S. (2018). “HyperNTM: Evolving Scalable Neural Turing Machines through HyperNEAT.” Intemational Conference on the Applications of Evolutionary Computation, pp. 750-766. [cited by applicant]
Daftry, S., Zeng, S., Bagnell, J.A., and Hebert, M. (2016). “Introspective Perception: Learning to Predict Failures in Vision Systems.” In 2016 IEEE/RSJ Intemational Conference on Intelligent Robots and Systems (IROS), … [cited by applicant]
Mnih, V., Kavukcuoglu, K., Silver, D., Rusu, A.A., Veness, J., Bellemare, M.G., Graves, A., Riedmiller, M., Fidjeland, A.K., Ostrovski, G. and Petersen, S. (2015). “Human-Level Control Through Deep Reinforcement Leaming… [cited by applicant]
Miikkulainen, R., Liang, J., Meyerson, E., Rawal, A., Fink, D., Francon, O., and Hodjat, B. (2019). “Evolving Deep Neural Networks.” In Artificial Intelligence in the Age of Neural Networks and Brain Computing, pp. 293-… [cited by applicant]
Pilly, P.K., Howard, M.D., and Bhattacharyya, R. (2018). “Modeling Contextual Modulation of Memory Associations in the Hippocampus.” Frontiers in Human Neuroscience, 12, pp. 1-20. [cited by applicant]
Liou, Cheng-Yuan; Huang, Jau-Chi; Yang, Wen-Chie. (2008). “Modeling Word Perception Using the Elman Network”. Neurocomputing 71 (2008), pp. 3150-3157. [cited by applicant]
Notification of Transmittal, the International Search Report, and the Written Opinion of the International Searching Authority for PCT/US2020/037468; date of mailing Sep. 23, 2020. [cited by applicant]
Brett W Israelsen, et al: “a Dave . . . I can assure you . . . that ita s going to be all right . . . a a Definition, Case for, and Survey of Algorithmic Assurances in Human-Autonomy Trust Relationships,” ACM Computing … [cited by applicant]
Yunqi Zhao, et al: “Winning Isn't Everything: Enhancing Game Development with Intelligent Agents,” Arxiv .Org, Cornell University Library, 201 Olin Library Cornell University Ithaca, NY 14853, Mar. 25, 2019 (Mar. 25, 20… [cited by applicant]
Notification of the International Preliminary Report on Patentability Chapter I for PCT/US2020/037468; date of mailing Apr. 7, 2022. [cited by applicant]
The International Preliminary Report on Patentability Chapter I for PCT/US2020/037468; date of mailing Apr. 7, 2022. [cited by applicant]
Ketz, N., Kolouri, S., & Pilly, P. (2019). Using World Models for Pseudo-Rehearsal in Continual Learning. ArXiv:1903.02647 [Cs, Stat], pp. 1-16. [cited by applicant]
Kansky K, Silver T, Mély DA, Eldawy M, Lázaro-Gredilla M, Lou X, Dorfman N, Sidor S, Phoenix S, George D. (2017). Schema networks: Zero-shot transfer with a generative causal model of intuitive physics. In Proceedings o… [cited by applicant]
Kolouri, Soheil, Charles E. Martin, and Heiko Hoffmann (2017). “Explaining distributed neural activations via unsupervised learning.” In CVPR Workshop on Explainable Computer Vision and Job Candidate Screening Competiti… [cited by applicant]
Merrild, J., Rasmussen, M. A., & Risi, S. (2018). HyperNTM: Evolving scalable neural Turing machines through HyperNEAT. International Conference on the Applications of Evolutionary Computation, pp. 750-766. Springer. [cited by applicant]
Pilly, P. K., Howard, M. D., & Bhattacharyya, R. (2018). Modeling Contextual Modulation of Memory Associations in the Hippocampus. Frontiers in Human Neuroscience, Nov. 2018, vol. 12, Article 442, pp. 1-20. [cited by applicant]
Miikkulainen, R., Liang, J., Meyerson, E., Rawal, A., Fink, D., Francon, O., Raju, B., . . . & Hodjat, B. (2019) Evolving Deep Neural Networks, arXiv, 1703.00548, pp. 1-8. [cited by applicant]
Mnih, V., Badia, A.P., Mirza, M., Graves, A., Lillicrap, T., Harley, T., Silver, D. and Kavukcuoglu, K., 2016, June. Asynchronous methods for deep reinforcement leaming. In International Conference on Machine Leaming, p… [cited by applicant]
Robins, A., 1995. Catastrophic forgetting, rehearsal and pseudorehearsal. Connection Science, 7(2), pp. 123-146. [cited by applicant]
Kaiser, L., Babaeizadeh, M., Milos, P., Osinski, B., Campbell, R.H., Czechowski, K., Erhan, D., Finn, C., Kozakowski, P., Levine, S. and Mohiuddin, A., 2019. Model-based reinforcement leaming for Atari. arXiv preprint a… [cited by applicant]
Dutordoir, V., Salimbeni, H., Deisenroth, M., & Hensman, J. (2018). Gaussian Process Conditional Density Estimation. ArXiv:1810.12750 [Cs, Stat], pp. 1-19. [cited by applicant]
Fawcett, Tom (2006). “An Introduction to ROC Analysis”. Pattern Recognition Letters 27 (2006), pp. 861-874. doi:10.1016/j.patrec.2005.10.010. [cited by applicant]
Mnih, V, Kavukcuoglu, K, Silver, D, Rusu, A, Veness, J, Bellemare, M, Graves, A, Riedmiller, M, Fidjeland, A, Ostrovski, G, et al. (2015). Human-level control through deep reinforcement leaming. Nature, 518(7540): pp. 5… [cited by applicant]
Office Action 1 for U.S. Appl. No. 17/590,726, Date mailed: Jul. 12, 2023. [cited by applicant]
Response to Office Action 1 for U.S. Appl. No. 17/590,726, Date mailed: Oct. 11, 2023. [cited by applicant]
Notice of Allowance for U.S. Appl. No. 17/590,726, Date mailed: Oct. 23, 2023. [cited by applicant]