IP Library Granted Patent US 12,437,205
Granted Patent B2
US 12,437,205 · App. 18/299,942 · Granted Oct 7, 2025

Focused hyperparameter tuning using attribution

Inventors: Wanlin Xie (Shelton, CT); Mauro Joseph Sanchirico, III (Marlton, NJ)
Assignee: Lockheed Martin Corporation
G06N3/0985G06V10/776G06V10/82
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,437,205
App. No.
18/299,942
Granted
Oct 7, 2025
Kind
B2
Abstract

According to some embodiments, a method includes determining test predictions by performing an inference process using a checkpointed learning model and test vectors. The checkpointed learning model includes hyperparameters and weights. The method further includes determining an attribution map by performing one or more attribution processes using the test predictions and the test vectors. The method further includes determining a score for each particular hyperparameter by analyzing the attribution map using an association classifier. The method further includes determining, based on the analysis by the association classifier, whether each particular hyperparameter should be frozen or tuned again. The method further includes updating the hyperparameters and weights of the neural network when it is determined that at least one particular hyperparameter should be tuned again.

Claims (37)

1. A system for training a neural network, the system comprising:

one or more memory units; and

a processor communicatively coupled to the one or more memory units, the processor configured to:

determine a plurality of test predictions by performing an inference process using a checkpointed learning model of the neural network and a plurality of test vectors, wherein the checkpointed learning model comprises a plurality of hyperparameters and a plurality of weights;

determine an attribution map by performing one or more attribution processes using the plurality of test predictions and the plurality of test vectors;

determine a score for each particular hyperparameter of the plurality of hyperparameters by analyzing the attribution map using an association classifier;

determine, based on the analysis by the association classifier, whether each particular hyperparameter of the plurality of hyperparameters should be frozen or tuned again; and

when it is determined that at least one particular hyperparameter should be tuned again, update the plurality of hyperparameters and the plurality of weights of the neural network.

2. The system of claim 1 , wherein the attribution map comprises a plurality of learned features.

3. The system of claim 1 , wherein the plurality of test vectors comprises a plurality of images.

4. The system of claim 1 , wherein the one or more attribution processes comprises a gradient ascent process that comprises an integrated gradient algorithm.

5. The system of claim 1 , wherein determining whether each particular hyperparameter of the plurality of hyperparameters should be frozen or tuned again comprises comparing the score of each particular hyperparameter to a predetermined threshold.

6. The system of claim 1 , wherein the association classifier comprises a second neural network.

7. A method by a computing system, the method comprising:

determining a plurality of test predictions by performing an inference process using a checkpointed learning model of the neural network and a plurality of test vectors, wherein the checkpointed learning model comprises a plurality of hyperparameters and a plurality of weights;

determining an attribution map by performing one or more attribution processes using the plurality of test predictions and the plurality of test vectors;

determining a score for each particular hyperparameter of the plurality of hyperparameters by analyzing the attribution map using an association classifier;

determining, based on the analysis by the association classifier, whether each particular hyperparameter of the plurality of hyperparameters should be frozen or tuned again; and

when it is determined that at least one particular hyperparameter should be tuned again, updating the plurality of hyperparameters and the plurality of weights of the neural network.

8. The method of claim 7 , wherein the attribution map comprises a plurality of learned features.

9. The method of claim 7 , wherein the plurality of test vectors comprises a plurality of images.

10. The method of claim 7 , wherein determining the attribution map further comprises utilizing one or more gradient ascent processes.

11. The method of claim 10 , wherein the one or more gradient ascent processes comprises an integrated gradient algorithm.

12. The method of claim 7 , wherein determining whether each particular hyperparameter of the plurality of hyperparameters should be frozen or tuned again comprises comparing the score of each particular hyperparameter to a predetermined threshold.

13. The method of claim 7 , wherein the association classifier comprises a second neural network.

14. One or more computer-readable non-transitory storage media embodying instructions that, when executed by a processor, cause the processor to perform operations comprising:

determining a plurality of test predictions by performing an inference process using a checkpointed learning model of the neural network and a plurality of test vectors, wherein the checkpointed learning model comprises a plurality of hyperparameters and a plurality of weights;

determining an attribution map by performing one or more attribution processes using the plurality of test predictions and the plurality of test vectors;

determining a score for each particular hyperparameter of the plurality of hyperparameters by analyzing the attribution map using an association classifier;

determining, based on the analysis by the association classifier, whether each particular hyperparameter of the plurality of hyperparameters should be frozen or tuned again; and

when it is determined that at least one particular hyperparameter should be tuned again, updating the plurality of hyperparameters and the plurality of weights of the neural network.

15. The one or more computer-readable non-transitory storage media of claim 14 , wherein the attribution map comprises a plurality of learned features.

16. The one or more computer-readable non-transitory storage media of claim 14 , wherein the plurality of test vectors comprises a plurality of images.

17. The one or more computer-readable non-transitory storage media of claim 14 , wherein determining the attribution map further comprises utilizing one or more gradient ascent processes.

18. The one or more computer-readable non-transitory storage media of claim 17 , wherein the one or more gradient ascent processes comprises an integrated gradient algorithm.

19. The one or more computer-readable non-transitory storage media of claim 14 , wherein determining whether each particular hyperparameter of the plurality of hyperparameters should be frozen or tuned again comprises comparing the score of each particular hyperparameter to a predetermined threshold.

20. The one or more computer-readable non-transitory storage media of claim 14 , wherein the association classifier comprises a second neural network.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Apr 13, 2023
From: XIE, WANLIN; SANCHIRICO, MAURO JOSEPH, III
To: LOCKHEED MARTIN CORPORATION
Reel/Frame 063315/0430 →
Continuity (1)
Related Publication 20240346331A1 · Oct 17, 2024
References Cited (19)
US 10600005B2 · Gunes et al. · 2020 [cited by applicant]
US 10810512B1 · Wubbels et al. · 2020 [cited by applicant]
US 11003992B2 · Wesolowski et al. · 2021 [cited by applicant]
US 11157812B2 · McCourt et al. · 2021 [cited by applicant]
US 20190122141A1 · Zhen et al. · 2019 [cited by applicant]
US 20190244139A1 · Varadarajan et al. · 2019 [cited by applicant]
US 20190318248A1 · Moreira-Matias et al. · 2019 [cited by applicant]
US 20210264263A1 · Walters · 2021 [cited by examiner]
US 20220392637A1 · Kollada · 2022 [cited by examiner]
US 20240273400A1 · Boué · 2024 [cited by examiner]
CN 110889450A · 2020 [cited by applicant]
CN 111047016A · 2020 [cited by applicant]
CN 111553482A · 2020 [cited by applicant]
CN 113723615A · 2021 [cited by applicant]
IN 201821025560 · 2018 [cited by applicant]
IN 113673174A · 2021 [cited by applicant]
KR 102107378B · 2020 [cited by applicant]
TW I7332708B · 2021 [cited by applicant]
Goh, Gary SW, et al. “Understanding integrated gradients with smoothtaylor for deep neural network attribution.” 2020 25th International Conference on Pattern Recognition (ICPR). IEEE, 2021 (Year: 2021). [cited by examiner]