IP Library › Granted Patent US 12,572,796
Granted Patent B2
US 12,572,796 · App. 17/138,890 · Granted Mar 10, 2026

Methods and systems for generating recommendations for counterfactual explanations of computer alerts that are automatically detected by a machine learning algorithm

Inventors: Brian Barr (McLean, VA); Jason Wittenbach (McLean, VA)
Assignee: Capital One Services, LLC
G06N3/08G06N5/045H04L63/1416H04L63/1425
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,572,796
App. No.
17/138,890
Granted
Mar 10, 2026
Kind
B2
Abstract

Methods and systems are described herein for generating recommendations for counterfactual explanations to computer alerts that are automatically detected by a machine learning algorithm. The methods and systems use an artificial neural network architecture that trains a hybrid classifier and autoencoder. For example, one model (or artificial neural network), which is a classifier, is trained to make predictions. A second model (or artificial neural network), which is an autoencoder, is trained to reconstruct its inputs. As the second model is trained to reconstruct its inputs means, the second model is implicitly trained to determine what in-sample data looks like. By combining these networks and train them jointly, the system generates predictions (e.g., counterfactual explanations) that are in-sample.

Claims (53)

1 . A system for generating recommendations for counterfactual explanations to computer security alerts that are automatically detected by machine learning algorithms monitoring network activity, comprising:

cloud-based memory configured to store a trained artificial neural network comprising: (i) a trained classifier used to detect a class of a known alert status based on labeled inputted feature vectors from a training data set corresponding to a first class of the known alert status and a second class of the known alert status, (ii) a first trained autoencoder trained on the first class, (iii) and a second trained autoencoder trained on multiple classes of known alert statuses including the first class and the second class, and wherein the trained classifier is trained to classify an input into the first class and the second class, wherein the known alert status comprises a detected cyber incident;

cloud-based control circuitry configured to:

receive a first feature vector associated with an unknown alert status, wherein the first feature vector represents values corresponding to a plurality of computer states in a computer system, and wherein the first feature vector occupies a feature space having a first dimensionality;

input the first feature vector into the trained artificial neural network;

responsive to inputting the first feature vector into the trained classifier:

receive, from the trained classifier, a first prediction that the unknown alert status corresponds to the first class, and a latent representation of the first feature vector in a latent space, wherein the latent space has a smaller dimensionality than the feature space;

responsive to inputting the latent representation into the first trained autoencoder and the second trained autoencoder:

receive a first output from the first trained autoencoder comprising a first latent vector and a second output from the second trained autoencoder comprising a second latent vector;

apply, in the latent space, linear interpolation to the first latent vector of the first output and the second latent vector of the second output to generate a decoder input comprising an interpolated latent vector based on a weighted combination of the first latent vector and the second latent vector; and

apply, in the latent space, gradient descent to the interpolated latent vector to obtain an adjusted latent vector to correspond to the second class;

input the adjusted latent vector into a unified decoder trained on the multiple classes of the known alert statuses to map latent representations back to feature vectors in the feature space; and

receive an output from the unified decoder comprising a second feature vector representing one or more updates to the values corresponding to a smallest change to the first feature vector that causes the trained classifier to output a second prediction corresponding to the second class; and

cloud-based I/O circuitry configured to generate for display, on a user interface, a recommendation for a counterfactual explanation to the first prediction based on the second feature vector.

2 . A method, comprising:

receiving, using control circuitry, a first feature vector with an unknown alert status, wherein the first feature vector represents values corresponding to a plurality of computer states in a computer system;

inputting, using the control circuitry, the first feature vector into a trained artificial neural network comprising (i) a trained classifier used to detect a class of a known alert status, (ii) a first autoencoder trained on a first class of the known alert status, and (iii) a second autoencoder trained on multiple classes of known alert statuses, wherein the multiple classes of known alert statuses comprise the first class and at least a second class of the known alert status;

receiving, using the control circuitry, a first prediction from the trained classifier indicating that the unknown alert status corresponds to the first class and a latent representation of the first feature vector in a latent space having a smaller dimensionality than a feature space of the first feature vector;

receiving, using the control circuitry, a first output from the first autoencoder comprising a first latent vector and a second output from the second autoencoder comprising a second latent vector;

generating an interpolated latent vector based on linear interpolation being applied, in the latent space, to the first latent vector and the second latent vector;

generating an adjusted latent encoding corresponding to the second class based on gradient descent being applied to the interpolated latent vector;

inputting the adjusted latent encoding into a unified decoder trained on the multiple classes to obtain a second feature vector representing one or more updates to the values that causes the trained classifier to output a second prediction corresponding to the second class; and

generating, for display on a user interface, a recommendation for a counterfactual explanation to the first prediction based on the second prediction.

3 . The method of claim 2 , wherein the first feature vector has an isotropic gaussian distribution, and wherein the counterfactual explanation to the known alert status indicates a minimal change to the first feature vector that would cause the trained artificial neural network to change the first prediction.

4 . The method of claim 2 , wherein the multiple classes comprise multiple classes of the known alert statuses, generating the recommendation comprises:

receiving an output from the unified decoder, the output comprising the second prediction; and

determining the counterfactual explanation based on the output.

5 . The method of claim 4 , wherein the gradient descent being applied to the interpolated latent vector comprises:

applying, to the interpolated latent vector, a gradient descent loss function to determine a probability that the known alert status corresponds to the class of the known alert status as compared to the output from the unified decoder, wherein the recommendation is generated based on the probability.

6 . The method of claim 5 , further comprising:

determining a minimum at a decision boundary between the first class and the second class for the gradient descent loss function; and

determining a distance to the decision boundary.

7 . The method of claim 2 , wherein the known alert status comprises a detected fraudulent transaction, and wherein the values corresponding to the plurality of computer states in the computer system indicate a transaction history of a user.

8 . The method of claim 2 , wherein the known alert status comprises a detected cyber incident, and wherein the values corresponding to the plurality of computer states in the computer system indicate networking activity of a user.

9 . The method of claim 2 , wherein the known alert status comprises a refusal of a credit application, and wherein the values corresponding to the plurality of computer states in the computer system indicate credit history of a user.

10 . The method of claim 2 , wherein the known alert status comprises a detected identity theft, and wherein the values corresponding to the plurality of computer states in the computer system indicate a user transaction history.

11 . One or more non-transitory computer-readable media comprising instructions that, when executed by one or more processors, cause operations comprising:

receiving a first feature vector with an unknown alert status, wherein the first feature vector represents values corresponding to a plurality of computer states in a computer system;

inputting the first feature vector into a trained artificial neural network comprising (i) an agnostic classifier used to detect a class of a known alert status, (ii) a first autoencoder trained on a first class of the known alert status, and (ii) a second autoencoder trained on multiple classes of known alert statuses comprising the first class and at least a second class of the known alert status;

receiving (a) from the agnostic classifier, (a)(i) a first prediction from the agnostic classifier indicating that the unknown alert status corresponds to the first class and (a)(ii) a latent representation of the first feature vector in a latent space having a smaller dimensionality than a feature space of the first feature vector, (b) a first output from the first autoencoder comprising a first latent vector, and (c) second output from the second autoencoder comprising a second latent vector;

generating an interpolated latent vector based on linear interpolation being applied, in the latent space, to the first latent vector and the second latent vector;

inputting an adjusted latent vector corresponding to the second class, generated based on gradient descent being applied to the interpolated latent vector, into a unified decoder trained on the multiple classes to obtain a second feature vector representing one or more updates to the values that causes the agnostic classifier to output a second prediction corresponding to the second class; and

generating, for display on a user interface, a recommendation for a counterfactual explanation to the first prediction based on the second prediction.

12 . The one or more non-transitory computer-readable media of claim 11 , wherein the counterfactual explanation to the known alert status indicates a minimal change to the first feature vector that would cause the trained artificial neural network to change the first prediction.

13 . The one or more non-transitory computer-readable media of claim 11 , wherein the multiple classes comprise multiple classes of the known alert statuses, generating the recommendation comprises:

receiving an output from the unified decoder comprising the second prediction; and

determining the counterfactual explanation based on the output.

14 . The one or more non-transitory computer-readable media of claim 13 , wherein the gradient descent being applied to the interpolated latent vector comprises:

applying, to the interpolated latent vector a gradient descent loss function to determine a probability that the known alert status corresponds to the class of the known alert status as compared to the output from the unified decoder, and wherein the gradient descent loss function has a minimum at a decision boundary between the first class and the second class, wherein the recommendation is generated based on the probability.

15 . The one or more non-transitory computer-readable media of claim 11 , wherein the known alert status comprises a detected fraudulent transaction, and wherein the values corresponding to the plurality of computer states in the computer system indicate a transaction history of a user.

16 . The one or more non-transitory computer-readable media of claim 11 , wherein the known alert status comprises a detected cyber incident, and wherein the values corresponding to the plurality of computer states in the computer system indicate networking activity of a user.

17 . The one or more non-transitory computer-readable media of claim 11 , wherein the known alert status comprises a refusal of a credit application, and wherein the values corresponding to the plurality of computer states in the computer system indicate credit history of a user.

18 . The one or more non-transitory computer-readable media of claim 11 , wherein the known alert status comprises a detected identity theft, and wherein the values corresponding to the plurality of computer states in the computer system indicate a user transaction history.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jan 31, 2021
From: BARR, BRIAN; WITTENBACH, JASON
To: CAPITAL ONE SERVICES, LLC
Reel/Frame 055089/0591 →
Continuity (1)
Related Publication 20220207353A1 · Jun 30, 2022
References Cited (21)
US 10825028B1 · Kramme · 2020 [cited by examiner]
US 11403538B1 · Verma · 2022 [cited by examiner]
US 20170004397A1 · Yumer et al. · 2017 [cited by applicant]
US 20180176243A1 · Arnaldo et al. · 2018 [cited by applicant]
US 20180248893A1 · Israel et al. · 2018 [cited by applicant]
US 20190043126A1 · Billman · 2019 [cited by examiner]
US 20190068627A1 · Thampy · 2019 [cited by examiner]
US 20200151326A1 · Patrich et al. · 2020 [cited by applicant]
US 20200184316A1 · Kavukcuoglu et al. · 2020 [cited by applicant]
US 20200280573A1 · Johnson et al. · 2020 [cited by applicant]
US 20210390974A1 · Kawaguchi · 2021 [cited by examiner]
Seo, Eunbi, Hyun Min Song, and Huy Kang Kim. GIDS: GAN based intrusion detection system for in-vehicle network (Year: 2018). [cited by examiner]
Qian, Sheng, et al. “Improving representation learning in autoencoders via multidimensional interpolation and dual regularizations.” IJCAI (Year: 2019). [cited by examiner]
Hamon, Ronan, Henrik Junklewitz, and Ignacio Sanchez. Robustness and explainability of artificial intelligence (Year: 2020). [cited by examiner]
Wang, et al. “A parametric top-view representation of complex road scenes.” (Year: 2019). [cited by examiner]
Yang, et al., “Network intrusion detection based on supervised adversarial variational auto-encoder with regularization.”, (Year: 2020). [cited by examiner]
Seo, et al., “GIDS: GAN based intrusion detection system for in-vehicle network.” (Year: 2018). [cited by examiner]
Qian, et al. “Improving representation learning in autoencoders via multidimensional interpolation and dual regularizations.” (Year: 2019). [cited by examiner]
International Search Report and Written Opinion issued in corresponding International Application PCT/US2021/063089 on Apr. 15, 2022 (9 pages). [cited by applicant]
US Final Office Action on U.S. Appl. No. 17/138,886 Dated Sep. 18, 2024 (40 pages). [cited by applicant]
Yang, Yanqing, et al. “Network intrusion detection based on supervised adversarial variational auto-encoder with regularization.” (Year: 2020) (16 pages). [cited by applicant]