IP Library › Granted Patent US 12,579,479
Granted Patent B2
US 12,579,479 · App. 17/937,356 · Granted Mar 17, 2026

Counterfactual samples for maintaining consistency between machine learning models

Inventors: Samuel Sharpe (Cambridge, MA); Christopher Bayan Bruss (Washington, DC); Brian Barr (Schenectady, NY)
Assignee: Capital One Services, LLC
G06N20/20G06N3/045G06N3/08G06N5/045G06Q30/0631G06Q40/03
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,579,479
App. No.
17/937,356
Granted
Mar 17, 2026
Kind
B2
Abstract

In some aspects, a computing system may aggregating multiple counterfactual samples so that machine learning explanations can be generated for sub-populations. In addition, methods and systems described herein use machine learning and counterfactual samples to determine text to use in an explanation for a model's prediction. A computing system may also train machine learning models to not only determine whether a request to perform an action should be accepted, but also to generate output that is consistent with output generated by previous machine learning models. Further, a computing system may generate counterfactual samples based on user preferences. A computing system may obtain preferences and then apply a penalty or adjustment parameter such that when a counterfactual sample is created, the computing system is forced to change one or more features indicated by the preferences to create the counterfactual sample.

Claims (38)

1 . A system for training a machine learning model to minimize generation of new output that conflicts with previous output generated by a previous machine learning model, the system comprising:

one or more processors and non-transitory media having instructions recorded thereon that, when executed by the one or more processors, cause operations comprising:

based on a first output, from a first machine learning model, indicating a communication corresponding to a first data sample should not be accepted, executing a counterfactual sample generation system to generate a first counterfactual sample comprising a modification for the first data sample, wherein the modification causes the first machine learning model to generate output indicating that the communication should be accepted;

configuring (i) a first training of a second machine learning model to use a first loss function that compares a first predicted probability of the second machine learning model with a first class label of first training data and (ii) a second training of the second machine learning model to use a second loss function that minimizes a difference between a second predicted probability of the second machine learning model and a threshold for accepting the communication;

after the first machine learning model outputs the first output indicating that the communication should not be accepted, and based on the configuration, executing the second machine learning model and alternating between the first training of the second machine learning model with the first loss function on the first training data and the second training of the second machine learning model with the second loss function on second training data comprising the first counterfactual sample; and

in connection with obtaining a second communication after completion of the first training and the second training, inputting data corresponding to the second communication to the second machine learning model to obtain a second output, from the second machine learning model, indicating that the second communication should be accepted.

2 . The system of claim 1 , wherein alternating between the first training and the second training comprises, after the first machine learning model outputs the first output indicating that the communication should not be accepted, alternating between the first training of the second machine learning model with the first loss function on the first training data comprising one or more other counterfactual samples and the second training of the second machine learning model with the second loss function on the second training data comprising the first counterfactual sample.

3 . The system of claim 1 , the operations further comprising:

in connection with the first training and the second training of the second machine learning model respectively with the first loss function and the second loss function, displaying, on a user interface, a first loss value corresponding to the first loss function, a second loss value corresponding to the second loss function, and a plurality of counterfactual samples.

4 . The system of claim 1 , wherein the first training data corresponds to a first set of users and the second training data corresponds to a second set of users that do not overlap with the first set of users.

5 . A method comprising:

performing, via one or more processors, operations comprising:

based on a first output, from a first machine learning model, indicating a communication corresponding to a first data sample should not be accepted, executing a counterfactual sample generation system to generate a first counterfactual sample comprising a modification for the first data sample, wherein the modification causes the first machine learning model to generate output indicating that the communication should be accepted;

after the first machine learning model outputs the first output indicating that the communication should not be accepted, executing a second machine learning model and alternating between a first training of the second machine learning model with a first loss function on first training data and a second training of the second machine learning model with a second loss function on second training data comprising the first counterfactual sample, wherein the first loss function compares a first predicted probability of the second machine learning model with a first class label of the first training data, and wherein the second loss function minimizes a difference between a second predicted probability of the second machine learning model and a threshold for accepting the communication; and

in connection with obtaining a second communication after completion of the first training and the second training, inputting data corresponding to the second communication to the second machine learning model to obtain a second output, from the second machine learning model, indicating that the second communication should be accepted.

6 . The method of claim 5 , wherein alternating between the first training and the second training comprises, after the first machine learning model outputs the first output indicating that the communication should not be accepted, alternating between the first training of the second machine learning model with the first loss function on the first training data comprising one or more other counterfactual samples and the second training of the second machine learning model with the second loss function on the second training data comprising the first counterfactual sample.

7 . The method of claim 5 , wherein alternating between the first training and the second training comprises, after the first machine learning model outputs the first output indicating that the communication should not be accepted, alternating between the first training of the second machine learning model with a binary cross-entropy function on the first training data and the second training of the second machine learning model with the second loss function on the second training data comprising the first counterfactual sample.

8 . The method of claim 5 , wherein alternating between the first training and the second training comprises, after the first machine learning model outputs the first output indicating that the communication should not be accepted, alternating between the first training of the second machine learning model with a cross-entropy loss function on the first training data and the second training of the second machine learning model with the second loss function on the second training data comprising the first counterfactual sample.

9 . The method of claim 5 , further comprising:

in connection with the first training and the second training of the second machine learning model respectively with the first loss function and the second loss function, displaying, on a user interface, a first loss value corresponding to the first loss function, a second loss value corresponding to the second loss function, and a plurality of counterfactual samples.

10 . The method of claim 5 , wherein:

executing the counterfactual sample generation system comprises executing the counterfactual sample generation system to generate a plurality of counterfactual samples for a plurality of data samples corresponding to communications that were not accepted; and

the second training data comprises the plurality of counterfactual samples.

11 . The method of claim 5 , wherein the first training data corresponds to a first set of users, and the second training data corresponds to a second set of users that do not overlap with the first set of users.

12 . One or more non-transitory computer-readable media comprising instructions that, when executed by one or more processors, causes operations comprising:

based on a first output, from a first machine learning model, indicating a communication corresponding to a first data sample should not be accepted, executing a counterfactual sample generation system to generate a first counterfactual sample comprising a modification for the first data sample, wherein the modification causes the first machine learning model to generate output indicating that the communication should be accepted;

executing a second machine learning model and alternating between a first training of the second machine learning model with a first loss function on first training data and a second training of the second machine learning model with a second loss function on second training data comprising the first counterfactual sample, wherein the first loss function compares a first predicted probability of the second machine learning model with a first class label of the first training data, and wherein the second loss function minimizes a difference between a second predicted probability of the second machine learning model and a threshold for accepting the communication; and

in connection with obtaining a second communication after completion of the first training and the second training, inputting data corresponding to the second communication to the second machine learning model to obtain a second output, from the second machine learning model, indicating that the second communication should be accepted.

13 . The one or more non-transitory computer-readable media of claim 12 , wherein alternating between the first training and the second training comprises, after the first machine learning model outputs the first output indicating that the communication should not be accepted, alternating between the first training of the second machine learning model with the first loss function on the first training data and the second training of the second machine learning model with the second loss function on the second training data.

14 . The one or more non-transitory computer-readable media of claim 12 , wherein alternating between the first training and the second training comprises, after the first machine learning model outputs the first output indicating that the communication should not be accepted, alternating between the first training of the second machine learning model with the first loss function on the first training data comprising one or more other counterfactual samples and the second training of the second machine learning model with the second loss function on the second training data comprising the first counterfactual sample.

15 . The one or more non-transitory computer-readable media of claim 12 , wherein alternating between the first training and the second training comprises, after the first machine learning model outputs the first output indicating that the communication should not be accepted, alternating between the first training of the second machine learning model with a binary cross-entropy function on the first training data and the second training of the second machine learning model with the second loss function on the second training data comprising the first counterfactual sample.

16 . The one or more non-transitory computer-readable media of claim 12 , wherein alternating between the first training and the second training comprises, after the first machine learning model outputs the first output indicating that the communication should not be accepted, alternating between the first training of the second machine learning model with a cross-entropy loss function on the first training data and the second training of the second machine learning model with the second loss function on the second training data comprising the first counterfactual sample.

17 . The one or more non-transitory computer-readable media of claim 12 , the operations further comprising:

in connection with the first training and the second training of the second machine learning model respectively with the first loss function and the second loss function, displaying, on a user interface, a first loss value corresponding to the first loss function, a second loss value corresponding to the second loss function, and a plurality of counterfactual samples.

18 . The one or more non-transitory computer-readable media of claim 12 , wherein:

executing the counterfactual sample generation system comprises executing the counterfactual sample generation system to generate a plurality of counterfactual samples for a plurality of data samples corresponding to communications that were not accepted; and

the second training data comprises the plurality of counterfactual samples.

19 . The one or more non-transitory computer-readable media of claim 12 , wherein the first training data corresponds to a first set of users, and the second training data corresponds to a second set of users that do not overlap with the first set of users.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Sep 30, 2022
From: SHARPE, SAMUEL; BRUSS, CHRISTOPHER BAYAN; BARR, BRIAN
To: CAPITAL ONE SERVICES, LLC
Reel/Frame 061277/0759 →
Continuity (1)
Related Publication 20240112092A1 · Apr 4, 2024
References Cited (49)
US 10861028B2 · Silberman · 2020 [cited by examiner]
US 11403538B1 · Verma · 2022 [cited by examiner]
US 11762950B1 · Stupp · 2023 [cited by examiner]
US 11922495B1 · Hernandez · 2024 [cited by examiner]
US 12254388B2 · McGrath · 2025 [cited by examiner]
US 12333775B2 · Alon · 2025 [cited by examiner]
US 20200097997A1 · Li · 2020 [cited by examiner]
US 20200279140A1 · Pai · 2020 [cited by examiner]
US 20200380571A1 · Ramakrishnan · 2020 [cited by examiner]
US 20210089895A1 · Munoz Delgado · 2021 [cited by examiner]
US 20210201184A1 · Scheepens · 2021 [cited by examiner]
US 20210287273A1 · Janakiraman · 2021 [cited by examiner]
US 20210326661A1 · Munoz Delgado · 2021 [cited by examiner]
US 20220012613A1 · Datta · 2022 [cited by examiner]
US 20220083871A1 · Nemirovsky · 2022 [cited by examiner]
US 20220114399A1 · Castiglione · 2022 [cited by examiner]
US 20220114464A1 · Yang · 2022 [cited by examiner]
US 20220130143A1 · Rodriguez Lopez · 2022 [cited by examiner]
US 20220147876A1 · Dalli · 2022 [cited by examiner]
US 20220188645A1 · Nia · 2022 [cited by examiner]
US 20220198498A1 · Huang · 2022 [cited by examiner]
US 20220207352A1 · Barr · 2022 [cited by examiner]
US 20220207353A1 · Barr · 2022 [cited by examiner]
US 20220253721A1 · Xu · 2022 [cited by examiner]
US 20220253733A1 · Zuo · 2022 [cited by examiner]
US 20220318639A1 · Upadhyay · 2022 [cited by examiner]
US 20220358594A1 · Zhu · 2022 [cited by examiner]
US 20220374782A1 · Spooner · 2022 [cited by examiner]
US 20220398460A1 · Dalli · 2022 [cited by examiner]
US 20230025731A1 · Vinov · 2023 [cited by examiner]
US 20230045950A1 · Gao · 2023 [cited by examiner]
US 20230080235A1 · Gil Ramos · 2023 [cited by examiner]
US 20230214695A1 · Wu · 2023 [cited by examiner]
GB 2589828A · 2021 [cited by examiner]
WO WO2022016556A1 · 2022 [cited by examiner]
Alaa, Ahmed M., Michael Weisz, and Mihaela Van Der Schaar. “Deep counterfactual networks with propensity-dropout.” arXiv preprint arXiv: 1706.05966 (2017). (Year: 2017). [cited by examiner]
Neal, Lawrence, et al. “Open set learning with counterfactual images.” Proceedings of the European conference on computer vision (ECCV). 2018. (Year: 2018). [cited by examiner]
White, Adam, and Artur d'Avila Garcez. “Measurable counterfactual local explanations for any classifier.” arXiv preprint arXiv: 1908.03020v2 (2019). (Year: 2019). [cited by examiner]
Hamon, Ronan, Henrik Junklewitz, and Ignacio Sanchez. “Robustness and explainability of artificial intelligence.” Publications Office of the European Union 207.40 (2020). (Year: 2020). [cited by examiner]
Mahajan, Divyat, Chenhao Tan, and Amit Sharma. “Preserving causal constraints in counterfactual explanations for machine learning classifiers.” arXiv preprint arXiv:1912.03277v3 (2020). (Year: 2020). [cited by examiner]
Wexler, James, et al. “The what-if tool: Interactive probing of machine learning models.” IEEE transactions on visualization and computer graphics 26.1 (2019): 56-65 (Year: 2019). [cited by examiner]
Poyiadzi, Rafael, et al. “FACE: feasible and actionable counterfactual explanations.” Proceedings of the AAAI/ACM Conference on AI, Ethics, and Society. 2020. (Year: 2020). [cited by examiner]
Van Looveren, Arnaud, and Janis Klaise. “Interpretable counterfactual explanations guided by prototypes.” Joint European Conference on Machine Learning and Knowledge Discovery in Databases. Cham: Springer International … [cited by examiner]
de Oliveira, Raphael Mazzine Barbosa, and David Martens. “A framework and benchmarking study for counterfactual generating methods on tabular data.” Applied Sciences 11.16 (2021): 7274. (Year: 2021). [cited by examiner]
Rasouli, Peyman, and Ingrid Chieh Yu. “Analyzing and improving the robustness of tabular classifiers using counterfactual explanations.” 2021 20th IEEE International Conference on Machine Learning and Applications (ICML… [cited by examiner]
Sujatha, C. N., et al. “Loan prediction using machine learning and its deployement on web application.” 2021 Innovations in Power and Advanced Computing Technologies (i-PACT). IEEE, 2021. (Year: 2021). [cited by examiner]
Tan, Juntao, et al. “Counterfactual Explainable Recommendation.” arXiv preprint arXiv:2108.10539v1 (2021). (Year: 2021). [cited by examiner]
Pham, David, and Yongfeng Zhang. “Counterfactual based reinforcement learning for graph neural networks.” Annals of Operations Research (2022): 1-17. (Year: 2022). [cited by examiner]
Yang, Fan, et al. “Generative counterfactuals for neural networks via attribute-informed perturbation.” ACM SIGKDD Explorations Newsletter 23.1 (2021): 59-68. (Year: 2021). [cited by examiner]