IP Library › Granted Patent US 12,731,007
Granted Patent B2
US 12,731,007 · App. 18/331,600 · Granted Sep 8, 2026

Temporal sequence causal transformer machine learning model

Inventors: Dominik Roman Christian Dahlem (Dublin, IE); Vijay S. Nori (Roswell, GA); Eran Halperin (Santa Monica, CA); Nadav Rakocz (Los Angeles, CA)
Assignee: UnitedHealth Group Incorporated
G06N3/0455G06N3/047G06N3/092G06N20/00G16H10/60
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,731,007
App. No.
18/331,600
Granted
Sep 8, 2026
Kind
B2
Abstract

Various embodiments of the present disclosure provide methods, apparatus, systems, computing devices, computing entities, and/or the like for generating a prediction output comprising one or more actions by receiving data associated with encounters in a tuple form, tokenizing the encounters, training a causal transformer machine learning model configured to predict outcomes of actions by translating action tokens from the tokenized encounters into one or more embedding spaces, and training a causal transformer machine learning model to select the one or more actions based on embeddings from the one or more embedding spaces.

Claims (56)

1 . A computer-implemented method comprising:

receiving, by one or more processors, an input temporal sequence, wherein:

(i) the input temporal sequence comprises a set of one or more input tuples,

(ii) an input tuple of the set of one or more input tuples comprises a plurality of tuple data objects comprising data representative of (a) one or more states, (b) one or more combinations of actions, (c) one or more outcomes, and (d) a cumulative discounted future outcome, and

(iii) the cumulative discounted future outcome is associated with the input tuple;

generating, by the one or more processors, a plurality of input tokens associated with the plurality of tuple data objects, wherein the plurality of input tokens are generated according to a plurality of respective tuple data object types associated with the plurality of tuple data objects;

generating, by the one or more processors and using a causal transformer machine learning model, a prediction output based on the plurality of input tokens and a conditional distribution of actions, the prediction output comprising a plurality of output tokens, wherein training the causal transformer machine learning model comprises:

(a) projecting a plurality of training tokens into a plurality of respective embedding spaces using an embedding layer, wherein (i) at least one of the plurality of respective embedding spaces comprises a plurality of embedding sets associated with the plurality of training tokens, (ii) the plurality of embedding sets comprises a temporal embedding set, a structural embedding set, and a positional embedding set, (iii) the plurality of training tokens is associated with a plurality of training temporal sequences, and (iv) a training temporal sequence of the plurality of training temporal sequences comprises a set of one or more training tuples, wherein a training tuple of the set of one or more training tuples comprises a plurality of training tuple data objects comprising data representative of one or more training states, one or more training combinations of actions, one or more training outcomes, and a training cumulative discounted future outcome associated with the training tuple,

(b) inputting the plurality of respective embedding spaces into the causal transformer machine learning model, and

(c) for the training tuple, generating a context dependent representation based on one or more of the plurality of respective embedding spaces associated with one or more sequentially prior tuples of the set of one or more training tuples with respect to the training tuple;

generating, by the one or more processors, one or more policy scores based on the prediction output; and

initiating, by the one or more processors, the performance of one or more prediction-based actions based on the one or more policy scores and the prediction output.

2 . The computer-implemented method of claim 1 , wherein the temporal embedding set comprises one or more embeddings associated with a relative time between the training tuple and a sequentially first training tuple of the set of one or more training tuples.

3 . The computer-implemented method of claim 1 , wherein the structural embedding set comprises one or more embeddings associated with the plurality of respective tuple data object types of the plurality of training tuple data objects.

4 . The computer-implemented method of claim 1 , wherein the positional embedding set comprises one or more embeddings associated with a sequential position of the training tuple.

5 . The computer-implemented method of claim 1 further comprising discarding at least one sequentially first training tuple of the set of one or more training tuples from a subset of training temporal sequences of the plurality of training temporal sequences.

6 . The computer-implemented method of claim 1 , wherein the causal transformer machine learning model is trained based on teacher-forcing training by using one or more ground-truth tokens as training feedback input to the causal transformer machine learning model.

7 . The computer-implemented method of claim 1 , wherein the plurality of output tokens comprises one or more output state tokens, one or more output action tokens, one or more output outcome tokens, and one or more output cumulative discounted future outcome tokens.

8 . The computer-implemented method of claim 7 , wherein generating the prediction output further comprises generating one or more log- likelihood scores of one or more output actions associated with the one or more output action tokens, the one or more log-likelihood scores representative of a likelihood of the one or more output actions most likely to follow based on the input temporal sequence.

9 . The computer-implemented method of claim 7 , wherein generating the prediction output further comprises generating one or more predictive scores, the one or more predictive scores comprising (i) one or more action predictive scores of one or more output actions associated with the one or more output action tokens based on the one or more states, and (ii) one or more outcome predictive scores associated with one or more output cumulative discounted future outcomes associated with the one or more output cumulative discounted future outcome tokens based on the one or more output actions.

10 . The computer-implemented method of claim 9 , wherein generating the prediction output further comprises generating one or more expected predicted outcomes based on the one or more output cumulative discounted future outcomes and the one or more predictive scores.

11 . The computer-implemented method of claim 7 , wherein initiating the performance of the one or more prediction-based actions further comprises selecting one or more output actions associated with the one or more output action tokens based on the one or more policy scores.

12 . The computer-implemented method of claim 1 , further comprising excluding one or more action combination tokens of a plurality of action combination tokens from the conditional distribution of actions based on the one or more action combination tokens comprising one or more probability scores that are below a threshold.

13 . The computer-implemented method of claim 1 , wherein generating the plurality of input tokens further comprises:

receiving an action space data object comprising a plurality of possible individual actions; and

assigning a plurality of action combination tokens to a plurality of combinations comprising selected ones a subset of possible individual actions of the plurality of possible individual actions.

14 . The computer-implemented method of claim 1 further comprising generating the conditional distribution of actions based on the one or more states.

15 . A system comprising one or more processors and

one or more non-transitory computer readable media storing processor-executable instructions that, when executed by the one or more processors, cause the one or more processors to perform operations comprising:

receiving an input temporal sequence, wherein:

(i) the input temporal sequence comprises a set of one or more input tuples,

(ii) an input tuple of the set of one or more input tuples comprises a plurality of tuple data objects comprising data representative of (a) one or more states, (b) one or more combinations of actions, (c) one or more outcomes, and (d) a cumulative discounted future outcome, and

(iii) the cumulative discounted future outcome is associated with the input tuple;

generating a plurality of input tokens associated with the plurality of tuple data objects, wherein the plurality of input tokens are generated according to a plurality of respective tuple data object types associated with the plurality of tuple data objects;

generating, using a causal transformer machine learning model, a prediction output based on the plurality of input tokens and a conditional distribution of actions, the prediction output comprising a plurality of output tokens, wherein training the causal transformer machine learning model comprises:

(a) projecting a plurality of training tokens into a plurality of respective embedding spaces using an embedding layer, wherein (i) at least one of the plurality of respective embedding spaces comprises a plurality of embedding sets associated with the plurality of training tokens, (ii) the plurality of embedding sets comprises a temporal embedding set, a structural embedding set, and a positional embedding set, (iii) the plurality of training tokens is associated with a plurality of training temporal sequences, and (iv) a training temporal sequence of the plurality of training temporal sequences comprises a set of one or more training tuples, wherein a training tuple of the set of one or more training tuples comprises a plurality of training tuple data objects comprising data representative of one or more training states, one or more training combinations of actions, one or more training outcomes, and a training cumulative discounted future outcome associated with the training tuple,

(b) inputting the plurality of respective embedding spaces into the causal transformer machine learning model, and

(c) for the training tuple, generating a context dependent representation based on one or more of the plurality of respective embedding spaces associated with one or more sequentially prior tuples of the set of one or more training tuples with respect to the training tuple;

generating one or more policy scores based on the prediction output; and

initiating the performance of one or more prediction-based actions based on the one or more policy scores and the prediction output.

16 . The system of claim 15 , wherein the plurality of output tokens comprises one or more output state tokens, one or more output action tokens, one or more output outcome tokens, and one or more output cumulative discounted future outcome tokens.

17 . The system of claim 16 , wherein the operations further comprise generating the prediction output by generating one or more log-likelihood scores of one or more output actions associated with the one or more output action tokens, the one or more log-likelihood scores representative of a likelihood of the one or more output actions most likely to follow based on the input temporal sequence.

18 . The system of claim 16 , wherein the operations further comprise generating the prediction output by generating one or more predictive scores, the one or more predictive scores comprising (i) one or more action predictive scores of one or more output actions associated with the one or more output action tokens based on the one or more states, and (ii) one or more outcome predictive scores associated with one or more output cumulative discounted future outcomes associated with the one or more output cumulative discounted future outcome tokens based on the one or more output actions.

19 . The system of claim 18 , wherein the operations further comprise generating the prediction output by generating one or more expected predicted outcomes based on the one or more output cumulative discounted future outcomes and the one or more predictive scores.

20 . One or more non-transitory computer-readable storage media storing instructions that, when executed by one or more processors, cause the one or more processors to perform operations comprising:

receiving an input temporal sequence, wherein:

(i) the input temporal sequence comprises a set of one or more input tuples,

(ii) an input tuple of the set of one or more input tuples comprises a plurality of tuple data objects comprising data representative of (a) one or more states, (b) one or more combinations of actions, (c) one or more outcomes, and (d) a cumulative discounted future outcome, and

(iii) the cumulative discounted future outcome is associated with the input tuple;

generating a plurality of input tokens associated with the plurality of tuple data objects, wherein the plurality of input tokens are generated according to a plurality of respective tuple data object types associated with the plurality of tuple data objects;

generating, using a causal transformer machine learning model, a prediction output based on the plurality of input tokens and a conditional distribution of actions, the prediction output comprising a plurality of output tokens, wherein training the causal transformer machine learning model comprises:

(a) projecting a plurality of training tokens into a plurality of respective embedding spaces using an embedding layer, wherein (i) at least one of the plurality of respective embedding spaces comprises a plurality of embedding sets associated with the plurality of training tokens, (ii) the plurality of embedding sets comprises a temporal embedding set, a structural embedding set, and a positional embedding set, (iii) the plurality of training tokens is associated with a plurality of training temporal sequences, and (iv) a training temporal sequence of the plurality of training temporal sequences comprises a set of one or more training tuples, wherein a training tuple of the set of one or more training tuples comprises a plurality of training tuple data objects comprising data representative of one or more training states, one or more training combinations of actions, one or more training outcomes, and a training cumulative discounted future outcome associated with training tuple,

(b) inputting the plurality of respective embedding spaces into the causal transformer machine learning model, and

(c) for the training tuple, generating a context dependent representation based on one or more of the plurality of respective embedding spaces associated with one or more sequentially prior tuples of the set of one or more training tuples with respect to the training tuple;

generating one or more policy scores based on the prediction output; and

initiating the performance of one or more prediction-based actions based on the one or more policy scores and the prediction output.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jun 8, 2023
From: DAHLEM, DOMINIK ROMAN CHRISTIAN; NORI, VIJAY S.; HALPERIN, ERAN; RAKOCZ, NADAV
To: UNITEDHEALTH GROUP INCORPORATED
Reel/Frame 063897/0945 →
Continuity (2)
Provisional Application 63384941 · Nov 23, 2022
Related Publication 20240169264A1 · May 23, 2024
References Cited (76)
US 5550734A · Tarter et al. · 1996 [cited by applicant]
US 6073104A · Field · 2000 [cited by applicant]
US 8370172B2 · Shell et al. · 2013 [cited by applicant]
US 9697469B2 · Mcmahon et al. · 2017 [cited by applicant]
US 10133250B2 · Kohn et al. · 2018 [cited by applicant]
US 10366451B2 · Kaznady · 2019 [cited by applicant]
US 10755816B2 · Bennett et al. · 2020 [cited by applicant]
US 10885432B1 · Dulac-Arnold · 2021 [cited by examiner]
US 11107159B2 · Hails et al. · 2021 [cited by applicant]
US 11177025B2 · Bettencourt-Silva et al. · 2021 [cited by applicant]
US 11295861B2 · Al Hasan et al. · 2022 [cited by applicant]
US 11490966B2 · Roche et al. · 2022 [cited by applicant]
US 11600387B2 · Peng et al. · 2023 [cited by applicant]
US 11615208B2 · Truong et al. · 2023 [cited by applicant]
US 12334192B2 · Singal · 2025 [cited by examiner]
US 12626050B2 · Mallinson et al. · 2026 [cited by applicant]
US 20050216315A1 · Andersson · 2005 [cited by applicant]
US 20070073685A1 · Thibodeau et al. · 2007 [cited by applicant]
US 20090192831A1 · Raynes et al. · 2009 [cited by applicant]
US 20120324113A1 · Prince et al. · 2012 [cited by applicant]
US 20130275144A1 · Bain · 2013 [cited by applicant]
US 20140081652A1 · Klindworth · 2014 [cited by applicant]
US 20150339769A1 · Deoliveira et al. · 2015 [cited by applicant]
US 20170024643A1 · Lillicrap · 2017 [cited by examiner]
US 20170076201A1 · van Hasselt · 2017 [cited by examiner]
US 20170091861A1 · Bianchi et al. · 2017 [cited by applicant]
US 20170286622A1 · Cox et al. · 2017 [cited by applicant]
US 20180260891A1 · Merrill et al. · 2018 [cited by applicant]
US 20190372644A1 · Chen · 2019 [cited by examiner]
US 20200134716A1 · Lahrichi et al. · 2020 [cited by applicant]
US 20200226690A1 · Gulati et al. · 2020 [cited by applicant]
US 20200286056A1 · Martino et al. · 2020 [cited by applicant]
US 20210158085A1 · Budzik · 2021 [cited by applicant]
US 20210202055A1 · Xiao et al. · 2021 [cited by applicant]
US 20210295130A1 · Rasoolinejad · 2021 [cited by examiner]
US 20210319054A1 · Glass et al. · 2021 [cited by applicant]
US 20210387330A1 · Mavrin · 2021 [cited by examiner]
US 20220019888A1 · Aggarwal et al. · 2022 [cited by applicant]
US 20220084664A1 · Ginsburg · 2022 [cited by applicant]
US 20220374608A1 · Shazeer et al. · 2022 [cited by applicant]
US 20220398460A1 · Dalli et al. · 2022 [cited by applicant]
US 20230040705A1 · Loganathan et al. · 2023 [cited by applicant]
US 20230121711A1 · Chhaya et al. · 2023 [cited by applicant]
US 20240028907A1 · Shi et al. · 2024 [cited by applicant]
US 20240104379A1 · Laskin · 2024 [cited by examiner]
Chen, X., Yao, L., McAuley, J., Zhou, G., & Wang, X. (2021). A survey of deep reinforcement learning in recommender systems: A systematic review and future directions. arXiv preprint arXiv:2109.03540. (Year: 2021) . [cited by examiner]
Chen, Lilli et al. “Decision Transformer: Reinforcement Learning via Sequence Modeling”, Advances in Neural Information Processing Systems, 21 pages, Jun. 24, 2021, https://arxiv.org/abs/2106.01345. [cited by applicant]
Cortaredona, Sébastien et al. “The Extra Cost of Comorbidity: Multiple Illnesses and the Economic Burden of Non-Communicable Diseases”, BMC Medicine, 17 pages, Dec. 8, 2017, https://bmcmedicine.biomedcentral.com/article… [cited by applicant]
Festor, Paul et al. “Assuring the Safety of AI-based Clinical Decision Support Systems: A Case Study of the AI Clinician for Sepsis Treatment,” BMJ Health and Care Informatics, 9 pages, Jul. 17, 2022, ISSN 2632-1009. Ht… [cited by applicant]
Fringuellotti et al., “Insurance Companies and the Growth of Corporate Loans' Securitization”, Federal Reserve Bank of New York, 67 pages, Aug. 2021, http://hdl.handle.net/10419/247898. [cited by applicant]
Hernan, Miguel A. et al. “Causal Inference”, 102 pages, Mar. 19, 2014, ISBN 978-1-315-37493-2. [cited by applicant]
Hu, Shengchao, et al. “On Transforming Reinforcement Learning by Transformer: The Development Trajectory”, arXiv preprint arXiv:2212.14164, 26 pages, Jan. 20, 2023, https://arxiv.org/abs/2212.14164. [cited by applicant]
Huang, Yong et al., “Reinforcement Learning For Sepsis Treatment: A Continuous Action Space Solution”, Machine Learning for Healthcare, 17 pages, Dec. 31, 2022, https://www.mlforhc.org/s/93Reinforcement_learning_for_sep… [cited by applicant]
Humphrey, Kyle, “Using Reinforcement Learning to Personalize Dosing Strategies in a Simulated Cancer Trial with High Dimensional Data”, The University of Arizona, 50 pages, (2017), https://repository.arizona.edu/handle/… [cited by applicant]
Imbens, Guido W., “The Role of the Propensity Score in Estimating Dose-Response Functions”, Biometrika, vol. 87, No. 3, 22 pages, Sep. 2000, https://www.jstor.org/stable/2673642. [cited by applicant]
Janner, Michael et al., “Offline Reinforcement Learning as One Big Sequence Modeling Problem”, Advances in Neural Information Processing Systems, vol. 4, 17 pages, Nov. 29, 2021, https://arxiv.org/abs/2106.02039. [cited by applicant]
Komorowski, Matthleu, et al. “The Artificial Intelligence Clinician Learns Optimal Treatment Strategies for Sepsis in Intensive Care”, Nature Medicine, vol. 24, 11 pages, (2018), https://doi.org/10.1038/s41591-018-0213-… [cited by applicant]
Krajna, Agneza et al., “Explainability in Reinforcement Learning: Perspective and Position,” arXiv preprint arXiv:2203.11547, vol. 1, 18 pages, Mar. 22, 2022, https://arxiv.org/pdf/2203.11547. [cited by applicant]
Liu, Mingyang et al. “Deep Reinforcement Learning for Personalized Treatment Recommendation,” Statistics in Medicine, vol. 41, 23 pages, May 22, 2022, https://onlinelibrary.wiley.com/doi/epdf/10.1002/sim.9491. [cited by applicant]
Michalowski, Martin, et al. “MitPlan 2.0: Enhanced Support for Multi-morbid Patient Management Using Planning”, Artificial Intelligence in Medicine, 11 pages, (2021), https://doi.org/10.1007/978-3-030-77211-6_31. [cited by applicant]
Peng, Xuefeng et al. “Improving Sepsis Treatment Strategies by Combining Deep and Kernel-Based Reinforcement Learning”, MIT, Institute for Medical Engineering & Science, 10 pages, Dec. 5, 2018, https://www.ncbi.nlm.nih.… [cited by applicant]
Petropoulos ,Anastasios, et al. “A Robust Machine Learning Approach for Credit Risk Analysis of Large Loan Level Datasets Using Deep Learning and Extreme Gradient Boosting”, Irving Fisher Committee On Central Bank Stati… [cited by applicant]
Prasad, Niranjani et al. “A Reinforcement Learning Approach to Weaning of Mechanical Ventilation in Intensive Care Units”, arXiv preprint arXiv:1704.06300, 10 pages, Apr. 20, 2017, https://arxiv.org/abs/1704.06300. [cited by applicant]
Raghu, Aniruddh et al. “Deep Reinforcement Learning for Sepsis Treatment”, Machine Learning For Health, vol. 1, 9 pages, Nov. 27, 2017, https://arxiv.org/abs/1711.09602. [cited by applicant]
Rosenbaum, Paul et al. “The Central Role of the Propensity Score in Observational Studies for Causal Effects”, Biometrika, vol. 70, Issue 1, pp. 41-55, Apr. 1, 1983, https://academic.oup.com/biomet/article/70/1//41/2408… [cited by applicant]
Tang, Shengpu, et al. “Leveraging Factored Action Spaces for Efficient Offline Reinforcement Learning in Healthcare,” Advances in Neural Information Processing Systems, vol. 35, 15 pages, (2022), https://papers.nips.cc/… [cited by applicant]
Tseng, Huan-Hsin et al. “Deep Reinforcement Learning for Automated Radiation Adaptation in Lung Cancer,” American Association of Physicists in Medicine, vol. 44, 16 pages, Dec. 2017, ISSN 2473-4209. doi: 10.1002/mp.1262… [cited by applicant]
Watts, Jeremy, et al. “Optimizing Individualized Treatment Planning for Parkinson's Disease Using Deep Reinforcement Learning”, 42nd Annual International Conference of the IEEE Engineering in Medicine & Biology Society,… [cited by applicant]
Wennberg, John E., et al. “Unwarranted Variations in Healthcare Delivery: Implications for Academic Medical Centres”, National Library of Medicine, vol. 325, 4 pages, Oct. 26, 2002, https://www.ncbi.nlm.nih.gov/pmc/arti… [cited by applicant]
Yauney, Gregory, et al. “Reinforcement Learning with Action-Derived Rewards for Chemotherapy and Clinical Trial Dosing Regimen Selection”, Machine Learning for Healthcare, 65 pages, (2018), https://proceedings.mlr.press… [cited by applicant]
Zhao, Siyan, et al. “Decision Stacks: Flexible Reinforcement Learning via Modular Generative Models”, NeurIPS, 18 pages, Oct. 29, 2023, https://arxiv.org/abs/2306.06253. [cited by applicant]
Zhao, Yufan, et al. “Reinforcement Learning Strategies for Clinical Trials in Nonsmall Cell Lung Cancer”, Biometrics, vol. 67, 22 pages, Dec. 2011, https://www.ncbi.nlm.nih.gov/pmc/articles/PMC3138840/. [cited by applicant]
Ho Oh, et al., “Reinforcement Learning-Based Expanded Personalized Diabetes Treatment Recommendation Using South Korea Electronic Health Records”, Expert Systems with Applications, vol. 206, Jun. 23, 2022, (12 pages), d… [cited by applicant]
Wallace, et al., “Optum Labs: Building A Novel Node In The Learning Health Care System”, Health Affairs, vol. 33, No. 7, pp. 1187-1194, (2014), doi: 10.1377/hlthaff.2014.0038. [cited by applicant]
Melnychuk, Valentyn, et al., “Causal Transformer for Estimating Counterfactual Outcomes”, Proceedings of the 39th International Conference on Machine Learning, Jul. 17-23, 2022, 37 pages, Baltimore, Maryland. [cited by applicant]
Non-Final Rejection Mailed on Jun. 30, 2026 for U.S. Appl. No. 18/492,189, 44 page(s). [cited by applicant]