IP Library › Granted Patent US 12,525,354
Granted Patent B2
US 12,525,354 · App. 17/374,540 · Granted Jan 13, 2026

Machine learning techniques for future occurrence code prediction

Inventors: Rama Krishna Singh (Greater Noida, IN); Ravi Pande (Noida, IN); Priyank Jain (Noida, IN)
Assignee: Optum Technology, Inc.
G16H50/20G06F18/23G06F40/20G06N20/00
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,525,354
App. No.
17/374,540
Granted
Jan 13, 2026
Kind
B2
Abstract

Various embodiments of the present invention provide methods, apparatus, systems, computing devices, computing entities, and/or the like for performing predictive structural analysis. Certain embodiments of the present invention utilize systems, methods, and computer program products that perform predictive structural analysis using at least one of techniques using time bound code transition likelihood data objects, techniques using cross-code relationship values, techniques using augmented entity-code occurrence data objects, techniques using per-pathway text representations of inferred occurrence pathways of a one or more individual historic code occurrences, techniques using polygenic risk score (PRS) measures, and/or the like.

Claims (79)

1 . A computer-implemented method comprising:

identifying, by one or more processors, a plurality of time bound code transition likelihood data objects for an entity cluster that is associated with an identifiable data entity, wherein: (i) a time bound code transition likelihood data object of the plurality of time bound code transition likelihood data objects is associated with a defined time bound of a plurality of defined time bounds, (ii) the time bound code transition likelihood data object for the defined time bound describes, for a code pair comprising a first defined occurrence code of a plurality of defined occurrence codes and a second defined occurrence code of the plurality of defined occurrence codes, an inferred likelihood that the second defined occurrence code occurs within the defined time bound of an assumed occurrence of the first defined occurrence code, and (iii) the time bound code transition likelihood data object is generated for the defined time bound by:

identifying a plurality of per-entity code occurrences for the entity cluster, wherein a per-entity code occurrence of the plurality of per-entity code occurrences describes that a corresponding defined occurrence code occurs in relation to a corresponding identifiable data entity in the entity cluster at a corresponding timestamp, and

generating the time bound code transition likelihood data object based at least in part on the plurality of per-entity code occurrences;

generating, by the one or more processors and based at least in part on the plurality of time bound code transition likelihood data objects and a plurality of individual historic code occurrences for the identifiable data entity, a future occurrence code prediction, for the identifiable data entity, comprising a selected subset of the plurality of defined occurrence codes; and

initiating, by the one or more processors, one or more prediction-based actions based at least in part on the future occurrence code prediction.

2 . The computer-implemented method of claim 1 , further comprising:

determining an entity-code occurrence data object for the entity cluster based at least in part on the plurality of per-entity code occurrences; and

generating an augmented entity-code occurrence data object based at least in part on the entity-code occurrence data object.

3 . The computer-implemented method of claim 2 , wherein generating the augmented entity-code occurrence data object comprises:

performing singular value decomposition on the entity-code occurrence data object to generate a plurality of target lower-ranked entity-code occurrence data objects;

determining at least one missing value in the entity-code occurrence data object based at least in part on the plurality of target lower-ranked entity-code occurrence data objects; and

generating the augmented entity-code occurrence data object based at least in part on the at least one missing value.

4 . The computer-implemented method of claim 2 , wherein generating the time bound code transition likelihood data object based at least in part on the plurality of per-entity code occurrences comprises:

determining a plurality of augmented per-entity code occurrences for the entity cluster based at least in part on the plurality of per-entity code occurrences and the augmented entity-code occurrence data object;

determining a time bound code transition document that describes all co-occurrences of code pairs from the plurality of defined occurrence codes associated with the plurality of augmented per-entity code occurrences that occur within the defined time bound; and

generating the time bound code transition likelihood data object based at least in part on the time bound code transition document.

5 . The computer-implemented method of claim 4 , wherein generating the time bound code transition likelihood data object based at least in part on the time bound code transition document comprises:

determining a time bound code transition likelihood object for the code pair based at least in part on a per-document occurrence of the code pair as described by the time bound code transition document for the defined time bound and a cross-document occurrence of the code pair as described by the time bound code transition document; and

generating the time bound code transition likelihood data object based at least in part on the time bound code transition likelihood object.

6 . The computer-implemented method of claim 5 , wherein the time bound code transition likelihood object is a term-frequency-inverse-domain-frequency measure.

7 . The computer-implemented method of claim 1 , wherein generating the selected subset comprises:

generating a first subset of the plurality of defined occurrence codes based at least in part on the plurality of time bound code transition likelihood data objects and the plurality of individual historic code occurrences for the identifiable data entity;

generating a second subset of the plurality of defined occurrence codes based at least in part on a cross-code relationship value for a code pair; and

generating the selected subset based at least in part on the first subset and the second subset.

8 . The computer-implemented method of claim 1 , wherein generating the selected subset comprises:

determining, based at least in part on the plurality of time bound code transition likelihood data objects, a plurality of inferred occurrence pathways of the plurality of individual historic code occurrences;

generating a plurality of per-pathway text representations corresponding to the plurality of inferred occurrence pathways;

determining, based at least in part on a sequence of the plurality of per-pathway text representations and using a language-based machine learning model, an inferred subset of the plurality of defined occurrence codes for a predictive entity; and

generating the selected subset based at least in part on the inferred subset.

9 . A system comprising:

one or more processors; and

at least one memory storing processor-executable instructions that, when executed by any one or more of the one or more processors, cause the one or more processors to perform operations comprising:

identifying a plurality of time bound code transition likelihood data objects for an entity cluster that is associated with an identifiable data entity, wherein: (i) a time bound code transition likelihood data object of the plurality of time bound code transition likelihood data objects is associated with a defined time bound of a plurality of defined time bounds, (ii) the time bound code transition likelihood data object for the defined time bound describes, for a code pair comprising a first defined occurrence code of a plurality of defined occurrence codes and a second defined occurrence code of the plurality of defined occurrence codes, an inferred likelihood that the second defined occurrence code occurs within the defined time bound of an assumed occurrence of the first defined occurrence code, and (iii) the time bound code transition likelihood data object is generated for the defined time bound by:

identifying a plurality of per-entity code occurrences for the entity cluster, wherein a per-entity code occurrence of the plurality of per-entity code occurrences describes that a corresponding defined occurrence code occurs in relation to a corresponding identifiable data entity in the entity cluster at a corresponding timestamp, and

generating the time bound code transition likelihood data object based at least in part on the plurality of per-entity code occurrences;

generating, based at least in part on the plurality of time bound code transition likelihood data objects and a plurality of individual historic code occurrences for the identifiable data entity, a future occurrence code prediction, for the identifiable data entity, comprising a selected subset of the plurality of defined occurrence codes; and

initiating one or more prediction-based actions based at least in part on the future occurrence code prediction.

10 . The system of claim 9 , wherein the operations further comprise:

determining an entity-code occurrence data object for the entity cluster based at least in part on the plurality of per-entity code occurrences; and

generating an augmented entity-code occurrence data object based at least in part on the entity-code occurrence data object.

11 . The system of claim 10 , wherein generating the augmented entity-code occurrence data object comprises:

performing singular value decomposition on the entity-code occurrence data object to generate a plurality of target lower-ranked entity-code occurrence data objects;

determining at least one missing value in the entity-code occurrence data object based at least in part on the plurality of target lower-ranked entity-code occurrence data objects; and

generating the augmented entity-code occurrence data object based at least in part on the at least one missing value.

12 . The system of claim 10 , wherein generating the time bound code transition likelihood data object based at least in part on the plurality of per-entity code occurrences comprises:

determining a plurality of augmented per-entity code occurrences for the entity cluster based at least in part on the plurality of per-entity code occurrences and the augmented entity-code occurrence data object;

determining a time bound code transition document that describes all co-occurrences of code pairs from the plurality of defined occurrence codes associated with the plurality of augmented per-entity code occurrences that occur within the defined time bound; and

generating the time bound code transition likelihood data object based at least in part on the time bound code transition document.

13 . The system of claim 12 , wherein generating the time bound code transition likelihood data object based at least in part on the time bound code transition document comprises:

determining a time bound code transition likelihood object for the code pair based at least in part on a per-document occurrence of the code pair as described by the time bound code transition document for the defined time bound and a cross-document occurrence of the code pair as described by the time bound code transition document; and

generating the time bound code transition likelihood data object based at least in part on the time bound code transition likelihood object.

14 . The system of claim 13 , wherein the time bound code transition likelihood object is a term-frequency-inverse-domain-frequency measure.

15 . The system of claim 9 , wherein generating the selected subset comprises:

generating a first subset of the plurality of defined occurrence codes based at least in part on the plurality of time bound code transition likelihood data objects and the plurality of individual historic code occurrences for the identifiable data entity;

generating a second subset of the plurality of defined occurrence codes based at least in part on a cross-code relationship value for a code pair; and

generating the selected subset based at least in part on the first subset and the second subset.

16 . The system of claim 9 , wherein generating the selected subset comprises:

determining, based at least in part on the plurality of time bound code transition likelihood data objects, a plurality of inferred occurrence pathways of the plurality of individual historic code occurrences;

generating a plurality of per-pathway text representations corresponding to the plurality of inferred occurrence pathways;

determining, based at least in part on a sequence of the plurality of per-pathway text representations and using a language-based machine learning model, an inferred subset of the plurality of defined occurrence codes for a predictive entity; and

generating the selected subset based at least in part on the inferred subset.

17 . One or more non-transitory computer-readable storage media including instructions that, when executed by one or more processors, cause the one or more processors to perform operations comprising:

identifying a plurality of time bound code transition likelihood data objects for an entity cluster that is associated with an identifiable data entity, wherein: (i) a time bound code transition likelihood data object of the plurality of time bound code transition likelihood data objects is associated with a defined time bound of a plurality of defined time bounds, (ii) the time bound code transition likelihood data object for the defined time bound describes, for a code pair comprising a first defined occurrence code of a plurality of defined occurrence codes and a second defined occurrence code of the plurality of defined occurrence codes, an inferred likelihood that the second defined occurrence code occurs within the defined time bound of an assumed occurrence of the first defined occurrence code, and (iii) the time bound code transition likelihood data object is generated for the defined time bound by:

identifying a plurality of per-entity code occurrences for the entity cluster, wherein a per-entity code occurrence of the plurality of per-entity code occurrences describes that a corresponding defined occurrence code occurs in relation to a corresponding identifiable data entity in the entity cluster at a corresponding timestamp, and

generating the time bound code transition likelihood data object based at least in part on the plurality of per-entity code occurrences;

generating, based at least in part on the plurality of time bound code transition likelihood data objects and a plurality of individual historic code occurrences for the identifiable data entity, a future occurrence code prediction, for the identifiable data entity, comprising a selected subset of the plurality of defined occurrence codes; and

initiating one or more prediction-based actions based at least in part on the future occurrence code prediction.

18 . The one or more non-transitory computer-readable storage media of claim 17 , wherein the operations further comprise:

determining an entity-code occurrence data object for the entity cluster based at least in part on the plurality of per-entity code occurrences; and

generating an augmented entity-code occurrence data object based at least in part on the entity-code occurrence data object.

19 . The one or more non-transitory computer-readable storage media of claim 18 , wherein generating the augmented entity-code occurrence data object comprises:

performing singular value decomposition on the entity-code occurrence data object to generate a plurality of target lower-ranked entity-code occurrence data objects;

determining at least one missing value in the entity-code occurrence data object based at least in part on the plurality of target lower-ranked entity-code occurrence data objects; and

generating the augmented entity-code occurrence data object based at least in part on the at least one missing value.

20 . The one or more non-transitory computer-readable storage media of claim 18 , wherein generating the time bound code transition likelihood data object based at least in part on the plurality of per-entity code occurrences comprises:

determining a plurality of augmented per-entity code occurrences for the entity cluster based at least in part on the plurality of per-entity code occurrences and the augmented entity-code occurrence data object;

determining a time bound code transition document that describes all co-occurrences of code pairs from the plurality of defined occurrence codes associated with the plurality of augmented per-entity code occurrences that occur within the defined time bound; and

generating the time bound code transition likelihood data object based at least in part on the time bound code transition document.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jul 13, 2021
From: SINGH, RAMA KRISHNA; PANDE, RAVI; JAIN, PRIYANK
To: OPTUM TECHNOLOGY, INC.
Reel/Frame 056840/0936 →
Continuity (1)
Related Publication 20230017734A1 · Jan 19, 2023
References Cited (44)
US 6110109A · Hu et al. · 2000 [cited by applicant]
US 8700337B2 · Dudley et al. · 2014 [cited by applicant]
US 9910962B1 · Fakhrai-Rad et al. · 2018 [cited by applicant]
US 10867702B2 · Athey · 2020 [cited by examiner]
US 11322260B1 · Jain · 2022 [cited by examiner]
US 12154039B2 · Singh · 2024 [cited by examiner]
US 12165769B2 · Mitteldorf · 2024 [cited by examiner]
US 20060095853A1 · Amyot · 2006 [cited by examiner]
US 20080183454A1 · Barabasi · 2008 [cited by examiner]
US 20100070455A1 · Halperin et al. · 2010 [cited by applicant]
US 20150213225A1 · Amarasingham et al. · 2015 [cited by applicant]
US 20160283679A1 · Hu · 2016 [cited by examiner]
US 20180260925A1 · Gotz · 2018 [cited by examiner]
US 20190065689A1 · O'Malley · 2019 [cited by examiner]
US 20210343362A1 · Bryan · 2021 [cited by examiner]
US 20220028550A1 · Ng · 2022 [cited by examiner]
US 20220082574A1 · Watanabe · 2022 [cited by examiner]
US 20220093272A1 · Lerner · 2022 [cited by examiner]
US 20220188664A1 · Singh · 2022 [cited by examiner]
US 20220300835A1 · Ghosh · 2022 [cited by examiner]
US 20220383982A1 · Bridges · 2022 [cited by examiner]
US 20230016569A1 · Koyejo · 2023 [cited by examiner]
US 20230017734A1 · Singh · 2023 [cited by examiner]
US 20230225947A1 · Zuleta · 2023 [cited by examiner]
US 20240029896A1 · Nguyen · 2024 [cited by examiner]
CN 106778014B · 2020 [cited by applicant]
CN 113053503A · 2021 [cited by examiner]
CN 113537709A · 2021 [cited by examiner]
KR 102087613B1 · 2020 [cited by applicant]
WO 2020056389A1 · 2020 [cited by applicant]
Christensen et al., “Machine Learning Methods for Disease Prediction with Claims Data”, 2018, IEEE International Conference of Healthcare Informatics, pp. 467-471. (Year: 2018). [cited by examiner]
Wartelle et al., “Clustering of a Health Dataset Using Diagnosis Co-Occurrences”, Mar. 7, 2021, Applied Sciences, pp. 1-19 (Year: 2021). [cited by examiner]
Cao et al., “Mining a Clinical Data Warehouse to discover Disease-finding Associations using Co-Ooccurrence Statistics”, AMIA 2005 Symposium Proceedings, pp. 106-110. (Year: 2005). [cited by examiner]
Hanaeur et al., “Modeling temporal relationships in large scale clinical associations”, Sep. 17, 2012, Journal Am Med Inform Association, pp. 1-10 (Year: 2012). [cited by examiner]
Razavian et al., “Temporal Convolutional Neural Networks for Diagnosis from Lab Tests”, Mar. 11, 2016, arXiv.com, pp. 1-19 (Year: 2016). [cited by examiner]
“icd10-E083: Diabetes Mellitus Due to Underlying Condition With Ophthalmic Complications,” 1UPhEALTH, (article, online), (3 pages), [Retrieved from the Internet Sep. 18, 2021] <https://1up.health/health-data/icd10/id/E0… [cited by applicant]
“Understanding the ICD-10 Code Structure,” Health Network Solutions, (article, online), (6 pages), [Retrieved from the Internet Sep. 18, 2021] <https://www.healthnetworksolutions.net/index.php/understanding-the-icd-10-c… [cited by applicant]
Chawla, Nitesh V. et al. “Bringing Big Data To Personalized Healthcare: A Patient-Centered Framework,” Journal of General Internal Medicine, Sep. 1, 2013, vol. 28, No. pp. S660-S665, (Published Online: Jun. 25, 2013). [cited by applicant]
Choi, Edward et al. “Multi-Layer Representation Learning For Medical Concepts,” arXiv Preprint arXiv:1602.05568v1, Feb. 17, 2016, pp. 1-20. [cited by applicant]
Futoma, Joseph et al. “Predicting Disease Progression with a Model for Multivariate Longitudinal Clinical Data,” Proceedings of Machine Learning For Healthcare, vol. 56, Dec. 10, 2016, (12 pages). [cited by applicant]
Hane, Christopher A. et al. “Predicting Onset of Dementia Using Clinical Notes and Machine Learning: Case-Control Study,” Journal of Medical Internet Research, vol. 8, No. 6:e17819, Jun. 3, 2020, DOI: 10.2196/17819, PMI… [cited by applicant]
Kartchner, David et al. “Code2Vec: Embedding and Clustering Medical Diagnosis Data,” 2017 IEEE International Conference on Healthcare Informatics (ICHI), Aug. 23, 2017, DOI: 10.1109/ICHI.2017.94. [cited by applicant]
Kaushik, Kulvaibhav et al. “Disease Management: Clustering-Based Disease Prediction,” International Journal Of Collaborative Enterprise, vol. 4, Nos. 102, Jan. 2014, pp. 69-82, DOI: 10.1504/IJCENT.2014.065047. [cited by applicant]
Krzanowski, W.J. “Missing Value Imputation In Multivariate Data Using The Singular Value Decomposition Of A Matrix,” Listy Biometryczne—Biometrical Letters, vol. XXV, No. 1,2, (1988), pp. 31-39. [cited by applicant]