IP Library › Granted Patent US 12,266,431
Granted Patent B2
US 12,266,431 · App. 17/221,981 · Granted Apr 1, 2025

Machine learning engine and rule engine for document auto-population using historical and contextual data

Inventors: Tanmoy Mukherjee (Kolkata, IN); Shrilata Mondal (Guskara Bardhaman, IN); Debendra Kar (Marathalli Bangalore, IN); Damodar Reddy Karra (Telangana, IN)
Assignee: Cerner Innovation, Inc.
G16H10/60G06F40/174G06F40/20
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,266,431
App. No.
17/221,981
Filed
Apr 5, 2021
Granted
Apr 1, 2025
Kind
B2
Art Unit
3687
USPC
705/2
Abstract

Methods, systems, and computer-readable media are disclosed herein for a machine learning engine and rule engine that leverage historical and contextual data to intelligently identify, score, and suggest one or more documents for auto-population of a graphical user interface. The machine learning and rule engine employ vectorization and clustering technique to identify, score, and suggest the most factually accurate and contextually relevant documents as selectable candidates for electronic documentation. Further, the selection, rejection, or modification of the candidate documents are ingested by the machine learning engine and/or the rule engine to update a clustering algorithm and/or to update a relevance scoring algorithm, which are then utilized for subsequent instances.

Claims (90)

1. One or more non-transitory media having instructions that, when executed by one or more processors, cause the one or more processors to facilitate a plurality of operations, the operations comprising:

generating a combined vector, associated with healthcare information, based on a first set of historical text data and based further on a second set of historical contextual data, the first set of historical text data differing from the second set of historical contextual data;

generating a plurality of clusters based on an electronic machine-learning clustering model using the combined vector as input, wherein:

(a) the electronic machine-learning clustering model is trained based on data associated with instances of first information selected from a group comprising the first set of historical text data and the second set of historical contextual data, and

(b) the plurality of clusters includes the first set of historical text data and further include the second set of historical contextual data;

identifying a primary cluster in the plurality of clusters that is a best match to the combined vector;

performing natural language processing on second information associated with the primary cluster;

based at least on the natural language processing:

assigning a relevance score, for at least a portion of a plurality of text blocks in the primary cluster, corresponding to a respective relevance of at least a portion of the plurality of text blocks to information associated with at least a portion of the combined vector;

identifying a primary text block in the plurality of text blocks of the primary cluster having a highest relevance score; and

in response to the identifying of the primary text block in the plurality of text blocks of the primary cluster having the highest relevance score:

communicating the primary text block for display as a recommended selection for automatic population into at least a narrative field, of a healthcare related electronic document, configured to receive free-form narrative text.

2. The one or more non-transitory media of claim 1 , wherein generating the combined vector comprises:

identifying categorical data in one or more of the first set of historical text data or the second set of historical contextual data, wherein the categorical data is incompatible for input to the electronic machine-learning clustering model; and

transforming the categorical data into dummy variables, wherein the dummy variables are used to generate the combined vector.

3. The one or more non-transitory media of claim 1 , wherein the operations further comprise reducing data sparsity of the combined vector by:

applying Principle Component Analysis (PCA) to the combined vector to identify sparse columns of the combined vector, wherein the combined vector has n dimensions, one dimension for each of a plurality of features; and

projecting the combined vector having n dimensions to a k dimensional vector space to reduce a total quantity of the sparse columns, wherein k<n.

4. The one or more non-transitory media of claim 1 , wherein the electronic machine-learning clustering model includes a Density-Based Spatial Clustering of Applications with Noise (DBSCAN).

5. The one or more non-transitory media of claim 1 , wherein when generating the plurality of clusters, a total quantity of the plurality of clusters is determined by one or more characteristics of the first set of text data and the second set of historical contextual data in the combined vector, wherein a plurality of cluster tags are attached to the combined vector, and wherein the plurality of cluster tags map features of the first set of text data and the second set of historical contextual data to one or more of the plurality of clusters.

6. The one or more non-transitory media of claim 1 , wherein the plurality of clusters includes one or more portions of physician-specific patient-encounter data or patient-specific context data.

7. The one or more non-transitory media of claim 6 , wherein the primary cluster is identified, subsequent to reducing a data sparsity, by:

identifying one cluster in the plurality of clusters that corresponds to a set of text blocks in one or more of the physician-specific patient-encounter data or the patient-specific context data, wherein the set of text blocks in the one cluster have a greatest similarity to the combined vector relative to other clusters in the plurality of clusters; and

designating the one cluster as the primary cluster based on the one cluster have the greatest similarity to the combined vector relative.

8. The one or more non-transitory media of claim 1 , wherein the primary text block includes text selected from a group comprising unstructured text, text other than a medical or billing code, and text other than a medical concept.

9. The one or more non-transitory media of claim 1 , wherein the narrative field is configured to receive narrative text from a clinician describing a patient encounter, and wherein the primary text block includes unstructured text.

10. The one or more non-transitory media of claim 1 , wherein the primary text block comprises a narrative note.

11. The one or more non-transitory media of claim 1 , wherein the primary text block includes a narrative note describing a patient encounter.

12. The one or more non-transitory media of claim 1 , wherein the primary text block includes a narrative note created by a clinician and describing a patient encounter.

13. The one or more non-transitory media of claim 1 , wherein the primary text block includes content other than a medical code, billing code, or medical concept.

14. The one or more non-transitory media of claim 1 , wherein the primary text block includes content other than a medical concept.

15. The one or more non-transitory media of claim 1 , wherein the primary text block includes first text of a first type and second text of a second type, wherein the first type excludes structured text, medical and billing codes, and medical concepts, and wherein the second type includes content selected from a group comprising structured text, a medical or billing code, and a medical concept.

16. The one or more non-transitory media of claim 1 , wherein the primary text block includes first text of a first type and second text of a second type, wherein the first type excludes structured text, and wherein the second type includes structured text.

17. The one or more non-transitory media of claim 1 , wherein generating the combined vector comprises generating a physician-specific vector from the first set of historical text data and a patient-specific vector from the second set of historical contextual data.

18. The one or more non-transitory media of claim 17 , wherein the combined vector is generated based on the physician-specific vector and the patient-specific vector.

19. The one or more non-transitory media of claim 1 , wherein the primary text block includes a narrative input by a clinician documenting reasons and details regarding a clinical encounter with a patient, including a medical history of the patient and a chief complaint of the patient.

20. The one or more non-transitory media of claim 1 , wherein the primary text block includes first text of a first type and second text of a second type, wherein the first type excludes structured text, medical and billing codes, and medical concepts, and wherein the second type includes a structured text element, a medical or billing code, and a medical concept.

21. The one or more non-transitory media of claim 1 , wherein the primary text block includes first text of a first type and second text of a second type, wherein the first type excludes structured text, medical and billing codes, and medical concepts, and wherein the second type includes a medical concept.

22. The one or more non-transitory media of claim 1 , wherein the primary text block includes a plurality of elements selected from a group comprising: a clinical fact, a clinical observation, and a contextual data item.

23. The one or more non-transitory media of claim 1 , wherein the primary text block includes a medical assessment describing a patient encounter.

24. The one or more non-transitory media of claim 1 , wherein the operations further comprise decreasing a dimensionality of the combined vector prior to inputting the combined vector to the electronic machine-learning clustering model to generate the plurality of clusters.

25. The one or more non-transitory media of claim 1 , wherein the operations further comprise communicating the primary text block for display as a recommended selection for automatic population of a text assessment field of an electronic document.

26. The one or more non-transitory media of claim 1 , wherein the primary text block includes a plurality of clinical assessments, and wherein the primary text block is communicated for display as a recommended selection for automatic population of a text narrative field of an electronic document.

27. The one or more non-transitory media of claim 1 , wherein the operations further comprise:

retraining the electronic machine-learning clustering model based on additional data associated with the instances; and

reinputting to the retrained electronic machine-learning clustering model information associated with an additional combined vector to generate an additional plurality of clusters.

28. A system having one or more processors configured to facilitate a plurality of operations, the operations comprising:

generating a combined vector, associated with healthcare information, based on a first set of historical text data and based further on a second set of historical contextual data, the first set of historical text data differing from the second set of historical contextual data;

generating a plurality of clusters based on an electronic machine-learning clustering model using the combined vector as input, wherein:

(a) the electronic machine-learning clustering model is trained based on data associated with instances of first information selected from a group comprising the first set of historical text data and the second set of historical contextual data, and

(b) the plurality of clusters includes the first set of historical text data and further include the second set of historical contextual data;

identifying a primary cluster in the plurality of clusters that is a best match to the combined vector;

performing natural language processing on second information associated with the primary cluster;

based at least on the natural language processing:

assigning a relevance score, for at least a portion of a plurality of text blocks in the primary cluster, corresponding to a respective relevance of at least a portion of the plurality of text blocks to information associated with at least a portion of the combined vector;

identifying a primary text block in the plurality of text blocks of the primary cluster having a highest relevance score; and

in response to the identifying of the primary text block in the plurality of text blocks of the primary cluster having the highest relevance score:

communicating the primary text block for display as a recommended selection for automatic population into at least a narrative field, of a healthcare related electronic document, configured to receive free-form narrative text.

29. The system of claim 28 , wherein generating the combined vector comprises:

identifying categorical data in one or more of the first set of historical text data or the second set of historical contextual data, wherein the categorical data is incompatible for input to the electronic machine-learning clustering model; and

transforming the categorical data into dummy variables, wherein the dummy variables are used to generate the combined vector.

30. The system of claim 28 , wherein the operations further comprise reducing data sparsity of the combined vector by:

applying Principle Component Analysis (PCA) to the combined vector to identify sparse columns of the combined vector, wherein the combined vector has n dimensions, one dimension for each of a plurality of features; and

projecting the combined vector having n dimensions to a k dimensional vector space to reduce a total quantity of the sparse columns, wherein k<n.

31. The system of claim 28 , wherein the electronic machine-learning clustering model includes a Density-Based Spatial Clustering of Applications with Noise (DBSCAN).

32. The system of claim 28 , wherein when generating the plurality of clusters, a total quantity of the plurality of clusters is determined by one or more characteristics of the first set of text data and the second set of historical data in the combined vector, wherein a plurality of cluster tags are attached to the combined vector, and wherein the plurality of cluster tags map features of the first set of text data and the second set of historical contextual data to one or more of the plurality of clusters.

33. The system of claim 28 , wherein the primary cluster is identified, subsequent to reducing a data sparsity, by:

identifying one cluster in the plurality of clusters that corresponds to a set of text blocks in one or more of physician-specific patient-encounter data or patient-specific context data, wherein the set of text blocks in the one cluster have a greatest similarity to the combined vector relative to other clusters in the plurality of clusters; and

designating the one cluster as the primary cluster based on the one cluster have the greatest similarity to the combined vector relative.

34. A method, comprising:

generating a combined vector, associated with healthcare information, based on a first set of historical text data and based further on a second set of historical contextual data, the first set of historical text data differing from the second set of historical contextual data;

generating a plurality of clusters based on an electronic machine-learning clustering model using the combined vector as input, wherein:

(a) the electronic machine-learning clustering model is trained based on data associated with instances of first information selected from a group comprising the first set of historical text data and the second set of historical contextual data, and

(b) the plurality of clusters includes the first set of historical text data and further include the second set of historical contextual data;

identifying a primary cluster in the plurality of clusters that is a best match to the combined vector;

performing natural language processing on second information associated with the primary cluster;

based at least on the natural language processing:

assigning a relevance score, for at least a portion of a plurality of text blocks in the primary cluster, corresponding to a respective relevance of at least a portion of the plurality of text blocks to information associated with at least a portion of the combined vector;

identifying a primary text block in the plurality of text blocks of the primary cluster having a highest relevance score; and

in response to the identifying of the primary text block in the plurality of text blocks of the primary cluster having the highest relevance score:

communicating the primary text block for display as a recommended selection for automatic population into at least a narrative field, of a healthcare related electronic document, configured to receive free-form narrative text.

35. The method of claim 34 , further comprising: reducing a data sparsity of the combined vector prior to generating the plurality of clusters using the combined vector as input to the electronic machine-learning clustering model.

36. The method of claim 35 , further comprising receiving a user input that includes a selection of the primary text block for automatic population of the healthcare related electronic document, communicating a value of one to a rule engine as feedback for subsequent cluster identifications.

37. The method of claim 34 , further comprising receiving user input that does not select the primary text block.

38. The method of claim 37 , wherein when the user input does not select the primary text block for automatic population of the electronic healthcare related document, communicating a value of zero to a rule engine, as feedback for subsequent cluster identifications.

39. The method of claim 34 , further comprising:

combining the first set of historical text data and the second set of historical contextual data to generate the combined vector, wherein:

(a) the first set of historical text data is stored at a first database of clinician notes, and

(b) the second set of historical contextual data is stored at a second database of patient demographic information.

40. The method of claim 34 , further comprising receiving user input that includes an edited version of the primary text block, wherein when the user input includes a selection of the edited version of the primary text block for automatic population of the healthcare related electronic document, communicating the edited version of the primary text block to a machine learning engine, wherein the edited version is stored as new historical data.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Apr 5, 2021
From: MUKHERJEE, TANMOY; MONDAL, SHRILATA; KAR, DEBENDRA; KARRA, DAMODAR REDDY
To: CERNER INNOVATION, INC.
Reel/Frame 055819/0790 →
Continuity (1)
Related Publication 20220319646A1 · Oct 6, 2022
References Cited (111)
US 6137911A · Zhilyaev · 2000 [cited by examiner]
US 7379946B2 · Carus et al. · 2008 [cited by applicant]
US 7881957B1 · Cohen et al. · 2011 [cited by applicant]
US 7885811B2 · Zimmerman et al. · 2011 [cited by applicant]
US 8170897B1 · Cohen et al. · 2012 [cited by applicant]
US 8185553B2 · Carus et al. · 2012 [cited by applicant]
US 8209183B1 · Patel et al. · 2012 [cited by applicant]
US 8255258B1 · Cohen et al. · 2012 [cited by applicant]
US 8374865B1 · Biadsy et al. · 2013 [cited by applicant]
US 8428940B2 · Kristjansson et al. · 2013 [cited by applicant]
US 8498892B1 · Cohen et al. · 2013 [cited by applicant]
US 8510340B2 · Carus et al. · 2013 [cited by applicant]
US 8612211B1 · Shires et al. · 2013 [cited by applicant]
US 8694335B2 · Yegnanarayanan · 2014 [cited by applicant]
US 8700395B2 · Zimmerman et al. · 2014 [cited by applicant]
US 8738403B2 · Flanagan et al. · 2014 [cited by applicant]
US 8756079B2 · Yegnanarayanan · 2014 [cited by applicant]
US 8768723B2 · Montyne et al. · 2014 [cited by applicant]
US 8782088B2 · Carus et al. · 2014 [cited by applicant]
US 8788289B2 · Flanagan et al. · 2014 [cited by applicant]
US 8799021B2 · Flanagan et al. · 2014 [cited by applicant]
US 8831957B2 · Taubman et al. · 2014 [cited by applicant]
US 8878773B1 · Bozarth · 2014 [cited by applicant]
US 8880406B2 · Santos-lang et al. · 2014 [cited by applicant]
US 8924211B2 · Ganong et al. · 2014 [cited by applicant]
US 8953886B2 · King et al. · 2015 [cited by applicant]
US 8972243B1 · Strom et al. · 2015 [cited by applicant]
US 8977555B2 · Torok et al. · 2015 [cited by applicant]
US 9058805B2 · Aleksic et al. · 2015 [cited by applicant]
US 9117451B2 · Fructuoso et al. · 2015 [cited by applicant]
US 9129013B2 · Delaney et al. · 2015 [cited by applicant]
US 9135571B2 · Delaney et al. · 2015 [cited by applicant]
US 9147054B1 · Beal et al. · 2015 [cited by applicant]
US 9152763B2 · Carus et al. · 2015 [cited by applicant]
US 9240187B2 · Torok et al. · 2016 [cited by applicant]
US 9257120B1 · Alvarez Guevara et al. · 2016 [cited by applicant]
US 9269012B2 · Fotland · 2016 [cited by applicant]
US 9292089B1 · Sadek · 2016 [cited by applicant]
US 9304736B1 · Whiteley et al. · 2016 [cited by applicant]
US 9318104B1 · Fructuoso et al. · 2016 [cited by applicant]
US 9324323B1 · Bikel et al. · 2016 [cited by applicant]
US 9343062B2 · Ganong et al. · 2016 [cited by applicant]
US 9378734B2 · Ganong et al. · 2016 [cited by applicant]
US 9384735B2 · White et al. · 2016 [cited by applicant]
US 9420227B1 · Shires et al. · 2016 [cited by applicant]
US 9424840B1 · Hart et al. · 2016 [cited by applicant]
US 9443509B2 · Ganong et al. · 2016 [cited by applicant]
US 9466294B1 · Tunstall-pedoe et al. · 2016 [cited by applicant]
US 9538005B1 · Nguyen et al. · 2017 [cited by applicant]
US 9542944B2 · Jablokov et al. · 2017 [cited by applicant]
US 9542947B2 · Schuster et al. · 2017 [cited by applicant]
US 9552816B2 · Vanlund et al. · 2017 [cited by applicant]
US 9563955B1 · Kamarshi et al. · 2017 [cited by applicant]
US 9569594B2 · Casella Dos Santos · 2017 [cited by applicant]
US 9570076B2 · Sierawski et al. · 2017 [cited by applicant]
US 9601115B2 · Chen et al. · 2017 [cited by applicant]
US 9678954B1 · Cuthbert et al. · 2017 [cited by applicant]
US 9679107B2 · Cardoza et al. · 2017 [cited by applicant]
US 9721570B1 · Beal et al. · 2017 [cited by applicant]
US 9786281B1 · Adams et al. · 2017 [cited by applicant]
US 9792914B2 · Alvarez Guevara et al. · 2017 [cited by applicant]
US 9805315B1 · Cohen et al. · 2017 [cited by applicant]
US 9898580B2 · Flanagan et al. · 2018 [cited by applicant]
US 9904768B2 · Yegnanarayanan · 2018 [cited by applicant]
US 9911418B2 · Chi · 2018 [cited by applicant]
US 9916420B2 · Cardoza et al. · 2018 [cited by applicant]
US 9922385B2 · Yegnanarayanan · 2018 [cited by applicant]
US 9971848B2 · D'souza et al. · 2018 [cited by applicant]
US 10032127B2 · Habboush et al. · 2018 [cited by applicant]
US 10347117B1 · Capurro · 2019 [cited by applicant]
US 10602974B1 · Govindjee et al. · 2020 [cited by applicant]
US 10951762B1 · Brandt et al. · 2021 [cited by applicant]
US 11024424B2 · Sun · 2021 [cited by examiner]
US 20080177537A1 · Ash et al. · 2008 [cited by applicant]
US 20090100358A1 · Lauridsen et al. · 2009 [cited by applicant]
US 20090138288A1 · Benja-athon · 2009 [cited by applicant]
US 20090268882A1 · Lee et al. · 2009 [cited by applicant]
US 20100161353A1 · Mayaud · 2010 [cited by applicant]
US 20110178931A1 · Kia · 2011 [cited by applicant]
US 20120066197A1 · Rana · 2012 [cited by examiner]
US 20140350961A1 · Csurka et al. · 2014 [cited by applicant]
US 20150006199A1 · Snider et al. · 2015 [cited by applicant]
US 20150142418A1 · Byron et al. · 2015 [cited by applicant]
US 20150169827A1 · Laborde · 2015 [cited by applicant]
US 20150340033A1 · Di Fabbrizio et al. · 2015 [cited by applicant]
US 20160119305A1 · Panchura et al. · 2016 [cited by applicant]
US 20170116373A1 · Ginsburg et al. · 2017 [cited by applicant]
US 20180012604A1 · Guevara et al. · 2018 [cited by applicant]
US 20180113676A1 · De Sousa Webber · 2018 [cited by examiner]
US 20180166076A1 · Higuchi et al. · 2018 [cited by applicant]
US 20180308490A1 · Lim et al. · 2018 [cited by applicant]
US 20180322110A1 · Rhodes et al. · 2018 [cited by applicant]
US 20180373844A1 · Ferrandez-escamez et al. · 2018 [cited by applicant]
US 20190121532A1 · Strader et al. · 2019 [cited by applicant]
US 20190130073A1 · Sun et al. · 2019 [cited by applicant]
US 20190206524A1 · Baldwin et al. · 2019 [cited by applicant]
US 20190272919A1 · Frandsen et al. · 2019 [cited by applicant]
US 20190287665A1 · Forsberg et al. · 2019 [cited by applicant]
US 20190311807A1 · Kannan et al. · 2019 [cited by applicant]
US 20190362846A1 · Vodencarevic · 2019 [cited by examiner]
US 20200034366A1 · Kivatinos · 2020 [cited by examiner]
US 20220335942A1 · Agassi et al. · 2022 [cited by applicant]
US 20220344049A1 · Hall · 2022 [cited by examiner]
WO WO2017205850A1 · 2017 [cited by examiner]
Final Office Action received for U.S. Appl. No. 16/720,641, mailed on Jan. 11, 2022, 18 pages. [cited by applicant]
Pre-Interview First Office action received for U.S. Appl. No. 16/720,632, mailed on Feb. 17, 2022, 4 pages. [cited by applicant]
Pre interview First Office Action received for U.S. Appl. No. 16/720,644, mailed on Aug. 4, 2022, 4 pages. [cited by applicant]
Chiu et al., “Speech Recognition for Medical Conversations”, Available Online at: <arXiv preprint arXiv:1711.07274>, Nov. 20, 2017, 5 pages. [cited by applicant]
Final Office Action received for U.S. Appl. No. 16/720,632, mailed on Oct. 19, 2022, 31 pages. [cited by applicant]
Non-Final Office Action received for U.S. Appl. No. 17/132,859, mailed on Nov. 22, 2022, 30 pages. [cited by applicant]
Preinterview First Office Action of U.S. Appl. No. 16/720,641, mailed on Oct. 14, 2021, 7 pages. [cited by applicant]
Cited By (1)
US 12,718,813