IP Library › Granted Patent US 12,229,512
Granted Patent B2
US 12,229,512 · App. 17/461,649 · Granted Feb 18, 2025

Significance-based prediction from unstructured text

Inventors: Ayan Sengupta (Noida, IN); Saransh Chauksi (Noida, IN); Zhijing J. Liu (Audubon, PA)
Assignee: Optum, Inc.
G06F40/284
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,229,512
App. No.
17/461,649
Granted
Feb 18, 2025
Kind
B2
Abstract

Various embodiments provide methods, apparatus, systems, computing entities, and/or the like, generating predictions based at least in part on recognizing significant words in unstructured text. In an embodiment, a method is provided. The method comprises: generating a plurality of word-level tokens for an input unstructured textual data object; and for each word-level token: determining a significance type and a significance subtype for the word-level token by using a significance recognition machine learning model, and assigning a significance token label or an insignificance token label to the word-level token. The method further comprises: generating a label-based feature data object based at least in part on a subset of word-level tokens associated with the significance token label; generating a prediction data object for the input unstructured textual data object by providing the label-based feature data object to a first prediction machine learning model; and performing one or more automated prediction-based actions.

Claims (81)

1. A computer-implemented method comprising:

receiving, by one or more processors and originating from a client computing entity, an application programming interface (API) request that identifies an input unstructured textual data object;

generating, by the one or more processors, a plurality of word-level tokens for the input unstructured textual data object;

determining, by the one or more processors using a significance recognition machine learning model, a significance type for a word-level token of the plurality of word-level tokens by:

(i) determining a word-level embedding for the word-level token,

(ii) determining a character-level embedding for the word-level token,

(iii) determining a fusion representation data object for the word-level token based at least in part on a combination of the word-level embedding and the character-level embedding, and

(iv) determining the significance type for the word-level token based at least in part on the fusion representation data object for the word-level token;

assigning, by the one or more processors, a significance token label to the word-level token based at least in part on the significance type;

generating, by the one or more processors, a label-based feature data object for the input unstructured textual data object based at least in part on a subset of word-level tokens from the plurality of word-level tokens that is associated with the significance token label;

generating, by the one or more processors and a prediction machine learning model, a prediction data object for the input unstructured textual data object based at least in part on the label-based feature data object;

providing, by the one or more processors and to the client computing entity, an API response that identifies the prediction data object; and

initiating, by the one or more processors, the performance of one or more automated prediction-based actions based at least in part on the prediction data object.

2. The computer-implemented method of claim 1 , wherein the significance recognition machine learning model determines the significance type of the word-level token by:

generating the word-level embedding for the word-level token based at least in part on a comparison between the word-level token and the plurality of word-level tokens;

generating the character-level embedding for the word-level token based at least in part on one or more characters within the word-level token; and

applying one or more attention mechanisms to the fusion representation data object to determine the significance type of the word-level token.

3. The computer-implemented method of claim 2 , wherein the significance recognition machine learning model comprises a first bidirectional long short term memory mechanism configured to generate the word-level embedding for the word-level token and a second bidirectional long short term memory mechanism configured to generate the fusion representation data object for the word-level token.

4. The computer-implemented method of claim 1 , wherein the significance recognition machine learning model is trained using a plurality of historical word-level tokens in a historical unstructured textual data object of a plurality of unstructured textual data objects by:

generating a ground-truth token label for a historical word-level token of the plurality of historical word-level tokens;

generating an inferred token label using the significance recognition machine learning model; and

updating the significance recognition machine learning model based at least in part on a comparison between the ground-truth token label and the inferred token label.

5. The computer-implemented method of claim 4 , wherein the ground-truth token label for the historical word-level token is based at least in part on a similarity value between the historical word-level token and each of a plurality of reference tokens respectively corresponding to a plurality of significance subtypes.

6. The computer-implemented method of claim 1 , wherein:

the prediction machine learning model comprises a plurality of graph-based prediction mechanisms, and

generating the prediction data object for the input unstructured textual data object comprises:

determining that the label-based feature data object is not one of a first plurality of different label-based feature data objects associated with a first graph-based prediction mechanism of the plurality of graph-based prediction mechanisms;

selecting a second graph-based prediction mechanism associated with a second plurality of different label-based feature data objects; and

generating the prediction data object for the input unstructured textual data object based at least in part on the second graph-based prediction mechanism.

7. The computer-implemented method of claim 6 , wherein each of the plurality of graph-based prediction mechanisms comprises a traversable k-nearest neighbor graph using Euclidean distance between each of a plurality of different label-based feature data objects, and wherein k is one.

8. The computer-implemented method of claim 6 , wherein the first plurality of different label-based feature data objects associated with the first graph-based prediction mechanism is obtained from a first external dataset, and the second plurality of different label-based feature data objects associated with the second graph-based prediction mechanism is obtained from a second external dataset.

9. The computer-implemented method of claim 1 , wherein the significance token label indicates that the word-level token describes demographic information for a healthcare provider rendering healthcare services in a medical encounter described by the input unstructured textual data object.

10. The computer-implemented method of claim 1 , wherein the prediction data object comprises an identifier for a healthcare provider rendering healthcare services in a medical encounter described by the input unstructured textual data object.

11. A system comprising memory and one or more processors communicatively coupled to the memory, the one or more processors configured to:

receive, originating from a client computing entity, an application programming interface (API) request that identifies an input unstructured textual data object;

generate a plurality of word-level tokens for the input unstructured textual data object;

determine, using a significance recognition machine learning model, a significance type for a word-level token of the plurality of word-level tokens by:

(i) determining a word-level embedding for the word-level token,

(ii) determining a character-level embedding for the word-level token,

(iii) determining a fusion representation data object for the word-level token based at least in part on a combination of the word-level embedding and the character-level embedding, and

(iv) determining the significance type for the word-level token based at least in part on the fusion representation data object for the word-level token;

assign a significance token label to the word-level token based at least in part on the significance type;

generate a label-based feature data object for the input unstructured textual data object based at least in part on a subset of word-level tokens from the plurality of word-level tokens that is associated with the significance token label;

generate, using a prediction machine learning model, a prediction data object for the input unstructured textual data object based at least in part on the label-based feature data object;

provide, to the client computing entity, an API response that identifies the prediction data object; and

initiate the performance of one or more automated prediction-based actions based at least in part on the prediction data object.

12. The system of claim 11 , wherein the significance recognition machine learning model determines the significance type of the word-level token by:

generating the word-level embedding for the word-level token based at least in part on a comparison between the word-level token and the plurality of word-level tokens;

generating the character-level embedding for the word-level token based at least in part on one or more characters within the word-level token; and

applying one or more attention mechanisms to the fusion representation data object to determine the significance type of the word-level token.

13. The system of claim 12 , wherein the significance recognition machine learning model comprises a first bidirectional long short term memory mechanism configured to generate the word-level embedding for the word-level token and a second bidirectional long short term memory mechanism configured to generate the fusion representation data object for the word-level token.

14. The system of claim 11 , wherein the significance recognition machine learning model is trained using a plurality of historical word-level tokens in a historical unstructured textual data object of a plurality of unstructured textual data objects by:

generating a ground-truth token label for a historical word-level token of the plurality of historical word-level tokens;

generating an inferred token label using the significance recognition machine learning model; and

updating the significance recognition machine learning model based at least in part on a comparison between the ground-truth token label and the inferred token label.

15. The system of claim 14 , wherein the ground-truth token label for the historical word-level token is based at least in part on a similarity value between the historical word-level token and each of a plurality of reference tokens respectively corresponding to a plurality of significance subtypes.

16. The system of claim 11 , wherein:

the prediction machine learning model comprises a plurality of graph-based prediction mechanisms, and

to generate the prediction data object for the input unstructured textual data object, the one or more processors are configured to:

determine that the label-based feature data object is not one of a first plurality of different label-based feature data objects associated with a first graph-based prediction mechanism of the plurality of graph-based prediction mechanisms;

select a second graph-based prediction mechanism associated with a second plurality of different label-based feature data objects; and

generate the prediction data object for the input unstructured textual data object based at least in part on the second graph-based prediction mechanism.

17. The system of claim 16 , wherein each of the plurality of graph-based prediction mechanisms comprises a traversable k-nearest neighbor graph using Euclidean distance between each of a plurality of different label-based feature data objects, and wherein k is one.

18. The system of claim 16 , wherein the first plurality of different label-based feature data objects associated with the first graph-based prediction mechanism is obtained from a first external dataset, and the second plurality of different label-based feature data objects associated with the second graph-based prediction mechanism is obtained from a second external dataset.

19. One or more non-transitory computer-readable storage media including instructions that, when executed by one or more processors, cause the one or more processors to:

receive, originating from a client computing entity, an application programming interface (API) request that identifies an input unstructured textual data object;

generate a plurality of word-level tokens for the input unstructured textual data object;

determine, using a significance recognition machine learning model, a significance type for a word-level token of the plurality of word-level tokens by:

(i) determining a word-level embedding for the word-level token,

(ii) determining a character-level embedding for the word-level token,

(iii) determining a fusion representation data object for the word-level token based at least in part on a combination of the word-level embedding and the character-level embedding, and

(iv) determining the significance type for the word-level token based at least in part on the fusion representation data object for the word-level token;

assign a significance token label to the word-level token based at least in part on the significance type;

generate a label-based feature data object for the input unstructured textual data object based at least in part on a subset of word-level tokens from the plurality of word-level tokens that is associated with the significance token label;

generate, using a prediction machine learning model, a prediction data object for the input unstructured textual data object based at least in part on the label-based feature data object;

provide, to the client computing entity, an API response that identifies the prediction data object; and

initiate the performance of one or more automated prediction-based actions based at least in part on the prediction data object.

20. The one or more non-transitory computer-readable storage media of claim 19 , wherein the significance recognition machine learning model determines the significance type of the word-level token by:

generating the word-level embedding for the word-level token based at least in part on a comparison between the word-level token and the plurality of word-level tokens;

generating the character-level embedding for the word-level token based at least in part on one or more characters within the word-level token; and

applying one or more attention mechanisms to the fusion representation data object to determine the significance type of the word-level token.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Sep 3, 2021
From: SENGUPTA, AYAN; CHAUKSI, SARANSH; LIU, ZHIJING J.
To: OPTUM, INC.
Reel/Frame 057386/0571 →
Continuity (1)
Related Publication 20230061731A1 · Mar 2, 2023
References Cited (47)
US 5664109A · Johnson et al. · 1997 [cited by applicant]
US 7505621B1 · Agrawal et al. · 2009 [cited by applicant]
US 8538780B1 · Sweeney et al. · 2013 [cited by applicant]
US 9961070B2 · Tang · 2018 [cited by applicant]
US 10861589B2 · Waits · 2020 [cited by applicant]
US 20050234740A1 · Krishnan et al. · 2005 [cited by applicant]
US 20110093293A1 · G. N. et al. · 2011 [cited by applicant]
US 20130191137A1 · Chen et al. · 2013 [cited by applicant]
US 20130197925A1 · Blue · 2013 [cited by applicant]
US 20130339060A1 · Delaney et al. · 2013 [cited by applicant]
US 20150193580A1 · Mosier, III et al. · 2015 [cited by applicant]
US 20150205846A1 · Aldridge et al. · 2015 [cited by applicant]
US 20150350148A1 · Kenney et al. · 2015 [cited by applicant]
US 20190206524A1 · Baldwin · 2019 [cited by examiner]
US 20200167871A1 · Basu et al. · 2020 [cited by applicant]
US 20230059494A1 · Hunter · 2023 [cited by examiner]
“Amazon Comprehend Medical,” Amazon Web Services (AWS), (7 pages), (online), [Retrieved from the Internet Nove. 27, 2021] <URL: https://aws.amazon.com/comprehend/medical/>. [cited by applicant]
“CLAMP—Clinical Language Annotation, Modeling, and Processing Toolkit,” (8 pages), (online), [Retrieved from the Internet Nov. 27, 2021] <URL: https://clamp.uth.edu/>. [cited by applicant]
“Industrial-Strength Natural Language Processing in Python,” spaCy, (9 pages), (online), [Retrieved from the Internet Nov. 27, 2021] <URL: https://spacy.io/>. [cited by applicant]
“SpaCy Models For biomedical Text Processing,” GitHub, (5 pages), (online), [Retrieved from the Internet Nov. 27, 2021] <URL: https://allenai.github.io/scispacy/>. [cited by applicant]
“Spark NLP—State of the Art Natural Language Processing,” John Snow Labs, (11 pages), (online), [Retrieved from the Internet Nov. 27, 2021] <URL: https://nlp.johnsnowlabs.com/>. [cited by applicant]
Alsentzer, Emily et al. “Publicly Available Clinical BERT Embeddings,” arXiv Preprint arXiv:1904.03323v3 [cs.CL] Jun. 20, 2019, (7 pages). [cited by applicant]
Arumae, Kristjan et al. “CALM: Continuous Adaptive Learning for Language Modeling,” arXiv Preprint arXiv: 2004.03794v1 [cs.CL] Apr. 8, 2020, (7 pages). [cited by applicant]
Bahdanau, Dzmitry et al. “Neural Machine Translation By Jointly Learning to Align and Translate,” arXiv Preprint arXiv: 1409.0473v1 [cs.CL] Sep. 1, 2014, pp. 1-15. [cited by applicant]
Bhatia, Parminder et al. “Joint Entity Extraction and Assertion Detection for Clinical Text,” arXiv Preprint, arXiv:1812.05270v5 [cs.CL] Jan. 22, 2020, (6 pages). [cited by applicant]
Bindman, Andrew B. “Using the National Provider Identifier for Health Care Workforce Evaluation,” Medicare & Medicaid Research Review, vol. 3, No. 3, (2013), pp. E1-E10, DOI: http://dx.doi.org/10.5600/mmrr.003.03.b03, I… [cited by applicant]
Bojanowski, Piotr et al. “Enriching Word Vectors with Subword Information,” Transactions of the Association for Computational Linguistics, vol. 5, Jun. 2017, pp. 135-146. [cited by applicant]
Cao, Shilei et al. “Knowledge Guided Short-Text Classification for Healthcare Applications,” in 2017 IEEE International Conference on Data Mining (ICDM), Nov. 18, 2017, pp. 31-40, IEEE. [cited by applicant]
Chiticariu, Laura et al. “Rule-Based Information Extraction Is Dead! Long Live Rule-Based Information Extraction Systems!,” Proceedings of the 2013 Conference on Empirical Methods In Natural Language Processing, Oct. 18… [cited by applicant]
Chiu, Jason P.C. et al. “Named Entity Recognition With Bidirectional LSTM-CNNs,” Transactions of the Association for Computational Linguistics, vol. 4, Jul. 2016, pp. 357-370. [cited by applicant]
Devlin, Jacob et al. “BERT: Pre-Training of Deep Bidirectional Transformers for Language Understanding,” arXiv Preprint arXiv: 1810.04805v1 [cs.CL] Oct. 11, 2018, (14 pages). [cited by applicant]
Hughes, Mark et al. “Medical Text Classification Using Convolutional Neural Networks,” Studies in Health Technology and Informatics, vol. 235, pp. 246-250, DOI: 10.3233/978-1-61499-753-5-246. [cited by applicant]
Jagannatha, Abhyuday N. et al. “Bidirectional RNN for Medical Event Detection in Electronic Health Records,” in Proceedings of the Conference. Association for Computational Linguistics. North American Chapter. Meeting, … [cited by applicant]
Lee, Jinhyuk et al. BioBERT: a Pre-Trained Biomedical Language Representation Model for Biomedical Text Mining, Bioinformatics, vol. 36, No. 4, Sep. 10, 2019, pp. 1234-1240, DOI: 10.1093/bioinformatics/btz682. [cited by applicant]
Li, Pengfei et al. “Improving Relation Extraction With Knowledge-Attention,” Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural L… [cited by applicant]
Luong, Minh-Thang et al. “Effective Approaches to Attention-Based Neural Machine Translation,” arXiv Preprint arXiv: 1508.04025v5 [cs.CL] Sep. 20, 2015, (11 pages). [cited by applicant]
Magge, Arjun et al. “Clinical NER and Relation Extraction Using Bi-Char-LSTMS and Random Forest Classifiers,” in International Workshop on Medication and Adverse Drug Event Detection (pp. 25-30), vol. 20, May 16, 2018, … [cited by applicant]
Mikolov, Tomas et al. “Distributed Representations of Words and Phrases and Their Compositionality,” Advances in Neural Information Processing Systems, vol. 26, Oct. 2013, (9 pages). [cited by applicant]
Mykowiecka, Agnieszka et al. “Rule-Based Information Extraction From Patients' Clinical Data,” Journal of Biomedical Informatics, vol. 42, Jul. 29, 2009, pp. 923-936, DOI: 10.1016/j.jbi.2009.07.007. [cited by applicant]
Singh, Guarav et al. “Relation Extraction Using Explicit Context Conditioning,” in Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Techno… [cited by applicant]
Vaswani, Ashish et al. “Attention Is All You Need,” 31st Conference on Advances in Neural Information Processing Systems (NIPS 2017), Dec. 4, 2017, (11 pages). [cited by applicant]
Vu, Thanh et al. “A Label Attention Model for ICD Coding from Clinical Text,” Proceedings of the Twenty-Ninth International Joint Conference on Artificial Intelligence (IJCAI-20), Jul. 9, 2020, pp. 3335-3341, DOI: 10.24… [cited by applicant]
Whang, Taesun et al. “Domain Adaptive Training BERT for Response Selection,” arXiv Preprint arXiv: 1908.04812v1 [cs.CL] Aug. 13, 2019, (7 pages). [cited by applicant]
Wu, Yonghui et al. “Clinical Named Entity Recognition Using Deep Learning Models,” American Medical Informatics Association Annual Symposium Proceedings, vol. 2017, Apr. 16, 2018, pp. 1812-1819. [cited by applicant]
Xu, Guohai et al. “Improving Clinical Named Entity Recognition With Global Neural Attention,” in Asia-Pacific Web (APWeb) and Web-Age Information Management (WAIM) Joint International Conference on Web and Big Data, Jul… [cited by applicant]
Zhang, Yijia et al. “BioWordVec, Improving Biomedical Word Embeddings With Subword Information and MeSH,” Scientific Data, vol. 6, No. 52, May 10, 2019, pp. 1-9, DOI: 10.1038/s41597-019-0055-0. [cited by applicant]
Zhu, Ming et al. “LATTE: Latent Type Modeling for Biomedical Entity Linking,” in Proceedings of the Thirty-Fourth AAAI Conference on Advancement of Artificial Intelligence (AAAI-20), vol. 34, No. 05, Apr. 3, 2020, pp. 9… [cited by applicant]