IP Library Granted Patent US 12,367,341
Granted Patent B2
US 12,367,341 · App. 17/808,199 · Granted Jul 22, 2025

Natural language processing machine learning frameworks trained using multi-task training routines

Inventors: Suman Roy (Bangalore, IN); Ayan Sengupta (Noida, IN); Michael Bridges (Dublin, IE); Amit Kumar (Gaya, IN)
Assignee: Optum Services (Ireland) Limited
G06F40/279G06F40/40G06N20/00G06F40/30
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,367,341
App. No.
17/808,199
Granted
Jul 22, 2025
Kind
B2
Abstract

Various embodiments of the present invention provide methods, apparatus, systems, computing devices, computing entities, and/or the like for performing natural language processing operations using an attention-based text encoder machine learning model that is trained using a multi-task training routine that is associated with two or more training tasks (e.g., a multi-task training routine that is associated with two or more sequential training tasks, a multi-training routine that is associated with two or more concurrent training tasks, and/or the like).

Claims (56)

1. A computer-implemented method comprising:

generating, by one or more processors and using an attention-based text encoder machine learning model and based at least in part on an unlabeled document data object, an unlabeled document word-wise embedded representation set that comprises a plurality of unlabeled document word-wise embedded representations, each of the plurality of unlabeled document word-wise embedded representations associated with a respective unlabeled document word of the unlabeled document data object, wherein the attention-based text encoder machine learning model comprises a tunable parameter value that is generated based at least in part on a concurrent learning loss by:

identifying one or more training input document data objects comprising one or more training unlabeled document data objects and a plurality of labeled document data objects,

identifying a training document pair, wherein the training document pair comprises (i) a training unlabeled document data object of the one or more training unlabeled document data objects and (ii) a labeled document data object of the plurality of labeled document data objects that is related to the training unlabeled document data object based at least in part on ground-truth cross-document relationships,

generating a cross-document distance measure for the training document pair,

generating the concurrent learning loss based at least in part on: (i) a language modeling loss model that is based at least in part on the one or more training input document data objects, and (ii) a similarity determination loss model that is determined based at least in part on the cross-document distance measure, and

generating, using the concurrent learning loss, the tunable parameter value;

generating, by the one or more processors and using the attention-based text encoder machine learning model and based at least in part on a labeled document data object of the plurality of the labeled document data objects, a labeled document word-wise embedded representation set that comprises a plurality of labeled document word-wise embedded representations, each of the plurality of labeled document word-wise embedded representations associated with a respective labeled document word of the labeled document data object;

generating, by the one or more processors and based at least in part on the plurality of unlabeled document word-wise embedded representations and the plurality of labeled document word-wise embedded representations, a cross-document similarity measure comparing the unlabeled document data object and the labeled document data object; and

generating, by the one or more processors and for the unlabeled document data object, a document classification based at least in part on the cross-document similarity measure.

2. The computer-implemented method of claim 1 , wherein the language modeling loss model is defined in accordance with a language modeling training task and the language modeling training task comprises a text reconstruction sub-task.

3. The computer-implemented method of claim 1 , wherein the attention-based text encoder machine learning model comprises an attentional autoencoder machine learning model.

4. The computer-implemented method of claim 1 , wherein the cross-document distance measure comprises one or more cosine distance measures.

5. The computer-implemented method of claim 1 , wherein generating the cross-document similarity measure comprises:

for a given word pair of a plurality of word pairs that comprises a given unlabeled document word of the unlabeled document data object and a given labeled document word of the labeled document data object, generating a pairwise word similarity measure,

for the given word pair, generating a pairwise flow indicator based at least in part on the pairwise word similarity measure of the given word pair relative to other pairwise word similarity measures in a subset of the plurality of word pairs that is associated with the given unlabeled document word in the given word pair; and

generating the cross-document similarity measure based at least in part on the pairwise word similarity measure and the pairwise flow indicator.

6. The computer-implemented method of claim 5 , wherein the pairwise word similarity measure for the given word pair comprises a cosine similarity measure between a given unlabeled document word-wise embedded representation for the given unlabeled document word and a given labeled document word-wise embedded representation for the given labeled document word.

7. The computer-implemented method of claim 5 further comprising determining the pairwise word similarity measure in a manner that is configured to maximize a word mover's similarity measure for the unlabeled document data object and the labeled document data object.

8. A system comprising one or more processors and at least one memory storing processor-executable instructions that, when executed by the one or more processors, cause the one or more processors to perform operations comprising:

generating, using an attention-based text encoder machine learning model and based at least in part on an unlabeled document data object, an unlabeled document word-wise embedded representation set that comprises a plurality of unlabeled document word-wise embedded representations, each of the plurality of unlabeled document word-wise embedded representations associated with a respective unlabeled document word of the unlabeled document data object, wherein the attention-based text encoder machine learning model comprises a tunable parameter value that is generated based at least in part on a concurrent learning loss by:

identifying one or more training input document data objects comprising one or more training unlabeled document data objects and a plurality of labeled document data objects,

identifying a training document pair, wherein the training document pair comprises (i) a training unlabeled document data object of the one or more training unlabeled document data objects and (ii) a labeled document data object of the plurality of labeled document data objects that is related to the training unlabeled document data object based at least in part on ground-truth cross-document relationships,

generating a cross-document distance measure for the training document pair,

generating the concurrent learning loss based at least in part on: (i) a language modeling loss model that is based at least in part on the one or more training input document data objects, and (ii) a similarity determination loss model that is determined based at least in part on the cross-document distance measure, and

generating, using the concurrent learning loss, the tunable parameter value;

generating, using the attention-based text encoder machine learning model and based at least in part on a labeled document data object of the plurality of labeled document data objects, a labeled document word-wise embedded representation set that comprises a plurality of labeled document word-wise embedded representations, each of the plurality of labeled document word-wise embedded representations associated with a respective labeled document word of the labeled document data object;

generating, based at least in part on the plurality of unlabeled document word-wise embedded representations and the plurality of labeled document word-wise embedded representations, a cross-document similarity measure comparing the unlabeled document data object and the labeled document data object; and

generating, for the unlabeled document data object, a document classification based at least in part on the cross-document similarity measure.

9. The system of claim 8 , wherein the language modeling loss model is defined in accordance with a language modeling training task and the language modeling training task comprises a text reconstruction sub-task.

10. The system of claim 8 , wherein the attention-based text encoder machine learning model comprises an attentional autoencoder machine learning model.

11. The system of claim 8 , wherein the cross-document distance measure comprises one or more cosine distance measures.

12. The system of claim 8 , wherein generating the cross-document similarity measure, comprises:

for a given word pair of a plurality of word pairs that comprises a given unlabeled document word of the unlabeled document data object and a given labeled document word of the labeled document data object, generating a pairwise word similarity measure,

for the given word pair, generating a pairwise flow indicator based at least in part on the pairwise word similarity measure of the given word pair relative to other pairwise word similarity measures in a subset of the plurality of word pairs that is associated with the given unlabeled document word in the given word pair; and

generating the cross-document similarity measure based at least in part on the pairwise word similarity measure and the pairwise flow indicator.

13. The system of claim 12 , wherein the pairwise word similarity measure for the given word pair comprises a cosine similarity measure between a given unlabeled document word-wise embedded representation for the given unlabeled document word and a given labeled document word-wise embedded representation for the given labeled document word.

14. The system of claim 12 , wherein the operations further comprise determining the pairwise word similarity measure in a manner that is configured to maximize a word mover's similarity measure for the unlabeled document data object and the labeled document data object.

15. One or more non-transitory computer-readable storage media including instructions that, when executed by one or more processors, cause the one or more processors to perform operations comprising:

generating, using an attention-based text encoder machine learning model and based at least in part on an unlabeled document data object, an unlabeled document word-wise embedded representation set that comprises a plurality of unlabeled document word-wise embedded representations, each of the plurality of unlabeled document word-wise embedded representations associated with a respective unlabeled document word of the unlabeled document data object, wherein the attention-based text encoder machine learning model comprises a tunable parameter value that is generated based at least in part on a concurrent learning loss by:

identifying one or more training input document data objects comprising one or more training unlabeled document data objects and a plurality of labeled document data objects,

identifying a training document pair, wherein the training document pair comprises (i) a training unlabeled document data object of the one or more training unlabeled document data objects and (ii) a labeled document data object of the plurality of labeled document data objects that is related to the training unlabeled document data object based at least in part on ground-truth cross-document relationships,

generating a cross-document distance measure for the training document pair,

generating the concurrent learning loss based at least in part on: (i) a language modeling loss model that is based at least in part on the one or more training input document data objects, and (ii) a similarity determination loss model that is determined based at least in part on the cross-document distance measure, and

generating, using the concurrent learning loss, the tunable parameter values;

generating, using the attention-based text encoder machine learning model and based at least in part on a labeled document data object of the plurality of labeled document data objects, a labeled document word-wise embedded representation set that comprises a plurality of labeled document word-wise embedded representations, each of the plurality of labeled document word-wise embedded representations associated with a respective labeled document word of the labeled document data object;

generating, based at least in part on the plurality of unlabeled document word-wise embedded representations and the plurality of labeled document word-wise embedded representations, a cross-document similarity measure comparing the unlabeled document data object and the labeled document data object; and

generating, for the unlabeled document data object, a document classification based at least in part on the cross-document similarity measure.

16. The one or more non-transitory computer-readable storage media of claim 15 , wherein the language modeling loss model is defined in accordance with a language modeling training task and the language modeling training task comprises a text reconstruction sub-task.

17. The one or more non-transitory computer-readable storage media of claim 15 , wherein the attention-based text encoder machine learning model comprises an attentional autoencoder machine learning model.

18. The one or more non-transitory computer-readable storage media of claim 15 , wherein the cross-document distance measure comprises one or more cosine distance measures.

19. The one or more non-transitory computer-readable storage media of claim 15 , wherein generating the cross-document similarity measure comprises:

for a given word pair of a plurality of word pairs that comprises a given unlabeled document word of the unlabeled document data object and a given labeled document word of the labeled document data object, generating a pairwise word similarity measure,

for the given word pair, generating a pairwise flow indicator based at least in part on the pairwise word similarity measure of the given word pair relative to other pairwise word similarity measures in a subset of the plurality of word pairs that is associated with the given unlabeled document word in the given word pair; and

generating the cross-document similarity measure based at least in part on the pairwise word similarity measure and the pairwise flow indicator.

20. The one or more non-transitory computer-readable storage media of claim 19 , wherein the pairwise word similarity measure for the given word pair comprises a cosine similarity measure between a given unlabeled document word-wise embedded representation for the given unlabeled document word and a given labeled document word-wise embedded representation for the given labeled document word.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jun 22, 2022
From: ROY, SUMAN; SENGUPTA, AYAN; BRIDGES, MICHAEL; KUMAR, AMIT
To: OPTUM SERVICES (IRELAND) LIMITED
Reel/Frame 060276/0703 →
Continuity (1)
Related Publication 20230419034A1 · Dec 28, 2023
References Cited (132)
US 6301577B1 · Matsumoto et al. · 2001 [cited by applicant]
US 7203679B2 · Agrawal et al. · 2007 [cited by applicant]
US 7610192B1 · Jamieson · 2009 [cited by applicant]
US 7844595B2 · Canright et al. · 2010 [cited by applicant]
US 8402030B1 · Pyle et al. · 2013 [cited by applicant]
US 8719197B2 · Schmidtler et al. · 2014 [cited by applicant]
US 9087236B2 · Dhoolia et al. · 2015 [cited by applicant]
US 9171057B2 · Botros · 2015 [cited by applicant]
US 10140288B2 · Pestian et al. · 2018 [cited by applicant]
US 10152648B2 · Filimonova · 2018 [cited by applicant]
US 10296846B2 · Csurka et al. · 2019 [cited by applicant]
US 10667794B2 · Beymer et al. · 2020 [cited by applicant]
US 10679738B2 · Ganesan et al. · 2020 [cited by applicant]
US 10754925B2 · D'Souza et al. · 2020 [cited by applicant]
US 10769381B2 · Tacchi et al. · 2020 [cited by applicant]
US 10824661B1 · Huang et al. · 2020 [cited by applicant]
US 11263523B1 · Duchon et al. · 2022 [cited by applicant]
US 11379665B1 · Edmund et al. · 2022 [cited by applicant]
US 11481689B2 · Dong et al. · 2022 [cited by applicant]
US 11734937B1 · Pushkin et al. · 2023 [cited by applicant]
US 20050216516A1 · Calistri-Yeh · 2005 [cited by examiner]
US 20080162455A1 · Daga et al. · 2008 [cited by applicant]
US 20120185275A1 · Loghmani · 2012 [cited by applicant]
US 20130031088A1 · Srikrishna et al. · 2013 [cited by applicant]
US 20130138641A1 · Korolev et al. · 2013 [cited by applicant]
US 20130254153A1 · Marcheret · 2013 [cited by applicant]
US 20150066974A1 · Winn · 2015 [cited by applicant]
US 20150227505A1 · Morimoto · 2015 [cited by applicant]
US 20150286629A1 · Abdel-Reheem et al. · 2015 [cited by applicant]
US 20160117589A1 · Scholtes · 2016 [cited by applicant]
US 20160300020A1 · Wetta et al. · 2016 [cited by applicant]
US 20170286869A1 · Zarosim et al. · 2017 [cited by applicant]
US 20170337334A1 · Stanczak et al. · 2017 [cited by applicant]
US 20180165554A1 · Zhang et al. · 2018 [cited by applicant]
US 20180349388A1 · Skiles et al. · 2018 [cited by applicant]
US 20190065986A1 · Witbrock et al. · 2019 [cited by applicant]
US 20190108175A1 · Sevenster et al. · 2019 [cited by applicant]
US 20200111019A1 · Goodsitt et al. · 2020 [cited by applicant]
US 20200125639A1 · Doyle · 2020 [cited by applicant]
US 20200134506A1 · Wang et al. · 2020 [cited by applicant]
US 20200202181A1 · Yadav et al. · 2020 [cited by applicant]
US 20200312431A1 · Zhang et al. · 2020 [cited by applicant]
US 20200327404A1 · Miotto et al. · 2020 [cited by applicant]
US 20200334416A1 · Vianu et al. · 2020 [cited by applicant]
US 20200356627A1 · Pablo · 2020 [cited by examiner]
US 20200364404A1 · Priestas et al. · 2020 [cited by applicant]
US 20200387668A1 · Yokote · 2020 [cited by examiner]
US 20200410157A1 · Van et al. · 2020 [cited by applicant]
US 20210034813A1 · Wu et al. · 2021 [cited by applicant]
US 20210065683A1 · Meng et al. · 2021 [cited by applicant]
US 20210149937A1 · Coulombe et al. · 2021 [cited by applicant]
US 20210157979A1 · Sheide et al. · 2021 [cited by applicant]
US 20210182479A1 · Kim et al. · 2021 [cited by applicant]
US 20210335469A1 · Xie et al. · 2021 [cited by applicant]
US 20210343410A1 · Zhang et al. · 2021 [cited by applicant]
US 20210358601A1 · Pillai et al. · 2021 [cited by applicant]
US 20210365676A1 · Valouch · 2021 [cited by examiner]
US 20220019741A1 · Roy et al. · 2022 [cited by applicant]
US 20220121823A1 · Lockett et al. · 2022 [cited by applicant]
US 20220139384A1 · Wu et al. · 2022 [cited by applicant]
US 20220207536A1 · Tian et al. · 2022 [cited by applicant]
US 20220318504A1 · Malkiel et al. · 2022 [cited by applicant]
US 20220368696A1 · Karpovsky et al. · 2022 [cited by applicant]
US 20220414330A1 · Roy et al. · 2022 [cited by applicant]
US 20230034401A1 · Weston · 2023 [cited by examiner]
US 20230103382A1 · Lu · 2023 [cited by examiner]
US 20230108863A1 · Gunasekara et al. · 2023 [cited by applicant]
US 20230119402A1 · Kumar et al. · 2023 [cited by applicant]
US 20230333518A1 · Oi et al. · 2023 [cited by applicant]
US 20230419035A1 · Roy et al. · 2023 [cited by applicant]
CN 109635109A · 2019 [cited by applicant]
CN 108733837B · 2021 [cited by applicant]
EP 3392780A2 · 2018 [cited by applicant]
WO 2021252419A1 · 2021 [cited by applicant]
WO 2022081812A1 · 2022 [cited by applicant]
Advisory Action for U.S. Appl. No. 16/930,862, dated Oct. 25, 2023, (3 pages), United States Patent and Trademark Office, US. [cited by applicant]
NonFinal Office Action for U.S. Appl. No. 16/930,862, dated Dec. 5, 2023, (29 pages), United States Patent and Trademark Office, US. [cited by applicant]
Notice of Allowance and Fees Due, for U.S. Appl. No. 17/355,731, dated Oct. 31, 2023, (7 pages), United States Patent and Trademark Office, US. [cited by applicant]
Cover, Thomas M. et al. “Elements of Information Theory,” John Wiley & Sons, Inc., (565 pages), (Year: 1991), Print ISBN: 0-471-06259-6, Online ISBN: 0-471-20061-1. [cited by applicant]
Goldstein, Ira et al. “Three Approaches to Automatic Assignment of ICD-9-CM Codes to Radiology Reports,” AMIA Annual Symposium Proceedings Archive, Oct. 11, 2007, pp. 279-283. PMID: 18693842; Pmcid: PMC2655861, availabl… [cited by applicant]
Hager, Gregory E. et al. “Multiple Kernel Tracking With SSD,” In Proceedings of the 2004 IEEE Computer Society Conference On Computer Vision and Pattern Recognition, vol. 1, pp. 1-790-1-797, Jun. 27, 2004, (Year: 2004),… [cited by applicant]
Kailath, Thomas. “The Divergence and Bhattacharyya Distance Measures in Signal Selection,” IEEE Transactions On Communication Technology, vol. COM-15, No. 1, Feb. 1967, pp. 52-60. [cited by applicant]
Kumar, Amit et al. “A Fast Unsupervised Assignment Of ICD Codes With Clinical Notes Through Explanations,” SAC '22: Proceedings of the 37th ACM/SIGAPP Symposium on Applied Computing, Apr. 2022, pp. 610-618, available on… [cited by applicant]
Niblack, W. et al. “QBIC Project: Querying Images By Content, Using Color, Texture, and Shape,” In Storage and Retrieval For Image and Video Databases, SPIE vol. 1908, pp. 173-187, Apr. 14, 1993. [cited by applicant]
Puzicha, Jan et al. “Non-Parametric Similarity Measures For Unsupervised Texture Segmentation and Image Retrieval,” In Proceedings / CVPR, IEEE Computer Society Conference On Computer Vision and Pattern Recognition, Jul… [cited by applicant]
Rubner, Yossi. “Perceptual Metrics For Image Database Navigation,” PhD Dissertation, Stanford University, May 1999, (177 pages). [cited by applicant]
Shuai, Zhao et al. “Comparison Of Different Feature Extraction Methods For Applicable Automated ICD Coding,” BMC Medical Informatics and Decision Making, vol. 22, No. 11, pp. 1-15, Dec. 2022, DOI: 10.1186/s12911-022-017… [cited by applicant]
Singaravelan, Anandakumar et al. “Predicting ICD-9 Codes Using Self-Report Of Patients,” Applied Sciences, vol. 11, No. 21:10046, pp. 1-18, Oct. 17, 2021, DOI: 10.3390/app112110046. [cited by applicant]
Swain, Michael J. et al. “Color Indexing,” International Journal of Computer Vision, vol. 7, No. 1, (Year: 1991), pp. 11-32. [cited by applicant]
Tang, Xiangru et al. “CONFIT: Toward Faithful Dialogue Summarization With Linguistically-Informed Contrastive Fine-Tuning,” arXiv: 2112.08713v1 [cs.CL] Dec. 16, 2021, (11 pages), available online at https://arxiv.org/pd… [cited by applicant]
Werman, Michael et al. “A Distance Metric For Multi-Dimensional Histograms,” Computer, Vision, Graphics, and Image Processing, vol. 32, pp. 328-336, (Year: 1985), available online at http://w3.cs.huji.ac.il/˜peleg/paper… [cited by applicant]
Zhao, Qi et al. “Differential Earth Mover's Distance with Its Applications to Visual Tracking,” IEEE Transactions on Pattern Analysis and Machine Intelligence, vol. 32, No. 2, Dec. 31, 2008, pp. 274-287. [cited by applicant]
Final Office Action for U.S. Appl. No. 16/930,862, dated Jun. 21, 2023, (28 pages), United States Patent and Trademark Office, US. [cited by applicant]
NonFinal Office Action for U.S. Appl. No. 17/355,731, dated Jun. 6, 2023, (9 pages), United States Patent and Trademark Office, US. [cited by applicant]
Bittencourt, Marciele M. et al. “ML-MDL Text: A Multi label Text Categorization Technique With Incremental Learning”, 2019 8th Brazilian Conference on Intelligent Systems (BRACIS). Oct. 15-18, 2019, pp. 580-585, DOI: 10… [cited by applicant]
Jiang, Shuo et al. “Deep Learning For Technical Document Classification”, IEEE Transactions On Engineering Management, vol. PP, Issue 99, Mar. 8, 2022, pp. 1-17, DOI: 10.1109/TEM.2022.3152216. [cited by applicant]
NonFinal Office Action for U.S. Appl. No. 17/808,223, dated Sep. 28, 2023, (29 pages), United States Patent and Trademark Office, US. [cited by applicant]
Wei, Guiying et al. “Study Of Text Classification Methods For Data Sets With Huge Features”, 2010 2nd International Conference on Industrial and Information Systems, vol. 1, pp. 433-436, Jul. 10-11, 2010, DOI: 10.1109/I… [cited by applicant]
Notice of Allowance and Fees Due (PTOL-85) Mailed on Feb. 6, 2024 for U.S. Appl. No. 17/808,223, 17 page(s). [cited by applicant]
Notice of Allowance and Fees Due (PTOL-85) Mailed on Feb. 28, 2024 for U.S. Appl. No. 17/355,731, 2 page(s). [cited by applicant]
Arumae, Kristjan et al., “CALM: Continuous Adaptive Learning for Language Modeling,” arXiv:2004.03794v1 [cs.CL], Apr. 8, 2020, available online at https://arxiv.org/pdf/2004.03794.pdf. [cited by applicant]
Atutxa, Aitziber et al. “Interpretable Deep Learning To Map Diagnostic Texts To ICD-10 Codes,” International Journal of Medical Informatics, vol. 129, Sep. 2019, pp. 49-59. [cited by applicant]
Bai, Tian et al. “Improving Medical Code Prediction From Clinical Text Incorporating Online Knowledge Sources,” In Proceedings of the 2019 World Wide Web Conference, pp. 72-82, May 13-17, 2019, San Francisco, CA, USA, D… [cited by applicant]
Baumel, Ted et al. “Multi-Label Classification Of Patient Notes: Case Study On ICD Code Assignment,” In Workshops at the Thirty-Second AAAI Conference On Artificial Intelligence, Jun. 20, 2018, pp. 409-416. [cited by applicant]
Burgess, Curt et al. “Explorations In Context Space: Words, Sentences, Discourse,” Discourse Processes, vol. 25, Nos. 2-3, (1998), pp. 211-257. DOI: 10.1080/01638539809545027. [cited by applicant]
Chen, Pei-Fu et al. Automatic ICD-10 Coding and Training System: Deep Neural Network Based On Supervised Learning, JMIR Medical Informatics, Aug. 31, 2021, vol. 9, No. 8:e23230, pp. 1-13. [cited by applicant]
Delvin, Jacob et al. “BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding,” NAACL-HLT (1), May 24, 2019, pp. 4171-4186, arXiv:1810.04805 [cs.CL], available online at https://arxiv.org/abs/18… [cited by applicant]
Dharmadhikari, Shweta C. et al. “A Novel Multi Label Text Classification Model Using Semi Supervised Learning,” International Journal of Data Mining & Knowledge Management Process (IJDKP), vol. 2, No. 4, Jul. 2021, pp. … [cited by applicant]
Goldstein, Ira et al. “Three Approaches to Automatic Assignment of ICD-9-CM Codes To Radiology Reports,” AMIA Annual Symposium Proceedings Archive, Oct. 11, 2007, pp. 279-283, PMID: 18693842, PMCID: PMC2655861. [cited by applicant]
Huang, Gao et al. “Supervised Word Mover's Distance,” In Advances In Neural Information Processing Systems, (2016), pp. 4862-4870. [cited by applicant]
Islam, Aminul et al. “Semantic Text Similarity Using Corpus-Based Word Similarity and String Similarity,” ACM Transactions on Knowledge Discovery from Data (TKDD), Issue 2, No. 2, Article 10, Jul. 2008, pp. 10:1-10:25. [cited by applicant]
Kavuluru, Ramakanth et al. “Unsupervised Extraction of Diagnosis Codes from EMRs Using Knowledge-Based and Extractive Text Summarization Techniques,” Advanced Artificial Intelligence, vol. 7884, May 2013, pp. 77-88, DOI… [cited by applicant]
Kusner, Matt J. et al. “From Word Embeddings to Document Distances,” Proceedings of the 32nd International Conference On Machine Learning, vol. 37, (ICML'15), Jun. 1, 2015, pp. 957-966. [cited by applicant]
Landauer, Thomas K. et al. “An Introduction To Latent Semantic Analysis,” Discourse Processes, (1998), vol. 25, pp. 259-284. [cited by applicant]
Li, Yuhua et al. “Sentence Similarity Based on Semantic Nets and Corpus Statistics,” IEEE Transactions On Knowledge and Data Engineering, vol. 18, No. 8, Jun. 26, 2006, pp. 1-35). [cited by applicant]
Liu, Bing et al. “Text Classification by Labeling Words,” American Association for Artificial Intelligence, Jul. 25, 2004, vol. 4, (6 pages). [cited by applicant]
Mikolov, Tomas et al. “Distributed Representations of Words and Phrases and Their Compositionality,” In Advances In Neural Information Processing Systems, (2013) pp. 1-9. [cited by applicant]
Mullenbach, James et al. “Explainable Prediction of Medical Codes From Clinical Text,” arXiv:1802.056952v2 [cs.CL] Apr. 16, 2018, (11 pages). [cited by applicant]
Nigam, Priyanka. Applying Deep Learning To ICD-9 Multi-Label Classification From Medical Records. Technical Report, Stanford University, (Year: 2016), pp. 1-8, available online: http://cs224d.stanford.edu/reports/priyan… [cited by applicant]
Okazaki, Naoaki et al. “Sentence Extraction By Spreading Activation Through Sentence Similarity,” IEICE Transactions Information and Systems, vol. E82, No. 1, Jan. 1999, pp. 1-9. [cited by applicant]
Patel, Kevin et al. “Adapting Pre-Trained Word Embeddings For Use In Medical Coding,” In Biomedical Natural Language Processing Workshop (BioNLP 2017), Aug. 2017, (5 pages). [cited by applicant]
Saxena, Nihit. “Word Mover's Distance For Text Similarity,” Aug. 26, 2019, (8 pages), [Article, Online]. [Retrieved from the Internet Oct. 15, 2020]<URL: https://towardsdatascience.com/word-movers-distance-for-text-simi… [cited by applicant]
Scheurwegs, Elyne et al. “Assigning Clinical Codes With Data Driven Concept Representation On Dutch Clinical Free Text,” Journal of Biomedical Informatics, vol. 69, Apr. 8, 2017, pp. 118-127, DOI: 10.1016/j.jbi.2017.04.… [cited by applicant]
Sonabend, W. Aaron et al. “Automated ICD Coding Via Unsupervised Knowledge Integration (UNITE),” International Journal of Medical Informatics, vol. 139, Jul. 2020, pp. 104135, ISSN: 1386-5056. [cited by applicant]
Werner, Matheus et al. “Speeding Up Word Mover's Distance and Its Variants Via Properties Of Distances Between Embeddings,” arXiv:1912.005092v2 [cs.CL] May 8, 2020, (8 pages). [cited by applicant]
Xie, Pengtao et al. “A Neural Architecture for Automated ICD Coding,” Proceedings of the 56th Annual Meeting of the Association For Computational Linguistics (Long Papers), vol. 1, Jul. 15-20, 2018, pp. 1066-1076. [cited by applicant]
Xu, Keyang et al. “Multimodal Machine Learning for Automated ICD Coding,” Proceedings of Machine Learning Research, vol. 106, (18 pages), Oct. 28, 2019, PMLR. [cited by applicant]
Zhang, Minghua et al. “An Unsupervised Model With Attention Autoencoders For Question Retrieval,” The Thirty-Second AAAI Conference On Artificial Intelligence (AAAI-18), vol. 32, No. 1, Apr. 26, 2018, pp. 4978-4986. [cited by applicant]
Zhang, Minghua et al. “Learning Universal Sentence Representations with Mean-Max Attention Autoencoder,” arXiv preprint arXiv: 1809.06590v1 [cs.CL], Sep. 18, 2018, (10 pages). [cited by applicant]
NonFinal Office Action for U.S. Appl. No. 16/930,862, dated Dec. 14, 2022, (29 pages), United States Patent and Trademark Office. [cited by applicant]
Notice of Allowance and Fees Due (PTOL-85) Mailed on Jun. 12, 2024 for U.S. Appl. No. 16/930,862, 9 page(s). [cited by applicant]
Notice of Allowance and Fees Due (PTOL-85) Mailed on May 22, 2024 for U.S. Appl. No. 17/808,214, 14 page(s). [cited by applicant]