IP Library Patent Application 18393312
Patent Application
App. No. 18/393,312

SYSTEMS, METHODS, AND ARTICLES FOR ENHANCING THE TRAINING OF NATURAL LANGUAGE PROCESSING MODELS IN BIOMEDICAL CONTEXT

Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US None
App. No.
18/393,312
Abstract

Technologies for enhancing natural language processing (NLP) model training in biomedical context are disclosed. An example method includes training an embedding model to capture semantic richness of biomedical context based on unstructured texts obtained from an electronic health record (EHR) system, obtaining a seed set of seed texts and an unlabeled set of unlabeled texts, using the trained embedding model to determine a vectorized semantic representation for each seed text and each unlabeled text, assigning classification labels to at least a subset of the unlabeled set by clustering the seed set with the unlabeled set based on the vectorized semantic representations for each seed text and unlabeled text, and providing NLP model training data including the assigned classification labels.

Claims (40)

1 . A method for enhancing natural language processing (NLP) model training in biomedical context, the method comprising:

training an embedding model to capture semantic richness of biomedical context based, at least in part, on a set of unstructured texts obtained from an electronic health record (EHR) system;

obtaining a seed set including seed texts each associated with a classification label;

obtaining an unlabeled set including unlabeled texts;

determining, via the trained embedding model, a vectorized semantic representation for each seed text of the seed set and for each unlabeled text of the unlabeled set;

assigning classification labels to at least a subset of the unlabeled set by clustering the seed set with the unlabeled set based, at least in part, on the vectorized semantic representations for each seed text and for each unlabeled text; and

providing NLP model training data including the assigned classification labels.

2 . The method of claim 1 , wherein the set of unstructured texts includes context snippets identified from clinical notes.

3 . The method of claim 2 , wherein the context snippets are identified based, at least in part, on at least one of demographics, medical history, diagnosis, severity of disease, medication, therapy, surgery, or associated outcome.

4 . The method of claim 1 , wherein training the embedding model comprises using at least the set of unstructured texts to finetune a large language model (LLM), wherein the LLM was pretrained for general-purpose language understanding.

5 . The method of claim 4 , further comprising, for each unstructured text of the set of unstructured texts:

extracting one or more entities of interest from the unstructured text; and

replacing at least a subset of the one or more extracted entities of interest with one or more types of mask tokens in the unstructured text to generate respective masked text.

6 . The method of claim 5 , wherein the at least a subset of the one or more extracted entities is selected based, at least in part, on a task of an NLP model to be trained with the NLP model training data.

7 . The method of claim 6 , wherein the task of the NLP model includes at least one of predicting diagnoses, identifying biomarker, recommending treatment, or determining medication intake.

8 . The method of claim 5 , further comprising labeling each masked text with at least one target label based, at least in part, on the one or more entities of interest extracted from the unstructured text.

9 . The method of claim 8 , wherein the labeling comprises normalizing the one or more entities of interest based, at least on part, on medical or clinical ontology.

10 . The method of claim 8 , further comprising adapting the LLM to use each masked text to predict its associated target label.

11 . The method of claim 10 , wherein the parameters of the LLM are adjusted during the adapting and fixed after the adapting is completed.

12 . The method of claim 1 , wherein given an input to the trained embedding model, a vectorized semantic representation of the input is generated based, at least in part, on output of one or more layers of the trained embedding model.

13 . The method of claim 12 , wherein the vectorized semantic representation of the input is generated by averaging the output of the one or more layers.

14 . The method of claim 1 , wherein assigning classification labels to at least a subset of the unlabeled set by clustering the seed set with the unlabeled set comprises iteratively performing the clustering while expanding the seed set.

15 . The method of claim 1 , wherein the classification labels assigned to the at least a subset of the unlabeled set is based, at least in part, on one or more nearest neighbors in a finalized seed set.

16 . The method of claim 1 , wherein at least a subset of the assigned classification labels is used to generate, modify, or supplement structured texts in the EHR system.

17 . The method of claim 1 , wherein at least a subset of the assigned classification labels is used in conjunction with data obtained from EHR system as input into at least one of a classification, prediction, or association model to produce output.

18 . The method of claim 1 , wherein the NLP model training data is used to train at least one of a support vector machine (SVM), neural network, or large language model (LLM).

19 . A computing system for enhancing natural language processing (NLP) model training in biomedical context, the computing system comprising:

one or more processors; and

one or more non-transitory computer-readable media collectively storing instructions that, when collectively executed by the one or more processors, cause the computing system to perform actions, the actions comprising:

training an embedding model to capture semantic richness of biomedical context based, at least in part, on a set of unstructured texts obtained from an electronic health record (EHR) system;

obtaining a seed set including seed texts each associated with a classification label;

obtaining an unlabeled set including unlabeled texts;

determining, via the trained embedding model, a vectorized semantic representation for each seed text of the seed set and for each unlabeled text of the unlabeled set; and

assigning classification labels to at least a subset of the unlabeled set by clustering the seed set with the unlabeled set based, at least in part, on the vectorized semantic representations for each seed text and for each unlabeled text.

20 . A non-transitory processor-readable storage medium that stores computer instructions that, when executed by one or more processors, cause the one or more processors to perform actions comprising:

training an embedding model to capture semantic richness of biomedical context based, at least in part, on a set of unstructured texts obtained from an electronic health record (EHR) system;

obtaining a seed set including seed texts each associated with a classification label;

obtaining an unlabeled set including unlabeled texts;

determining, via the trained embedding model, a vectorized semantic representation for each seed text of the seed set and for each unlabeled text of the unlabeled set; and

assigning classification labels to at least a subset of the unlabeled set by clustering the seed set with the unlabeled set based, at least in part, on the vectorized semantic representations for each seed text and for each unlabeled text.

Assignments (4)
RELEASE OF SECURITY INTEREST Recorded May 14, 2026
From: ARES CAPITAL CORPORATION, AS COLLATERAL AGENT
To: TEMPUS AI, INC. (F/K/A TEMPUS LABS, INC.)
Reel/Frame 075577/0513 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Aug 19, 2025
From: FARINA, NICHOLAS; MIZZI, JENNIFER; SMITH, VINCENT; LA ROCCA, SARA; PRINCE, MANISH; KUMAR, TAPAN
To: ARISTOCRAT TECHNOLOGIES, INC.
Reel/Frame 072062/0877 →
SECURITY INTEREST Recorded Jun 2, 2025
From: TEMPUS AI, INC.
To: ARES CAPITAL CORPORATION, AS COLLATERAL AGENT
Reel/Frame 071468/0107 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Mar 26, 2024
From: KANG, TIAN
To: TEMPUS AI, INC.
Reel/Frame 066905/0603 →