IP Library Granted Patent US 12,475,407
Granted Patent B2
US 12,475,407 · App. 17/081,277 · Granted Nov 18, 2025

Predicting topic sentiment using a machine learning model trained with observations in which the topics are masked

Inventors: Guilford T. Parsons (Seattle, WA); Kaitlin Claveau (Westminster, CO)
Assignee: Providence St. Joseph Health
G06N20/00G06F40/40G16H10/20G16H50/20
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,475,407
App. No.
17/081,277
Granted
Nov 18, 2025
Kind
B2
Abstract

A facility for determining sentiments expressed by a natural-language text string for each of one or more topics is described. In the natural-language text string, the facility identifies one or more topics. For each identified topic, the facility replaces the topic in the natural-language text string with a masking tag that occupies the same position in the natural-language text string as the topic. After the replacing, the facility applies a machine learning model to the natural-language text string to obtain a predicted sentiment for each of the identified topics.

Claims (44)

1 . One or more memories collectively having contents configured to cause a computing system to perform a method, the one or more memories not constituting a propagating transitory data signal, the method comprising:

receiving, by a processor of the computing system, a plurality of training natural-language text strings;

identifying, by the processor of the computing system, one or more noun phrases in each of the plurality of training natural-language text strings;

determining, by the processor of the computing system, one or more sentiment-qualified topics based on the one or more noun phrases;

determining, by the processor of the computing system, an entity class for each of the one or more sentiment-qualified topics;

modifying, by the processor of the computing system, each training natural-language text string to preserve a location of each sentiment-qualified topic in each training natural-language text string while replacing all information about an identity of each sentiment-qualified topic;

using, by the processor of the computing system, the modified plurality of training natural-language strings and the entity class for each of the one or more sentiment-qualified topics to train a machine learning model to determine a sentiment of each sentiment-qualified topic by at least performing a largest-first comparison of substrings of each of the plurality of training natural-language text strings to identify longer, multi-word topics to an exclusion of included shorter topics, and comparing the longer, multi-word topics to a list of topics or named entities specified for a domain; and

storing, by the processor of the computing system, the sentiment with each training natural-language text string by restoring the identity and using the location associated with each sentiment-qualified topic in each training natural-language text string.

2 . The one or more memories of claim 1 , the method further comprising:

initializing the machine learning model as a Bidirectional Encoder Representations from Transformers model.

3 . The one or more memories of claim 1 , the method further comprising:

initializing the machine learning model as a Clinical Bidirectional Encoder Representations from Transformers model.

4 . A method in a computing system, the method comprising:

receiving, by a processor of the computing system, a plurality of training natural-language text strings;

identifying, by the processor of the computing system, one or more noun phrases in each of the plurality of training natural-language text strings;

determining, by the processor of the computing system, one or more sentiment-qualified topics based on the one or more noun phrases;

determining, by the processor of the computing system, an entity class for each of the one or more sentiment-qualified topics;

modifying, by the processor of the computing system, each training natural-language text string to preserve a location of each sentiment-qualified topic in each training natural-language text string while replacing all information about an identity of each sentiment-qualified topic;

using, by the processor of the computing system, the modified plurality of training natural-language strings and the entity class for each of the one or more sentiment-qualified topics to train a machine learning model to determine a sentiment of each sentiment-qualified topic by at least performing a largest-first comparison of substrings of each of the plurality of training natural-language text strings to identify longer, multi-word topics to an exclusion of included shorter topics, and comparing the longer, multi-word topics to a list of topics or named entities specified for a domain; and

storing, by the processor of the computing system, the sentiment with each training natural-language text string by restoring the identity and using the location associated with each sentiment-qualified topic in each training natural-language text string.

5 . The method of claim 4 , the method further comprising:

initializing the machine learning model as a Bidirectional Encoder Representations from Transformers model.

6 . The method of claim 4 , the method further comprising:

initializing the machine learning model as a Clinical Bidirectional Encoder Representations from Transformers model.

7 . The method of claim 4 , wherein the topic to which the training natural-language text strings is related is a distinguished domain of expression, the method further comprising:

receiving a subject input text string related to the distinguished domain of expression;

masking an intent expressed by the subject text string for each of one or more masked noun phrases and one or more entity classes determined for each of the masked noun phrases for which an intent is identified; and

applying the trained model to the masked subject input text string to predict a sentiment for each masked noun phrase of the subject input text string.

8 . A system comprising:

at least one processor; and

at least one memory, the at least one memory not constituting a transitory propagating data signal, wherein the at least one memory is configured to cause the at least one processor to:

receive a plurality of training natural-language text strings;

identify one or more noun phrases in each of the plurality of training natural-language text strings;

determine one or more sentiment-qualified topics based on the one or more noun phrases;

determine an entity class for each of the one or more sentiment-qualified topics:

modify each training natural-language text string to preserve a location of each sentiment-qualified topic in each training natural-language text string while replacing all information about an identity of each sentiment-qualified topic;

use the modified plurality of training natural-language strings and the entity class for each of the one or more sentiment-qualified topics to train a machine learning model to determine a sentiment of each sentiment-qualified topic by at least performing a largest-first comparison of substrings of each of the plurality of training natural-language text strings to identify longer, multi-word topics to an exclusion of included shorter topics, and comparing the longer, multi-word topics to a list of topics or named entities specified for a domain; and

store the sentiment with each training natural-language text string by restoring the identity and using the location associated with each sentiment-qualified topic in each training natural-language text string.

9 . The system of claim 8 , wherein the at least one processor is further configured to initialize the machine learning model as a Bidirectional Encoder Representations from Transformers model.

10 . The system of claim 8 , wherein the at least one processor is further configured to initialize the machine learning model as a Clinical Bidirectional Encoder Representations from Transformers model.

11 . The system of claim 8 , wherein the topic to which the training natural-language text strings is related is a distinguished domain of expression, wherein the at least one processor is further configured to:

receive a subject input text string related to the distinguished domain of expression;

mask an intent expressed by the subject text string for each of one or more masked noun phrases and one or more entity classes determined for each of the masked noun phrases for which an intent is identified; and

apply the trained model to the masked subject input text string to predict a sentiment for each masked noun phrase of the subject input text string.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Feb 17, 2022
From: PARSONS, GUILFORD T.; CLAVEAU, KAITLIN
To: PROVIDENCE ST. JOSEPH HEALTH
Reel/Frame 059033/0174 →
Continuity (1)
Related Publication 20220129784A1 · Apr 28, 2022
References Cited (17)
US 12210836B2 · Batra · 2025 [cited by examiner]
US 20110112995A1 · Chang · 2011 [cited by examiner]
US 20170270096A1 · Sheafer · 2017 [cited by examiner]
US 20190356956A1 · Sheng · 2019 [cited by examiner]
US 20190370604A1 · Galitsky · 2019 [cited by examiner]
US 20190378179A1 · Cleverley · 2019 [cited by examiner]
US 20210312124A1 · Agarwal · 2021 [cited by examiner]
Huang, Kexin, Jaan Altosaar, and Rajesh Ranganath. “Clinicalbert: Modeling clinical notes and predicting hospital readmission.” arXiv preprint arXiv:1904.05342v2 (Apr. 2019). (Year: 2019). [cited by examiner]
Chen, Long. “Assertion detection in clinical natural language processing: a knowledge-poor machine learning approach.” 2019 IEEE 2nd International Conference on Information and Computer Technologies (ICICT). IEEE, May 2… [cited by examiner]
Zhang, Denghui, et al. “E-BERT: Adapting BERT to E-commerce with Adaptive Hybrid Masking and Neighbor Product Reconstruction.” arXiv preprint arXiv:2009.02835 (Sep. 2020). (Year: 2020). [cited by examiner]
Carrillo-de-Albornoz, Jorge, Javier Rodriguez Vidal, and Laura Plaza. “Feature engineering for sentiment analysis in e-health forums.” PloS one 13.11 (2018): e0207996. (Year: 2018). [cited by examiner]
Coroiu, Adriana Mihaela, Alina Delia Calin, and Maria Nutu. “Topic modeling in medical data analysis. case study based on medical records analysis.” 2019 International Conference on Software, Telecommunications and Comp… [cited by examiner]
Alsentzer et al., “Publicly Available Clinical BERT Embeddings,” [cited by applicant]
Delvin et al., “BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding,” [cited by applicant]
Harkema et al., “Context: An Algorithm for Determining Negation, Experiencer, and Temporal Status from Clinical Reports,” [cited by applicant]
Mehrabi et al., “Deepen: A negation detection system for clinical text incorporating dependency relation into NegEx,” [cited by applicant]
Yi et al., “Sentiment Analyzer: Extracting Sentiments about Given Topic using Natural Language Processing Techniques, ” [cited by applicant]