IP Library › Granted Patent US 12,423,528
Granted Patent B2
US 12,423,528 · App. 18/057,536 · Granted Sep 23, 2025

Generating sentiment models using weak labels generated from discourse markers

Inventors: Liat Ein-Dor (Tel Aviv, IL); Ilya Shnayderman (Jerusalem, IL); Artem Spector (Rishon Le-Zion, IL); Lena Dankin (Tel Aviv, IL); Ranit Aharonov (Ramat Hasharon, IL); Noam Slonim (Jerusalem, IL)
Assignee: International Business Machines Corporation
G06F40/40G06F40/205G06F40/295
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,423,528
App. No.
18/057,536
Granted
Sep 23, 2025
Kind
B2
Abstract

An example system includes a processor to receive a list of sentiment carrying discourse markers. The processor is to select sentences in a text corpus that begin with a discourse marker from the list of sentiment carrying discourse markers followed by a comma. The processor is to remove each discourse marker and comma from a beginning of the selected sentences and labeling each of the sentences with a sentiment associated with to a corresponding removed discourse marker to generate a weakly labeled dataset. The processor is to inter-train a pretrained language model using the generated weakly labeled dataset to generate a sentiment model.

Claims (37)

1. A system, comprising a processor to:

receive a list of sentiment carrying discourse markers;

select sentences in a text corpus that begin with a discourse marker from the list of sentiment carrying discourse markers followed by a comma;

remove each discourse marker and comma from a beginning of the selected sentences, and label each of the sentences with a sentiment associated with a corresponding removed discourse marker to generate a weakly labeled dataset;

inter-train a pretrained language model using the generated weakly labeled dataset to generate a sentiment model; and

identify whether a set of sentences is positive or negative sentiment, based on the sentiment model;

wherein the processor is to replace entity names with entity types in prefixes of sentences from a domain-specific corpus to generate a list of domain-specific discourse marker candidates, and select a domain-specific list of positive and negative discourse markers based on domain-specific discourse marker candidates associated with randomly sampled sentences having higher and lower sentiment scores.

2. The system of claim 1 , wherein the generated sentiment model comprises a general domain language model.

3. The system of claim 1 , wherein the processor is to automatically generate a domain-specific list of discourse markers using the generated sentiment model.

4. The system of claim 1 , wherein the processor is to extract sentences with domain-specific weak labels using an automatically generated domain-specific list of discourse markers.

5. The system of claim 1 , wherein the processor is to extract sentences with domain-specific discourse markers and require that a prediction of the sentiment model for the extracted sentences after removing the with domain-specific discourse markers is consistent with a prediction of the class associated with the discourse markers with of the class associated with the discourse markers with scores exceeding a threshold, and use the extracted sentences without the prefixes as weakly labeled domain-specific training data, where the weak labels are the classes associated with the discourse markers.

6. The system of claim 1 , wherein the processor is to fine-tune the generated sentiment model to generate a domain-specific sentiment model based on an weakly labeled domain-specific training data automatically generated using the sentiment model from a domain-specific corpus.

7. A computer-implemented method, comprising:

receiving, via a processor, a list of sentiment carrying discourse markers;

selecting, via the processor, sentences in a text corpus that begin with a discourse marker from the list of sentiment carrying discourse markers followed by a comma;

removing, via the processor, each discourse marker and comma from a beginning of the selected sentences and labeling each of the sentences with a sentiment associated with to a corresponding removed discourse marker to generate a weakly labeled dataset;

inter-training, via the processor, a pretrained language model using the generated weakly labeled dataset to generate a sentiment model; and

identifying whether a set of sentences is positive or negative sentiment, based on the sentiment model;

wherein the processor is to replace entity names with entity types in prefixes of sentences from a domain-specific corpus to generate a list of domain-specific discourse marker candidates, and select a domain-specific list of positive and negative discourse markers based on domain-specific discourse marker candidates associated with randomly sampled sentences having higher and lower sentiment scores.

8. The computer-implemented method of claim 7 , comprising automatically generating, via the processor, a domain-specific list of discourse markers using a domain-specific corpus and the sentiment model.

9. The computer-implemented method of claim 8 , wherein the automatically generating the domain-specific list of discourse markers comprises extracting sentence-opening n-grams followed by a comma from the domain-specific corpus.

10. The computer-implemented method of claim 8 , wherein the automatically generating the domain-specific list of discourse markers comprises grouping together discourse marker candidates using named entity recognition (NER).

11. The computer-implemented method of claim 8 , wherein the automatically generating the domain-specific list of discourse markers comprises randomly sampling sentences associated with a set of most frequently detected discourse marker candidates in the domain-specific corpus and running inference on the sentiment model for sentences without discourse marker prefixes to identify discourse markers associated with overrepresentation of confident predictions of the sentiment model in any of a plurality of sentiment classes.

12. The computer-implemented method of claim 7 , comprising generating, via the processor, domain-specific weakly labeled training data using an automatically generated list of domain-specific discourse markers.

13. The computer-implemented method of claim 12 , comprising fine-tuning, via the processor, the sentiment model using the generated domain-specific weakly labeled training data to generate a domain-specific sentiment model.

14. A computer program product for generating sentiment models, the computer program product comprising a computer-readable storage medium having program code embodied therewith, the program code executable by a processor to cause the processor to:

receive a list of sentiment carrying discourse markers;

select sentences in a text corpus that begin with a discourse marker from the list of sentiment carrying discourse markers followed by a comma;

remove each discourse marker and comma from a beginning of the selected sentences and labeling each of the sentences with a sentiment associated with to a corresponding removed discourse marker to generate a weakly labeled data;

inter-train a pretrained language model using the generated weakly labeled data to generate a sentiment model; and

identify whether a set of sentences is positive or negative sentiment, based on the sentiment model;

wherein the processor is to replace entity names with entity types in prefixes of sentences from a domain-specific corpus to generate a list of domain-specific discourse marker candidates, and select a domain-specific list of positive and negative discourse markers based on domain-specific discourse marker candidates associated with randomly sampled sentences having higher and lower sentiment scores.

15. The computer program product of claim 14 , further comprising program code executable by the processor to automatically generate a domain-specific list of discourse markers using a domain-specific corpus and the sentiment model.

16. The computer program product of claim 15 , further comprising program code executable by the processor to extract sentence-opening n-grams followed by a comma from the domain-specific corpus.

17. The computer program product of claim 15 , further comprising program code executable by the processor to group together discourse marker candidates using named entity recognition (NER).

18. The computer program product of claim 14 , further comprising program code executable by the processor to randomly sample sentences associated with a set of most frequently detected discourse marker candidates in the domain-specific corpus and running inference on the sentiment model for sentences without discourse marker prefixes to identify discourse markers associated with overrepresentation of confident predictions of the sentiment model in any of a plurality of sentiment classes.

19. The computer program product of claim 14 , further comprising program code executable by the processor to generate a domain-specific weakly labeled training dataset based on an automatically generated list of domain-specific discourse markers and fine-tune the sentiment model based on the domain-specific weakly labeled training dataset.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Nov 21, 2022
From: EIN-DOR, LIAT; SHNAYDERMAN, ILYA; SPECTOR, ARTEM; DANKIN, LENA; AHARONOV, RANIT; SLONIM, NOAM
To: INTERNATIONAL BUSINESS MACHINES CORPORATION
Reel/Frame 061843/0811 →
Continuity (1)
Related Publication 20240169160A1 · May 23, 2024
References Cited (19)
US 11556722B1 · Ben Shahar · 2023 [cited by examiner]
US 11734937B1 · Pushkin · 2023 [cited by examiner]
US 20140108006A1 · Vogel · 2014 [cited by examiner]
US 20150339288A1 · Baker · 2015 [cited by examiner]
US 20170161372A1 · Fernández · 2017 [cited by examiner]
US 20180189691A1 · Oehrle et al. · 2018 [cited by applicant]
US 20200279075A1 · Avedissian · 2020 [cited by examiner]
US 20210150140A1 · Galitsky · 2021 [cited by applicant]
US 20220406292A1 · Bratt · 2022 [cited by examiner]
CN 113705238A · 2021 [cited by applicant]
Jan Kocon et al., “AspectEmo: Multi-Domain Corpus of Consumer Reviews for Aspect-Based Sentiment Analysis”, 2021 International Conference on Data Mining Workshops (ICDMW), Jan. 20, 2022, 8 pages. [cited by applicant]
Liat Ein-Dor et al., “Fortunately, Discourse Markers Can Enhance Language Models for Sentiment Analysis”, IBM, Apr. 5, 2022, 10 pages. [cited by applicant]
Deriu et al. “Leveraging Large Amounts of Weakly Supervised Data for Multi-Language Sentiment Classification”, WWWW'17: Proceedings of the 26th International Conference on World Wide Web, Apr. 3, 2017, pp. 1045-1052. [cited by applicant]
Ke et al. “SentiLARE: Sentiment-Aware Language Representation Learning with Linguistic Knowledge” arXiv: 1911.02493v3 [cs.CL], Sep. 24, 2020, 14 pages. [cited by applicant]
Mukherjee et al. “Sentiment Analysis in Twitter with Lightweight Discourse Analysis”, COLING 2012, 24th International Conference on Computational Linguistics, Proceedings of the Conference: Technical Papers, Dec. 8-15, … [cited by applicant]
Severyn et al. “UNITN: Training Deep Convolutional Neural Network for Twitter Sentiment Classification”, ACL Anthology, Jun. 2015, pp. 464-469. [cited by applicant]
Sileo et al. “Mining Discourse Markers for Unsupervised Sentence Representation Learning”, Proceedings of NAACL-HLT, Jun. 2019, 10 pages. [cited by applicant]
Yin et al. “SentiBERT: A Transferable Transformer-Based Architecture for Compositional Sentiment Semantics”, arXiv:2005.04114v4 [cs.CL], May 21, 2020, 12 pages. [cited by applicant]
Zhou et al. “SentiX: A Sentiment-Aware Pre-Trained Model for Cross-Domain Sentiment Analysis”, ACL Anthology, Dec. 2020, 12 pages. [cited by applicant]