IP Library Granted Patent US 8,725,495
Granted Patent B2
US 8,725,495 · App. 13/082,963 · Granted May 13, 2014

Systems, methods and devices for generating an adjective sentiment dictionary for social media sentiment analysis

Inventors: Wei Peng (Sunnyvale, CA); Dae Hoon Park (Champaign, IL)
Assignee: Xerox Corporation
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 8,725,495
App. No.
13/082,963
Granted
May 13, 2014
Kind
B2
Abstract

Embodiments generally relate to systems and methods for generating a sentiment dictionary and calculating sentiment scores of adjectives within the sentiment dictionary. A set of seed words can be identified and expanded using synonyms and antonyms of the set of seed words. Social media data can be parse to identify adjectives that link to the set of seed words with the words “and” or “but.” Matrices representing the attraction and repulsion among the linked adjectives can be generated. A factorization algorithm can be minimized to determine an output matrix that comprises positive and negative sentiment scores for each of the adjectives. In embodiments, a sentiment score for part of all of the social media data can be calculated using the output matrix, and one or more parts of the social media data can be classified as a positive or negative sentiment.

Claims (60)

1. A method of processing data, the method comprising:

identifying a set of seed words comprising words defined as either positive words or negative words;

extracting, from a set of data, adjectives linked to the set of seed words with “and”;

extracting, from the set of data, adjectives linked to the set of seed words with “but”;

determining a first value indicating a first frequency with which the adjectives are linked to the set of seed words with “and”;

determining a second value indicating a second frequency with which the adjectives are linked to the set of seed words with “but”;

calculating a synonym score using an equation x ij =w ij + +Log(d ij + +1), wherein x ij corresponds to the synonym score, w ij + corresponds to a number of times that a word i and a word j appear to be synonyms, and d ij + corresponds to the first equency; and

calculating, by a processor, sentiment scores for each adjective of the set of data based on the synonym score and an antonym score calculated using the second value.

2. The method of claim 1 , wherein identifying the set of seed words comprises:

identifying the set of seed words; and

generating an expanded set of seed words by adding synonyms and antonyms of the set of seed words to the set of seed words.

3. The method of claim 1 , wherein the set of seed words comprises a positive set and a negative set.

4. The method of claim 1 , wherein each of the sentiment scores comprises a positive sentiment score and a negative sentiment score.

5. The method of claim 1 , further comprising calculating the antonym score using an equation c ij =I ij +Log(d ij − +1), wherein c ij corresponds to the antonym score, I ij equals 1 if the word i is in a positive set of the set of seed words and the word j is in a negative set of the set of seed words, or vice versa, and equals 0 otherwise, and d ij − corresponds to the second frequency of.

6. The method of claim 5 , wherein the sentiment scores can be calculated by minimizing an equation ∥X−HSH T ∥ F 2 +αTr(H T CH), wherein:

X is a nonnegative symmetric graph matrix formulated using the synonym score;

H is a matrix comprising the sentiment scores;

S is a modifier to allow H to be closer to a form of cluster indicators;

α is an input parameter;

C is a constraint weight matrix formulated using the antonym score; and

∥ ∥ F indicates a Frobenius norm.

7. The method of claim 1 , further comprising:

classifying each adjective of the set of data as either a positive word or a negative word.

8. The method of claim 1 , wherein the set of data is social media data.

9. The method of claim 1 , further comprising:

determining a sentiment score for at least part of the set of data.

10. A system for processing data, comprising:

an interface to a storage device configured to store a set of data; and

a processor, communicating with the storage device via the interface, the processor being configured to:

identify a set of seed words comprising words defined as either positive words or negative words;

extract, from the set of data, adjectives linked to the set of seed words with “and”;

extract, from the set of data, adjectives linked to the set of seed words with “but”;

determine a first value indicating a first frequency with which the adjectives are linked to the set of seed words with “and”;

determine a second value indicating a second frequency with which the adjectives are linked to the set of seed words with “but”;

calculate a synonym score using an equation x ij =w ij + +Log(d ij + +1), wherein x ij corresponds to the synonym score, w ij + corresponds to a number of times that a word i and a word j appear to be synonyms, and d ij + corresponds to the first frequency; and

calculate, by a processor, sentiment scores for each adjective of the set of data based on the synonym score and an antonym score calculated using the second value.

11. The system of claim 10 , wherein identifying the set of seed words comprises:

identifying the set of seed words; and

generating an expanded set of seed words by adding synonyms and antonyms of the set of seed words to the set of seed words.

12. The system of claim 10 , wherein the set of seed words comprises a positive set and a negative set.

13. The system of claim 10 , wherein each of the sentiment scores comprises a positive sentiment score and a negative sentiment score.

14. The system of claim 10 , wherein the processor is further configured to calculate the antonym score using an equation c ij =I ij +Log(d ij − +1), wherein c ij corresponds to the antonym score, I ij equals 1 if the word i is in a positive set of the set of seed words and the word j is in a negative set of the set of seed words, or vice versa, and equals 0 otherwise, and d ij − corresponds to the second frequency.

15. The system of claim 14 , wherein the sentiment scores can be calculated by minimizing an equation ∥X−HSH T ∥ F 2 +αTr(H T CH), wherein:

X is a nonnegative symmetric graph matrix formulated using the synonym score;

H is a matrix comprising the sentiment scores;

S is a modifier to allow H to be closer to a form of cluster indicators;

α is an input parameter;

C is a constraint weight matrix formulated using the antonym score; and

∥ ∥ F indicates a Frobenius norm.

16. The system of claim 10 , wherein the processor is further configured to:

classify each adjective of the set of data as either a positive word or a negative word.

17. The system of claim 10 , wherein the set of data is social media data.

18. The system of claim 10 , wherein the processor is further configured to:

determine a sentiment score for at least part of the set of data.

19. The method of claim 1 , wherein:

the sentiment scores are calculated using graph matrixes; and

the synonym score and the antonym score indicate graph edge weights.

20. The system of claim 10 , wherein:

the sentiment scores are calculated using graph matrixes; and

the synonym score and the antonym score indicate graph edge weights.

Assignments (9)
SECOND LIEN NOTES PATENT SECURITY AGREEMENT Recorded Jul 2, 2025
From: XEROX CORPORATION
To: U.S. BANK TRUST COMPANY, NATIONAL ASSOCIATION, AS COLLATERAL AGENT
Reel/Frame 071785/0550 →
FIRST LIEN NOTES PATENT SECURITY AGREEMENT Recorded Apr 11, 2025
From: XEROX CORPORATION
To: U.S. BANK TRUST COMPANY, NATIONAL ASSOCIATION, AS COLLATERAL AGENT
Reel/Frame 070824/0001 →
TERMINATION AND RELEASE OF SECURITY INTEREST IN PATENTS RECORDED AT RF 064760/0389 Recorded Feb 13, 2024
From: CITIBANK, N.A., AS COLLATERAL AGENT
To: XEROX CORPORATION
Reel/Frame 068261/0001 →
SECURITY INTEREST Recorded Feb 13, 2024
From: XEROX CORPORATION
To: CITIBANK, N.A., AS COLLATERAL AGENT
Reel/Frame 066741/0001 →
SECURITY INTEREST Recorded Nov 20, 2023
From: XEROX CORPORATION
To: JEFFERIES FINANCE LLC, AS COLLATERAL AGENT
Reel/Frame 065628/0019 →
SECURITY INTEREST Recorded Jun 22, 2023
From: XEROX CORPORATION
To: CITIBANK, N.A., AS COLLATERAL AGENT
Reel/Frame 064760/0389 →
RELEASE OF SECURITY INTEREST IN PATENTS AT R/F 062740/0214 Recorded May 18, 2023
From: CITIBANK, N.A., AS AGENT
To: XEROX CORPORATION
Reel/Frame 063694/0122 →
SECURITY INTEREST Recorded Nov 10, 2022
From: XEROX CORPORATION
To: CITIBANK, N.A., AS AGENT
Reel/Frame 062740/0214 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Apr 8, 2011
From: PENG, WEI; PARK, DAE HOON
To: XEROX CORPORATION
Reel/Frame 026097/0535 →
Continuity (1)
Related Publication 20120259616A1 · Oct 11, 2012