IP Library Granted Patent US 11,416,680
Granted Patent B2
US 11,416,680 · App. 15/241,040 · Granted Aug 16, 2022

Classifying social media inputs via parts-of-speech filtering

Inventors: Danqing Cai (Singapore, SG); Wei Tah Chai (Singapore, SG); Pek Gnee Ng (Singapore, SG); Subashini Rengarajan (Singapore, SG); Xin Zheng (Singapore, SG); Hang Guo (Singapore, SG); Weile Chen (Singapore, SG)
Assignee: SAP SE
G06F40/279G06F16/353
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,416,680
App. No.
15/241,040
Granted
Aug 16, 2022
Kind
B2
Abstract

Described herein is a framework for classifying social media inputs. In accordance with one aspect of the framework, one or more social media inputs is acquired from one or more social media platforms. The social media inputs are cleaned to remove redundant elements. One or more features are extracted from the cleaned social media inputs. The social media inputs are classified by a trained classifier into predefined categories using the extracted one or more features.

Claims (70)

1. A computer-implemented method for automatically classifying social media inputs comprising:

acquiring one or more social media inputs from one or more social media platforms;

cleaning the one or more social media inputs to remove redundant elements, wherein cleaning the one or more social media inputs comprises removing any nouns using part-of-speech (POS) tagging;

extracting one or more features from a cleaned social media input, wherein the one or more features are enhanced to define a list of input features including the one or more extracted features and at least one subsequent feature;

iteratively training a plurality of classifiers to classify social media input into a category from predefined categories, wherein the training comprises training using a training data set including social media input retrieved from a database, wherein each of the social media inputs of the training data set is labeled and assigned to one of the predefined categories;

dynamically selecting a trained classifier from among the plurality of trained classifiers, wherein the selected trained classifier is a current best fit classifier compared to one or more other classifiers from the iteratively trained plurality of classifiers, wherein the best fit classifier is a classifier that produces least misclassification compared to classifiers of the one or more other classifiers, wherein the plurality of trained classifiers are trained over same training data; and

automatically classifying, by the trained classifier, the one or more social media inputs into the predefined categories using corresponding one or more lists of input features, wherein the classifying comprises:

calculating probabilities that an extracted feature from a respective list of input features belongs to each of the predefined categories, wherein the extracted feature is extracted from a first social media input from the one or more social media inputs;

calculating probabilities that the extracted features as a subsequent feature of the at least one subsequent feature from the list of input features belongs to each of the predefined categories, wherein the probabilities are estimated based on a context of one or more n-grams of features preceding the subsequent features, wherein an n-gram includes two or more features from the list of input features and is considered as a single feature; and

selecting a predefined category for each feature, wherein the selected predefined category is a category having a highest probability that a respective feature belongs to the category;

wherein the predefined categories comprise:

comments, comprising consumers' opinions or feedbacks regarding a purchased product or service received from a provider;

complaints, comprising statements or expressions of unsatisfactory and discontent regarding a purchased product or service rendered by a provider; and

queries, comprising requests for information about a product or service;

wherein the extracted one or more features are variables associated with a feature space that is divided into different parts corresponding to the predefined categories, wherein collections of variables for each part of the feature space are mutually exclusive, wherein each predefined category of the predefined categories is associated with a corresponding list of predefined expressions, wherein a probability that a variable of the variables is associated with a first category from the predefined categories is the highest as compared to the other categories when the variable is associated with one of the expressions from a first list of predefined expressions corresponding to the first category.

2. The method of claim 1 comprising selecting the predefined categories.

3. The method of claim 1 wherein acquiring the one or more social media inputs comprises acquiring the social media inputs from the one or more social media platforms via an application programming interface (API).

4. The method of claim 1 wherein cleaning the one or more social media inputs comprises removing the redundant elements using one or more predefined filters.

5. The method of claim 1 wherein the redundant elements comprise emojis, special characters, Uniform Resource Locators (URLs), pictures, stop words and nouns.

6. The method of claim 1 wherein cleaning the one or more social media inputs comprises removing any stop words.

7. The method of claim 1 wherein extracting the one or more features comprises identifying the one or more features from the cleaned one or more social media inputs using term frequency and inverse document frequency (TF-IDF) weighting.

8. The method of claim 1 wherein the extracted one or more features comprises an n-gram.

9. The method of claim 1 comprising presenting a classification result to a user via a user interface.

10. The method of claim 9 wherein the presenting the classification result comprises displaying the selected predefined categories and the one or more social media inputs assigned to predefined categories.

11. The method of claim 10 wherein the method comprises:

receiving, through the user interface, input from the user to manually revise the assigned predefined category of the social media input to another predefined category;

retraining the plurality of classifiers based on the another predefined category revised by the user; and

re-selecting a new trained classifier from the retrained plurality of classifiers as a new best fit classifier compared to other classifiers associated with other trained classifiers, wherein the re-selected new trained classifier is used to automatically classify the one or more social media input.

12. The method of claim 11 wherein the social media input with a revised category is tagged and re-used in re-training of one or more classifiers.

13. The method of claim 1 wherein the one or more classifiers comprises Random Forest, Naive Bayes, Logistic Regression and Support Vector Machine (SVM).

14. The method of claim 13 wherein a classifier having the highest accuracy score is selected from the one or more trained classifiers.

15. The method of claim 1 wherein the each of the classifiers comprises a feature space which is divided into different parts, wherein each of the categories occupies a part of the feature space and each part of the feature spaces includes a collection of variables.

16. A system for classifying social media inputs comprising:

a non-transitory computer-readable medium for storing a database and computer-readable program code; and

one or more processors in communication with the non-transitory computer-readable medium, the one or more processors being operative with the computer-readable program code to perform operations including:

acquiring one or more social media inputs from one or more social media platforms;

cleaning the one or more social media inputs to remove redundant elements, wherein cleaning the one or more social media inputs comprises removing stop words and filtering out any nouns using part-of-speech (POS) tagging, wherein the part-of-speech tagging comprises assigning words a part-of-speech;

extracting one or more features from a cleaned social media input, wherein the one or more features are enhanced to define a list of input features including the one or more extracted features and at least one subsequent feature;

iteratively training a plurality of classifiers to classify social media input into a category from predefined categories, wherein the training comprises training using a training data set including social media input retrieved from a database, wherein each of the social media inputs of the training data set is labeled and assigned to one of the predefined categories;

dynamically selecting a trained classifier from among the plurality of trained classifiers, wherein the selected trained classifier is a current best fit classifier compared to one or more other classifiers from the iteratively trained plurality of classifiers, wherein the best fit classifier is a classifier that produces least misclassification compared to classifiers of the one or more other classifiers, wherein the plurality of trained classifiers are trained over same training data; and

automatically classifying, by the trained classifier, the one or more social media inputs into predefined categories using corresponding one or more lists of input features, wherein the classifying comprises:

calculating probabilities that an extracted feature from a respective list of input features belongs to each of the predefined categories, wherein the extracted feature is extracted from a first social media input from the one or more social media inputs;

calculating probabilities that the extracted features as a subsequent feature of the at least one subsequent feature from the list of input features belongs to each of the predefined categories, wherein the probabilities are estimated based on a context of one or more n-grams of features preceding the subsequent features, wherein an n-gram includes two or more features from the list of input features and is considered as a single feature; and

selecting a predefined category for each feature, wherein the selected predefined category is a category having a highest probability that a respective feature belongs to the category;

wherein the predefined categories comprise:

comments, comprising consumers' opinions or feedbacks regarding a purchased product or service received from a provider;

complaints, comprising statements or expressions of unsatisfactory and discontent regarding a purchased product or service rendered by a provider; and

queries, comprising requests for information about a product or service;

wherein the extracted one or more features are variables associated with a feature space that is divided into different parts corresponding to the predefined categories, wherein collections of variables for each part of the feature space are mutually exclusive, wherein each predefined category of the predefined categories is associated with a corresponding list of predefined expressions, wherein a probability that a variable of the variables is associated with a first category from the predefined categories is the highest as compared to the other categories when the variable is associated with one of the expressions from a first list of predefined expressions corresponding to the first category.

17. The system of claim 16 wherein the stop words comprise:

“is,”

“the”

“that,”

“at,” and

“which”.

18. A non-transitory computer-readable medium having stored thereon program code, the program code executable by a computer to perform steps comprising:

acquiring one or more social media inputs from one or more social media platforms;

cleaning the one or more social media inputs to remove redundant elements, wherein cleaning the one or more social media inputs comprises removing any nouns using Part-of-Speech (POS) tagging;

extracting one or more features from a cleaned social media input, wherein the one or more features are enhanced to define a list of input features including the one or more extracted features and at least one subsequent feature;

iteratively training a plurality of classifiers to classify social media input into a category from predefined categories, wherein the training comprises training using a training data set including social media input retrieved from a database, wherein each of the social media inputs of the training data set is labeled and assigned to one of the predefined categories;

dynamically selecting a trained classifier from among the plurality of trained classifiers, wherein the selected trained classifier is a current best fit classifier compared to one or more other classifiers from the iteratively trained plurality of classifiers, wherein the best fit classifier is a classifier that produces least misclassification compared to classifiers of the one or more other classifiers, wherein the plurality of trained classifiers are trained over same training data; and

automatically classifying, by the trained classifier, the one or more social media inputs into predefined categories using corresponding one or more lists of input features, wherein the classifying comprises:

calculating probabilities that an extracted feature from a respective list of input features belongs to each of the predefined categories, wherein the extracted feature is extracted from a first social media input from the one or more social media inputs;

calculating probabilities that the extracted features as a subsequent feature of the at least one subsequent feature from the list of input features belongs to each of the predefined categories, wherein the probabilities are estimated based on a context of one or more n-grams of features preceding the subsequent features, wherein an n-gram includes two or more features from the list of input features and is considered as a single feature; and

selecting a predefined category for each feature, wherein the selected predefined category is a category having a highest probability that a respective feature belongs to the category;

wherein the predefined categories comprise:

comments, comprising consumers' opinions or feedbacks regarding a purchased product or service received from a provider;

complaints, comprising statements or expressions of unsatisfactory and discontent regarding a purchased product or service rendered by a provider; and

queries, comprising requests for information about a product or service;

wherein the extracted one or more features are variables associated with a feature space that is divided into different parts corresponding to the predefined categories, wherein collections of variables for each part of the feature space are mutually exclusive, wherein each predefined category of the predefined categories is associated with a corresponding list of predefined expressions, wherein a probability that a variable of the variables is associated with a first category from the predefined categories is the highest as compared to the other categories when the variable is associated with one of the expressions from a first list of predefined expressions corresponding to the first category.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Aug 18, 2016
From: CAI, DANQING; CHAI, WEI TAH; NG, PEK GNEE; RENGARAJAN, SUBASHINI; ZHENG, XIN; GUO, HANG; CHEN, WEILE
To: SAP SE
Reel/Frame 039479/0878 →
Continuity (1)
Related Publication 20180053116A1 · Feb 22, 2018