IP Library › Granted Patent US 10,192,457
Granted Patent B2
US 10,192,457 · App. 13/408,547 · Granted Jan 29, 2019

Enhancing knowledge bases using rich social media

Inventors: Jitendra Ajmera (New Delhi, IN); Shantanu Ravindra Godbole (New Delhi, IN); Himabindu Lakkaraju (Bangalore, IN); Bernard Andrew Roden (Middleton, WI); Ashish Verma (New Delhi, IN)
Assignee: INTERNATIONAL BUSINESS MACHINES CORPORATION
G09B7/00
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 10,192,457
App. No.
13/408,547
Granted
Jan 29, 2019
Kind
B2
Abstract

Methods and arrangements for developing knowledge bases from social media. A question is obtained from social media. Social media are consulted, and a legitimacy of the question is ascertained. All the answers to the question are harvested from the social media including the rich media that is associated with these answers, and the question is filtered out if determined not to be legitimate.

Claims (47)

1. An apparatus comprising:

at least one processor; and

a computer readable storage medium having computer readable program code embodied therewith and executable by the at least one processor, the computer readable program code comprising:

computer readable program code configured to establish at least one legitimacy standard for filtering questions, wherein the at least one legitimacy standard includes presence of a question pattern and at least one exception relative to the question pattern, wherein a question pattern identifies data as requesting additional information;

computer readable program code configured to automatically obtain a question from at least one social media conversation, wherein to obtain a question comprises obtaining data from at least one social media forum, extracting domain-specific communications directed to a target domain by filtering the data using domain-specific keywords, and identifying questions from the extracted domain-specific communications by determining presence of a question pattern;

computer readable program code configured to ascertain a legitimacy of the question, based on the at least one legitimacy standard, via:

determining presence of a question pattern within the obtained data; and

determining presence of at least one exception to the question pattern,

wherein an exception indicates that the data identified as corresponding to a question pattern should not be answered;

wherein the determined at least one exception to the question pattern comprises at least one of: sentiment, author reputation, nature of one or more responses to the question, and a number of sentences relative to the question;

computer readable program code configured to classify, based upon the ascertained legitimacy, the automatically obtained question as legitimate or not legitimate, wherein a legitimate question comprises obtained data identified as containing a question pattern and as not containing at least one exception to the question pattern, wherein a not legitimate question comprises obtained data identified as containing a question pattern and containing at least one exception to the question pattern;

computer readable program code configured to filter out the automatically obtained questions classified as not legitimate;

computer readable program code configured to harvest, for the automatically obtained questions classified as legitimate, from at least one social media conversation an answer to the question, wherein the harvesting comprises:

harvesting an answer comprising at least one rich media component taken from the group consisting of: video content; audio content; picture content; and

harvesting text associated with the at least one rich media component; and

computer readable program code configured to augment an existing question knowledge base corresponding to the target domain using the questions classified as legitimate and including the harvested answer corresponding to the question; and

computer readable program code configured to automatically provide an answer the automatically obtained question using the harvested answer.

2. A computer program product comprising:

a non-transitory computer readable storage medium having computer readable program code embodied therewith, the computer readable program code comprising:

computer readable program code configured to establish at least one legitimacy standard for filtering questions,

wherein the at least one legitimacy standard includes presence of a question pattern and at least one exception relative to the question pattern, wherein a question pattern identifies data as requesting additional information;

computer readable program code configured to automatically obtain a question from at least one social media conversation, wherein to obtain a question comprises obtaining data from at least one social media forum, extracting domain-specific communications directed to a target domain by filtering the data using domain-specific keywords, and identifying questions from the extracted domain-specific communications by determining presence of a question pattern;

computer readable program code configured to ascertain a legitimacy of the question, based on the at least one legitimacy standard, via:

determining presence of a question pattern within the obtained data; and

determining presence of at least one exception to the question pattern,

wherein an exception indicates that the data identified as corresponding to a question pattern should not be answered;

wherein the determined at least one exception to the question pattern comprises at least one of: sentiment, author reputation, nature of one or more responses to the question, and a number of sentences relative to the question;

computer readable program code configured to classify, based upon the ascertained legitimacy, the automatically obtained question as legitimate or not legitimate, wherein a legitimate question comprises obtained data identified as containing a question pattern and as not containing at least one exception to the question pattern, wherein a not legitimate question comprises obtained data identified as containing a question pattern and containing at least one exception to the question pattern:,

computer readable program code configured to filter out the automatically obtained questions classified as not legitimate;

computer readable program code configured to harvest, for the automatically obtained questions classified as legitimate, from at least one social media conversation an answer to the question, wherein the harvesting comprises:

harvesting an answer comprising at least one rich media component taken from the group consisting of: video content; audio content; picture content; and

harvesting text associated with the at least one rich media component;

and

computer readable program code configured to augment an existing question knowledge base corresponding to the target domain using the questions classified as legitimate and including the harvested answer corresponding to the question; and

computer readable program code configured to automatically provide an answer the automatically obtained question using the harvested answer.

3. The computer program product according to claim 2 , wherein said computer readable program code is further configured to ascertain whether the question is a duplicate of a previously obtained question, and thereupon consulting an answer to the previously obtained question.

4. The computer program product according to claim 2 , wherein said computer readable program code is further configured to ascertain whether the question is similar to a previously obtained question, and thereupon consulting an answer to the previously obtained question.

5. The computer program product according to claim 2 , wherein said computer readable program code is further configured to apply media processing to the at least one rich media component.

6. The computer program product according to claim 5 , wherein:

the at least one rich media component comprises audio content; and

the media processing comprises speech recognition.

7. The computer program product according to claim 2 , wherein said computer readable program code is further configured to:

rank answers to the question; and

store the ranked answers in an enhanced knowledge database.

8. The computer program product according to claim 2 , wherein said computer readable program code is configured to obtain the question from social media.

9. The apparatus according to claim 1 , wherein the metadata include one or more of: date-of-posting, creator, location, and tags.

10. The computer program product according to claim 2 , wherein the metadata include one or more of: date-of-posting, creator, location, and tags.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Mar 5, 2012
From: AJMERA, JITENDRA; GODBOLE, SHANTANU RAVINDRA; LAKKARAJU, HIMABINDU; RODEN, BERNARD ANDREW; VERMA, ASHISH
To: INTERNATIONAL BUSINESS MACHINES CORPORATION
Reel/Frame 027804/0657 →
Continuity (1)
Related Publication 20130224713A1 · Aug 29, 2013