IP Library Granted Patent US 10,467,537
Granted Patent B2
US 10,467,537 · App. 14/877,970 · Granted Nov 5, 2019

System, method and apparatus for automatic topic relevant content filtering from social media text streams using weak supervision

View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 10,467,537
App. No.
14/877,970
Granted
Nov 5, 2019
Kind
B2
Abstract

Presented are a system, method, and apparatus for automatic topic relevant content filtering from social media text streams using weak supervision. A computing device utilizes heuristic rules allowing topic filtering and a data stream data chunk identifier. A plurality of messages are transmitted as streaming message data from a social media network in real-time. The messages are split into a plurality of data stream data chunks according to the data stream data chunk identifier. A rule-based labeled data set L 0 is built from one or more data instances in the first stream data chunk. An initial classifier is built based upon features of L 0 . The initial classifier is applied to a next data stream data chunk to build a labeled data set L 1 . A subset of representative instances S 1 is selected from labeled data set L 1 . A first representative classifier C 1 is constructed from representative instance S 1 .

Claims (50)

1. A method to filter relevant content from streaming message data streamed across social media networking software to a computing device utilizing a weak supervision strategy, the method comprising the steps of:

utilizing by the computing device one or more heuristic rules for use with a weakly supervised data stream filter utilizing the weak supervision strategy, the one or more heuristic rules allowing topic filtering;

utilizing by the computing device a data stream data chunk identifier;

receiving continuously by the computing device a plurality of messages as streaming message data in a data stream from the social media networking software in real time;

splitting, utilizing the data stream data chunk identifier, the plurality of messages in the data stream into a plurality of stream data chunks D i based on a timestamp or volume of data in the data stream;

loading one or more of the stream data chunks into memory associated with the computing device;

receiving by the computing device a topic for filtering;

determining by the computing device a stream data chunk D i where the topic is first discussed in the data stream;

setting D 0 equal to the determined D i where the topic is first discussed to identify the determined stream data chunk D i as first stream data chunk D 0 ;

building automatically by the computing device utilizing at least one of a plurality of heuristic rules a rule-based labeled data set L 0 from one or more data instances in the first stream data chunk D 0 ;

constructing an initial classifier C 0 based upon one or more features of the labeled data set L 0 ;

applying the initial classifier C 0 to stream data chunk D 1 to build automatically labeled data set L 1 ;

selecting a subset of representative instances S 1 from labeled data set L 1 ;

constructing a first representative classifier C 1 from the representative instances S 1 ;

applying the first representative classifier C 1 in combination with the initial classifier C 0 using one or more combination strategies to stream data chunk D 2 to build automatically a labeled data set L 2 ;

selecting a subset of representative instances S 2 from labeled data set L 2 ; and

constructing a second representative classifier C 2 from the representative instances S 2 .

2. The method of claim 1 , wherein applying the initial classifier C 0 or the first representative classifier C 1 comprises performing a determination of whether one or more words in the plurality of messages are relevant or not-relevant.

3. The method of claim 2 , wherein each relevant message is labeled as 1 and each nonrelevant message is labeled as 0.

4. The method of claim 2 , wherein the computing device returns each relevant message to a user.

5. The method of claim 1 , wherein the timestamp is based selectively on one or more of the following timeframes: a minute, an hour, a day, a week, and a month.

6. The method of claim 1 , wherein when constructing the initial classifier C 0 based on one or more features of the labeled data set L 0 , the one or more features are specially tailored to the social media networking software.

7. The method of claim 1 , wherein the computing device builds an ensemble classifier E i that ensembles the initial classifier C 0 and a latest built representative classifier C i R .

8. The method of claim 1 , wherein the computing device selects a subset of representative instances S i from each built labeled data set L i .

9. The method of claim 8 , where i=current+1.

10. A system to filter relevant content from streaming message data streamed across social media networking software to a computing device utilizing a weak supervision strategy, the system comprising a computing device which performs the following:

utilize one or more heuristic rules for use with a weakly supervised data stream filter utilizing the weak supervision strategy, the one or more heuristic rules allowing topic filtering;

utilize a data stream data chunk identifier;

receipt continuously of a plurality of messages as streaming message data in a data stream from the social media networking software in real-time;

split, utilizing the data stream data chunk identifier, the plurality of messages in the data stream into a plurality of stream data chunks D i based on a timestamp or volume of data in the data stream;

load one or more of the stream data chunks into memory associated with the computing device;

receive by the computing device a topic for filtering;

determine by the computing device a stream data chunk D i where the topic is first discussed in the data stream;

set D 0 equal to the determined D i where the topic is first discussed to identify the determined stream data chunk D i as first stream data chunk D 0 ;

build automatically utilizing at least one of a plurality of heuristic rules a rule-based labeled data set L 0 from one or more data instances in the first stream data chunk D 0 ;

construct an initial classifier C 0 based upon one or more features of the labeled data set L 0 ;

apply the initial classifier C 0 to stream data chunk D 1 to build automatically labeled data set L 1 ;

select a subset of representative instances S 1 from labeled data set L 1 ;

construct a first representative classifier C 1 from the representative instances S 1 ;

apply the first representative classifier C 1 in combination with the initial classifier C 0 using one or more combination strategies to stream data chunk to build automatically a labeled data set L 2 ;

select a subset of representative instances S 2 from labeled data set L 2 ; and

construct a second representative classifier C 2 the from representative instances S 2 .

11. The system of claim 10 , wherein the application of the initial classifier C 0 or the first representative classifier C 1 comprises performing a determination of whether one or more words in the plurality of messages are relevant or not-relevant.

12. The system of claim 11 , wherein each relevant message is labeled as 1 and each nonrelevant message is labeled as 0.

13. The system of claim 11 , wherein the computing device returns each relevant message to a user.

14. The system of claim 10 , wherein the timestamp is based selectively on one or more of the following timeframes: a minute, an hour, a day, a week, and a month.

15. The system of claim 10 , wherein when constructing the initial classifier C 0 based on one or more features of the labeled data set L 0 , the one or more features are specially tailored to the social media networking software.

16. The system of claim 10 , wherein building an ensemble classifier E i that ensembles the initial classifier C 0 and a latest built representative classifier C i R .

17. The system of claim 10 , wherein the computing device selects a subset of representative instances S i from each built labeled data set L i .

18. The system of claim 17 , where i=current+1.

Assignments (6)
SECURITY INTEREST Recorded Oct 19, 2021
From: CONDUENT BUSINESS SERVICES, LLC
To: U.S. BANK, NATIONAL ASSOCIATION
Reel/Frame 057969/0445 →
SECURITY INTEREST Recorded Oct 19, 2021
From: CONDUENT BUSINESS SERVICES, LLC
To: BANK OF AMERICA, N.A.
Reel/Frame 057970/0001 →
RELEASE OF SECURITY INTEREST Recorded Oct 18, 2021
From: JPMORGAN CHASE BANK, N.A.
To: CONDUENT BUSINESS SERVICES, LLC; CONDUENT STATE & LOCAL SOLUTIONS, INC.; CONDUENT TRANSPORT SOLUTIONS, INC.; ADVECTIS, INC.; CONDUENT COMMERCIAL SOLUTIONS, LLC; CONDUENT BUSINESS SOLUTIONS, LLC; CONDUENT CASUALTY CLAIMS SOLUTIONS, LLC; CONDUENT HEALTH ASSESSMENTS, LLC
Reel/Frame 057969/0180 →
SECURITY AGREEMENT Recorded Mar 19, 2020
From: CONDUENT BUSINESS SERVICES, LLC
To: JPMORGAN CHASE BANK, N.A.
Reel/Frame 052189/0698 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jan 28, 2017
From: XEROX CORPORATION
To: CONDUENT BUSINESS SERVICES, LLC
Reel/Frame 041542/0022 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Oct 8, 2015
From: AGARWAL, ARVIND; DONG, CAILING
To: XEROX CORPORATION
Reel/Frame 036753/0831 →