System, method and apparatus for automatic topic relevant content filtering from social media text streams using weak supervision
View Patent ↗Presented are a system, method, and apparatus for automatic topic relevant content filtering from social media text streams using weak supervision. A computing device utilizes heuristic rules allowing topic filtering and a data stream data chunk identifier. A plurality of messages are transmitted as streaming message data from a social media network in real-time. The messages are split into a plurality of data stream data chunks according to the data stream data chunk identifier. A rule-based labeled data set L 0 is built from one or more data instances in the first stream data chunk. An initial classifier is built based upon features of L 0 . The initial classifier is applied to a next data stream data chunk to build a labeled data set L 1 . A subset of representative instances S 1 is selected from labeled data set L 1 . A first representative classifier C 1 is constructed from representative instance S 1 .
1. A method to filter relevant content from streaming message data streamed across social media networking software to a computing device utilizing a weak supervision strategy, the method comprising the steps of:
utilizing by the computing device one or more heuristic rules for use with a weakly supervised data stream filter utilizing the weak supervision strategy, the one or more heuristic rules allowing topic filtering;
utilizing by the computing device a data stream data chunk identifier;
receiving continuously by the computing device a plurality of messages as streaming message data in a data stream from the social media networking software in real time;
splitting, utilizing the data stream data chunk identifier, the plurality of messages in the data stream into a plurality of stream data chunks D i based on a timestamp or volume of data in the data stream;
loading one or more of the stream data chunks into memory associated with the computing device;
receiving by the computing device a topic for filtering;
determining by the computing device a stream data chunk D i where the topic is first discussed in the data stream;
setting D 0 equal to the determined D i where the topic is first discussed to identify the determined stream data chunk D i as first stream data chunk D 0 ;
building automatically by the computing device utilizing at least one of a plurality of heuristic rules a rule-based labeled data set L 0 from one or more data instances in the first stream data chunk D 0 ;
constructing an initial classifier C 0 based upon one or more features of the labeled data set L 0 ;
applying the initial classifier C 0 to stream data chunk D 1 to build automatically labeled data set L 1 ;
selecting a subset of representative instances S 1 from labeled data set L 1 ;
constructing a first representative classifier C 1 from the representative instances S 1 ;
applying the first representative classifier C 1 in combination with the initial classifier C 0 using one or more combination strategies to stream data chunk D 2 to build automatically a labeled data set L 2 ;
selecting a subset of representative instances S 2 from labeled data set L 2 ; and
constructing a second representative classifier C 2 from the representative instances S 2 .
2. The method of claim 1 , wherein applying the initial classifier C 0 or the first representative classifier C 1 comprises performing a determination of whether one or more words in the plurality of messages are relevant or not-relevant.
3. The method of claim 2 , wherein each relevant message is labeled as 1 and each nonrelevant message is labeled as 0.
4. The method of claim 2 , wherein the computing device returns each relevant message to a user.
5. The method of claim 1 , wherein the timestamp is based selectively on one or more of the following timeframes: a minute, an hour, a day, a week, and a month.
6. The method of claim 1 , wherein when constructing the initial classifier C 0 based on one or more features of the labeled data set L 0 , the one or more features are specially tailored to the social media networking software.
7. The method of claim 1 , wherein the computing device builds an ensemble classifier E i that ensembles the initial classifier C 0 and a latest built representative classifier C i R .
8. The method of claim 1 , wherein the computing device selects a subset of representative instances S i from each built labeled data set L i .
9. The method of claim 8 , where i=current+1.
10. A system to filter relevant content from streaming message data streamed across social media networking software to a computing device utilizing a weak supervision strategy, the system comprising a computing device which performs the following:
utilize one or more heuristic rules for use with a weakly supervised data stream filter utilizing the weak supervision strategy, the one or more heuristic rules allowing topic filtering;
utilize a data stream data chunk identifier;
receipt continuously of a plurality of messages as streaming message data in a data stream from the social media networking software in real-time;
split, utilizing the data stream data chunk identifier, the plurality of messages in the data stream into a plurality of stream data chunks D i based on a timestamp or volume of data in the data stream;
load one or more of the stream data chunks into memory associated with the computing device;
receive by the computing device a topic for filtering;
determine by the computing device a stream data chunk D i where the topic is first discussed in the data stream;
set D 0 equal to the determined D i where the topic is first discussed to identify the determined stream data chunk D i as first stream data chunk D 0 ;
build automatically utilizing at least one of a plurality of heuristic rules a rule-based labeled data set L 0 from one or more data instances in the first stream data chunk D 0 ;
construct an initial classifier C 0 based upon one or more features of the labeled data set L 0 ;
apply the initial classifier C 0 to stream data chunk D 1 to build automatically labeled data set L 1 ;
select a subset of representative instances S 1 from labeled data set L 1 ;
construct a first representative classifier C 1 from the representative instances S 1 ;
apply the first representative classifier C 1 in combination with the initial classifier C 0 using one or more combination strategies to stream data chunk to build automatically a labeled data set L 2 ;
select a subset of representative instances S 2 from labeled data set L 2 ; and
construct a second representative classifier C 2 the from representative instances S 2 .
11. The system of claim 10 , wherein the application of the initial classifier C 0 or the first representative classifier C 1 comprises performing a determination of whether one or more words in the plurality of messages are relevant or not-relevant.
12. The system of claim 11 , wherein each relevant message is labeled as 1 and each nonrelevant message is labeled as 0.
13. The system of claim 11 , wherein the computing device returns each relevant message to a user.
14. The system of claim 10 , wherein the timestamp is based selectively on one or more of the following timeframes: a minute, an hour, a day, a week, and a month.
15. The system of claim 10 , wherein when constructing the initial classifier C 0 based on one or more features of the labeled data set L 0 , the one or more features are specially tailored to the social media networking software.
16. The system of claim 10 , wherein building an ensemble classifier E i that ensembles the initial classifier C 0 and a latest built representative classifier C i R .
17. The system of claim 10 , wherein the computing device selects a subset of representative instances S i from each built labeled data set L i .
18. The system of claim 17 , where i=current+1.