IP Library Granted Patent US 10,467,255
Granted Patent B2
US 10,467,255 · App. 14/982,711 · Granted Nov 5, 2019

Methods and systems for analyzing reading logs and documents thereof

Inventors: Tsung-Lin Tsai (Kaohsiung, TW); Meng-Yu Lee (New Taipei, TW); Shun-Chieh Lin (Tainan, TW)
Assignee: Industrial Technology Research Institute
G06F16/285G06F16/35G06F16/93
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 10,467,255
App. No.
14/982,711
Granted
Nov 5, 2019
Kind
B2
Abstract

Methods for analyzing reading log and documents corresponding thereof are provided, including: acquiring reading log and documents corresponding thereto, wherein the reading log at least includes reading-related information about the documents within a predetermined period of time, selecting interesting document sets from the documents according to the reading log in each time segment, performing a document content pre-processing on the interesting document sets to determine keyword sets corresponding thereto for each time segment according to the interesting document sets, performing cluster calculation on the keyword sets to obtain topics and calculating cohesion of each topic, deleting topics with insufficient cohesion to obtain multiple high-relevance topics and classifying each high-relevance topic into one of predetermined topic classes according to the respective keyword sets of the high-relevance topics, obtaining reading statistics for each topic class and calculating multiple degrees of interest for each topic class during each time segment.

Claims (50)

1. A method for analyzing reading logs and documents corresponding thereto, comprising:

acquiring reading logs related to webpages and documents corresponding thereto, wherein the reading logs at least includes reading-related information about the documents within a predetermined period of time and the reading-related information at least includes an interesting reading time and a number of interesting readings;

selecting a plurality of interesting document sets from the documents in each time segment of the predetermined period of time according to the interesting reading times and the number of interesting readings of the documents in the reading logs, each of the interesting document sets corresponding to one of the time segments of the predetermined period of time;

performing a document content pre-processing on the interesting document sets to determine keyword sets corresponding to the interesting document sets;

performing a cluster calculation on the keyword sets to obtain topics and calculating cohesion of each topic;

deleting topics with insufficient cohesion among the topics obtained to obtain a plurality of high-relevance topics and classifying each high-relevance topic into one of a plurality of predetermined topic classes by comparing the respective keyword sets of the high-relevance topics with a plurality of keyword sets of the predetermined topic classes;

obtaining reading statistics for documents of each predetermined topic class and calculating a plurality of degrees of interest for documents of each predetermined topic class during each time segment; and

determining a reading trend on each predetermined topic class according to changes in the degrees of interest,

wherein the document content pre-processing step further comprises the steps of performing the following steps on each document of the interesting document sets:

obtaining a plurality of keywords;

paragraphing the document and calculating a frequency at which the keywords appear in each paragraph to calculate a plurality of importance-weightings corresponding to all of the paragraphs and determining at least one key paragraph according to the importance-weightings; and

generating the set of keywords for the document based on the keywords within the at least one key paragraph.

2. The method as claimed in claim 1 , wherein the step of selecting the interesting document sets further comprises:

filtering out uninterested reading-related information among the reading-related information about the documents to obtain filtered reading-related information;

calculating the interesting reading time and the number of interesting readings for each document based on the filtered reading-related information; and

determining whether each document belongs to the interesting document sets based on the interesting reading time and the number of interesting readings of the document;

wherein the document is classified to the interesting document sets when the interesting reading time of the document has exceeded a time threshold value and the number of interesting readings of the document has exceeded a frequency threshold value.

3. The method as claimed in claim 1 , wherein the step of classifying each high-relevance topic into one of the predetermined topic classes further comprises:

when a degree of similarity of keyword sets between a first high-relevance topic of the high-relevance topics and a first predetermined topic class of the predetermined topic classes has exceeded a predetermined threshold degree, classifying the first high-relevance topic corresponding to the keyword set being compared into the first predetermined topic class.

4. The method as claimed in claim 3 , further comprising:

automatically updating the keyword set of the first predetermined topic class using the respective keyword set of the first high-relevance topic after classifying the first high-relevance topic into the first predetermined topic classes.

5. The method as claimed in claim 3 , further comprising:

after classifying the first high-relevance topic into the first predetermined topic class, comparing the degree of similarity of keyword sets between the first high-relevance topic and a keyword set of a first topic within the first predetermined topic class; and

relating the first high-relevance topic to the first topic when the degree of similarity of keyword sets between the first high-relevance topic and the first topic has exceeded a predetermined threshold degree.

6. The method as claimed in claim 1 , wherein the step of determining the reading trend on each predetermined topic class according to changes in the degrees of interest further comprises:

gathering a total number of readings for each predetermined topic class in each time segment;

determining a degree of interest for each predetermined topic class in each time segment according to the total numbers of readings for all of the predetermined topic classes; and

determining the reading trend on each predetermined topic class based on the changes in the degrees of interest for each predetermined topic class.

7. The method as claimed in claim 6 , wherein the reading trend of each predetermined topic class comprises at least one of: the trend of going from being interested to being uninterested in the predetermined topic class, the trend of staying interested in the predetermined topic class, and the trend of going from being uninterested to being interested in the predetermined topic class.

8. The method as claimed in claim 7 , further comprising:

providing a user interface to graphically show a result of the reading trend determined in the predetermined period of time;

wherein the reading trend indicates a trend of changing in document interest for each predetermined topic class.

9. A system, implemented by a processor, for analyzing reading logs and documents corresponding thereto, comprising:

a reading log extractor, acquiring reading logs related to webpages and documents corresponding thereto, wherein the reading logs at least includes reading-related information about the documents within a predetermined period of time and the reading-related information at least includes an interesting reading time and a number of interesting readings;

an interesting document filter coupled to the reading log extractor, selecting a plurality of interesting document sets from the documents in each time segment of the predetermined period of time according to the interesting reading times and the number of interesting readings of the documents in the reading logs, each of the interesting document sets corresponding to one of the time segments of the predetermined period of time;

a document pre-processor coupled to the interesting document filter, performing a document content pre-processing on the interesting document sets to determine keyword sets corresponding to the interesting document sets;

a topic cluster generator coupled to the document pre-processor, performing a cluster calculation on the keyword sets to obtain topics, calculating cohesion of each topic and deleting topics with insufficient cohesion among the topics obtained to obtain a plurality of high-relevance topics;

a topic classifier and combiner coupled to the topic cluster generator, classifying each high-relevance topic into one of a plurality of predetermined topic classes by comparing the respective keyword sets of the high-relevance topics with a plurality of keyword sets of the predetermined topic classes;

a degree of interest normalizer coupled to the topic classifier and combiner, obtaining reading statistics for documents of each predetermined topic class and calculating a plurality of degrees of interest for documents of each predetermined topic class during each time segment; and

a reading trend analyzer coupled to the degree of interest normalizer, determining a reading trend on each predetermined topic class according to changes in the degrees of interest,

wherein for each document of the interesting document sets, the document pre-processor further obtains a plurality of keywords, paragraphs the document and calculates a frequency at which the keywords appear in each paragraph to calculate a plurality of importance-weightings corresponding to all of the paragraphs and determines at least one key paragraph according to the importance-weightings, and generates the set of keywords for the document based on the keywords within the at least one key paragraph.

10. The system as claimed in claim 9 , wherein the interesting document filter further filters out uninterested reading-related information among the reading-related information about the documents to obtain filtered reading-related information, calculates the interesting reading time and the number of interesting readings for each document based on the filtered reading-related information and determines whether each document belongs to the interesting document sets based on the interesting reading time and the number of interesting readings of the document, wherein a document is classified to the interesting document sets when the interesting reading time of the document has exceeded a time threshold value and the number of interesting readings of the document has exceeded a frequency threshold value.

11. The system as claimed in claim 9 , wherein the topic classifier and combiner further classifies a first high-relevance topic of the high-relevance topics corresponding to the keyword set being compared into a first predetermined topic class of the predetermined topic classes when a degree of similarity of keyword sets between the first high-relevance topic and the first predetermined topic class has exceeded a predetermined threshold degree.

12. The system as claimed in claim 11 , wherein the topic classifier and combiner further automatically updates the keyword set of the first predetermined topic class using the respective keyword set of the first high-relevance topic after classifying the first high-relevance topic into the first predetermined topic classes.

13. The system as claimed in claim 11 , wherein the topic classifier and combiner further computes the degree of similarity of keyword sets between the first high-relevance topic and a keyword set of a first topic within the first predetermined topic class after classifying the first high-relevance topic into the first predetermined topic class and relates the first high-relevance topic to the first topic when the degree of similarity of keyword sets between the first high-relevance topic and the first topic has exceeded a predetermined threshold degree.

14. The system as claimed in claim 11 , wherein the degree of interest normalizer further gathers a total number of readings for each predetermined topic class in each time segment, determines a degree of interest for each predetermined topic class in each time segment according to the total numbers of readings for all of the predetermined topic classes, and determines the reading trend on each predetermined topic class based on the changes in the degrees of interest for each predetermined topic class.

15. The system as claimed in claim 14 , wherein the reading trend of each predetermined topic class comprises at least one of: the trend of going from being interested to being uninterested in the predetermined topic class, the trend of staying interested in the predetermined topic class, and the trend of going from being uninterested to being interested in the predetermined topic class.

16. The system as claimed in claim 9 , wherein the reading trend analyzer further provides a user interface to graphically show a result of the reading trend determined in the predetermined period of time, wherein the reading trend indicates a trend of changing in document interest for each predetermined topic class.

17. The method as claimed in claim 1 , wherein the reading trend for each predetermined topic class is one of going from being interested to being uninterested in documents of the predetermined topic class, staying interested in documents of the predetermined topic class, and going from being uninterested to being interested in documents of the predetermined topic class.

18. The system as claimed in claim 9 , wherein the reading trend for each predetermined topic class is one of going from being interested to being uninterested in documents of the predetermined topic class, staying interested in documents of the predetermined topic class, and going from being uninterested to being interested in documents of the predetermined topic class.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jan 12, 2016
From: TSAI, TSUNG-LIN; LEE, MENG-YU; LIN, SHUN-CHIEH
To: INDUSTRIAL TECHNOLOGY RESEARCH INSTITUTE
Reel/Frame 037464/0347 →
Priority Claims (1)
TW 104141664 A · Dec 11, 2015 · national
Continuity (1)
Related Publication 20170169096A1 · Jun 15, 2017