IP Library Granted Patent US 9,135,242
Granted Patent B1
US 9,135,242 · App. 13/832,339 · Granted Sep 15, 2015

Methods and systems for the analysis of large text corpora

Inventors: Xiaoyu Wang (Charlotte, NC); Wenwen Dou (Charlotte, NC); William Ribarsky (Charlotte, NC)
Assignee: The University of North Carolina at Charlotte
G06F17/28
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 9,135,242
App. No.
13/832,339
Granted
Sep 15, 2015
Kind
B1
Abstract

Computerized methods and systems for the analysis of textual data, including: receiving, from one or more memories at one or more processors, textual data; using the processors, formatting the textual data for analysis and applying a probabilistic topic model to the textual data to extract semantically meaningful topics that collectively describe it; using a keyword weighting module, generating a topic cloud view representing the topics as a tagcloud with each being associated with a plurality of keywords; using a topic ordering module, generating a document distribution view representing a distribution of the textual data across multiple topics; using a document entropy calculation module, generating a document scatterplot view representing how many topics are attributable to the textual data; using a temporal topic trend calculation module, generating a temporal view representing changes in the occurrence of topics over time; and displaying one or more of the views to a user.

Claims (36)

1. A computerized method for the analysis of textual data, comprising:

receiving, from one or more memories at one or more processors, textual data to be analyzed;

using the one or more processors, formatting the textual data for subsequent analysis;

using the one or more processors, applying a probabilistic topic model to the textual data to extract a set of semantically meaningful topics that collectively describe all or a portion of the textual data;

using a keyword weighting module executed on the one or more processors, generating a topic cloud view representing the topics as a tagcloud with each being associated with a plurality of keywords;

using a topic ordering module executed on the one or more processors, generating a document distribution view representing a distribution of all or a portion of the textual data across multiple topics;

using a document entropy calculation module executed on the one or more processors, generating a document scatterplot view representing how many topics are attributable to all or a portion of the textual data;

using a temporal topic trend calculation module executed on the one or more processors, generating a temporal view representing changes in the occurrence of topics over time in relation to all or a portion of the textual data; and

displaying one or more of the topic cloud view, the document distribution view, the document scatterplot view, and the temporal view to a user in the analysis of all or a portion of the textual data.

2. The computerized method of claim 1 , wherein the textual data comprises one or more of textual data derived from a plurality of documents, textual data derived from a plurality of files, textual data derived from one or more data storage repositories, and textual data derived from the Internet.

3. The computerized method of claim 1 , wherein formatting the textual data for subsequent analysis comprises one or more of stopword removal, duplicated-content removal, part-of-speech analysis, n-gram analysis of sentences to extract segments, entity extraction analysis to extract named entities, sentiment analysis of the basic sentiment of documents or paragraphs, and temporal and spatial indicator extraction.

4. The computerized method of claim 1 , wherein the probabilistic topic model generates a set of latent topics and represents each topic as a multinomial distribution over a plurality of keywords.

5. The computerized method of claim 4 , wherein the textual data is described as a probabilistic mixture of topics.

6. The computerized method of claim 1 , wherein the probabilistic topic model comprises Latent Dirichet Allocation (LDA).

7. The computerized method of claim 1 , wherein the keywords are ordered to indicate their importance to a given topic and relationship to one another.

8. The computerized method of claim 1 , wherein the keywords are highlighted to indicate their importance to multiple topics.

9. The computerized method of claim 1 , wherein topics are ordered to represent their relationships.

10. The computerized method of claim 1 , wherein the document entropy calculation module utilizes a Shannon entropy calculation.

11. A computerized system for the analysis of textual data, comprising:

one or more memories operable for storing and one or more processors operable for receiving textual data to be analyzed;

an algorithm executed on the one or more processors operable for formatting the textual data for subsequent analysis;

an algorithm executed on the one or more processors operable for applying a probabilistic topic model to the textual data to extract a set of semantically meaningful topics that collectively describe all or a portion of the textual data;

a keyword weighting module executed on the one or more processors operable for generating a topic cloud view representing the topics as a tagcloud with each being associated with a plurality of keywords;

a topic ordering module executed on the one or more processors operable for generating a document distribution view representing a distribution of all or a portion of the textual data across multiple topics;

a document entropy calculation module executed on the one or more processors operable for generating a document scatterplot view representing how many topics are attributable to all or a portion of the textual data;

a temporal topic trend calculation module executed on the one or more processors operable for generating a temporal view representing changes in the occurrence of topics over time in relation to all or a portion of the textual data; and

a display operable for displaying one or more of the topic cloud view, the document distribution view, the document scatterplot view, and the temporal view to a user in the analysis of all or a portion of the textual data.

12. The computerized system of claim 11 , wherein the textual data comprises one or more of textual data derived from a plurality of documents, textual data derived from a plurality of files, textual data derived from one or more data storage repositories, and textual data derived from the Internet.

13. The computerized system of claim 11 , wherein formatting the textual data for subsequent analysis comprises one or more of word binning, geo-spatial binning, temporal information binning, entity-level content binning, document similarity comparison, document probability distribution, entropy analysis, document segmentation, word frequency detection, data coordination, GUI design, direct visual manipulation, and data-visual-element transformation and correlation.

14. The computerized system of claim 11 , wherein the probabilistic topic model generates a set of latent topics and represents each topic as a multinomial distribution over a plurality of keywords.

15. The computerized system of claim 14 , wherein the textual data is described as a probabilistic mixture of topics.

16. The computerized system of claim 11 , wherein the probabilistic topic model comprises Latent Dirichet Allocation (LDA).

17. The computerized system of claim 11 , wherein the keywords are ordered to indicate their importance to a given topic and relationship to one another.

18. The computerized system of claim 11 , wherein the keywords are highlighted to indicate their importance to multiple topics.

19. The computerized system of claim 11 , wherein topics are ordered to represent their relationships.

20. The computerized system of claim 11 , wherein the document entropy calculation module utilizes a Shannon entropy calculation.

Assignments (4)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jul 2, 2024
From: STRATIFYD, INC.
To: STRATIFYD SOFTWARE, LLC
Reel/Frame 067900/0572 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Dec 28, 2023
From: DW CAPITAL, LLC
To: STRATIFYD, INC.
Reel/Frame 065967/0537 →
CONFIRMATORY LICENSE Recorded Jan 12, 2018
From: UNIVERSITY OF NORTH CAROLINA, CHARLOTTE
To: NATIONAL SCIENCE FOUNDATION
Reel/Frame 044608/0549 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Mar 15, 2013
From: WANG, XIAOYU; DOU, WENWEN; RIBARSKY, WILLIAM
To: THE UNIVERSITY OF NORTH CAROLINA AT CHARLOTTE
Reel/Frame 030009/0277 →
Continuity (2)
Continuation In Part 13645776 · Oct 5, 2012
Provisional Application 61545331 · Oct 10, 2011