IP Library Granted Patent US 10,599,700
Granted Patent B2
US 10,599,700 · App. 15/244,226 · Granted Mar 24, 2020

Systems and methods for narrative detection and frame detection using generalized concepts and relations

Inventors: Hasan Davulcu (Phoenix, AZ); Steven Corman (Chandler, AZ)
Assignee: Arizona Board of Regents on behalf of Arizona State University
G06F16/355G06F17/2785G06F17/2795
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 10,599,700
App. No.
15/244,226
Filed
Aug 23, 2016
Granted
Mar 24, 2020
Kind
B2
Art Unit
2154
USPC
707/737
Abstract

Co-clustering based on generalized conceptual relationships can automatically detect story forms incorporating archetypes/targets and actions. Co-clustering can help in identifying similarities that exist in low-dimensional sub-spaces of sparse data such as textual paragraphs. Through co-clustering, the clusters themselves and their characteristic features are identifiable which can be useful in describing and summarizing their contents. The residual error of factorization with concept-based features is significantly lower than the error with prior keyword-based features. Qualitative evaluations also suggest that concept-based features yield more coherent, distinctive and interesting story forms compared to those produced by using prior keyword-based features.

Claims (37)

1. A method for story form detection in a corpus, the method comprising:

preprocessing, in an electronic dataset comprising a plurality of paragraphs and by a story form detection system operative on a computer optimized for story form detection, each paragraph to prepare the paragraphs for feature extraction;

generating, from the dataset and by the story form detection system, a uni-gram feature set for the plurality of paragraphs;

generating, from the dataset and by the story form detection system, a bi-gram feature set for the plurality of paragraphs;

generating, from the dataset and by the story form detection system, a generalized concepts/relations feature set for the plurality of paragraphs;

creating, by the story form detection system, a feature matrix for the uni-gram feature set;

creating, by the story form detection system, a feature matrix for the bi-gram feature set;

creating, by the story form detection system, a binary feature matrix for the generalized concepts/relations feature set; and

co-clustering, by the story form detection system and via an algorithm, the uni-gram feature set, the bi-gram feature set, and the generalized concepts/relations feature set into a set of clusters, wherein the co-clustering comprises:

utilizing a t-distributed stochastic neighbor embedding technique to visualize block diagonal sub-structures in the feature sets, and

selecting, in the feature sets, a number of clusters between two and fourteen, the number of clusters configured to generate the highest variance ratio criterion among the possible number of clusters.

2. The method of claim 1 , wherein the co-clustering utilizes non-negative matrix factorization.

3. The method of claim 1 , wherein the preprocessing comprises determining, by the story form detection system, if a paragraph in the plurality of paragraphs comprises a story.

4. The method of claim 1 , further comprising generating, by the story form detection system, block diagonal sparsity structures for the feature sets by reordering indices for the feature sets in each row and column for each duster in the set of clusters; and displaying, on a computer system display associated with the story form detection system, the block diagonal sparsity structures.

5. The method of claim 1 , further comprising utilizing the set of clusters to generate a set of counter-narratives targeted to each duster in the set of clusters.

6. The method of claim 5 , further comprising disseminating, via at least one media channel, the set of counter narratives.

7. The method of claim 1 , wherein the co-clustering algorithm utilizes a syntactic criterion and a semantic criterion.

8. A method for story form detection in a corpus, the method comprising:

preprocessing, in an electronic dataset comprising a plurality of paragraphs and by a story form detection system operative on a computer optimized for story form detection, each paragraph to prepare the paragraphs for feature extraction;

generating, from the dataset and by the story form detection system, a uni-gram feature set for the plurality of paragraphs;

generating, from the dataset and by the story form detection system, a bi-gram feature set for the plurality of paragraphs;

generating, from the dataset and by the story form detection system, a generalized concepts/relations feature set for the plurality of paragraphs;

creating, by the story form detection system, a feature matrix for the uni-gram feature set;

creating, by the story form detection system, a feature matrix for the bi-gram feature set;

creating, by the story form detection system, a binary feature matrix for the generalized concepts/relations feature set;

co-clustering, by the story form detection system and via an algorithm, the rani-gram feature set, the bi-gram feature set, and the generalized concepts/relations feature set into a set of clusters;

generating, by the story form detection system, block diagonal sparsity structures for the feature sets by reordering indices for the feature sets in each row and column for each duster in the set of clusters; and

displaying, on a computer system display associated with the story form detection system, the block diagonal sparsity structures.

9. The method of claim 8 , wherein the co-clustering utilizes non-negative matrix factorization.

10. The method of claim 8 , wherein the preprocessing comprises determining, by the story form detection system, if a paragraph in the plurality of paragraphs comprises a story.

11. The method of claim 8 , wherein the co-clustering comprises:

utilizing a t-distributed stochastic neighbor embedding technique to visualize block diagonal sub-structures in the feature sets; and

selecting, in the feature sets, a number of clusters between two and fourteen, the number of clusters configured to generate the highest variance ratio criterion among the possible number of clusters.

12. The method of claim 8 , wherein the set of clusters comprises at least two clusters and less than 14 clusters.

13. The method of claim 8 , further comprising utilizing the set of clusters to generate a set of counter-narratives targeted to each cluster in the set of clusters.

14. The method of claim 13 , further comprising disseminating, via at least one media channel, the set of counter narratives.

15. The method of claim 8 , wherein the co-clustering algorithm utilizes a syntactic criterion and a semantic criterion.

Assignments (3)
CONFIRMATORY LICENSE Recorded May 13, 2019
From: ARIZONA STATE UNIVERSITY
To: NAVY, SECRETARY OF THE UNITED STATES OF AMERICA
Reel/Frame 049166/0467 →
CONFIRMATORY LICENSE Recorded Apr 24, 2019
From: ARIZONA STATE UNIVERSITY
To: NAVY, SECRETARY OF THE UNITED STATES OF AMERICA
Reel/Frame 048993/0028 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Aug 29, 2016
From: DAVULCU, HASAN; CORMAN, STEVEN
To: ARIZONA BOARD OF REGENTS ON BEHALF OF ARIZONA STATE UNIVERSITY
Reel/Frame 039562/0948 →
Continuity (2)
Provisional Application 62208966 · Aug 24, 2015
Related Publication 20170116204A1 · Apr 27, 2017
Cited By (1)
US 12,388,862