IP Library Granted Patent US 9,922,032
Granted Patent B2
US 9,922,032 · App. 14/558,224 · Granted Mar 20, 2018

Featured co-occurrence knowledge base from a corpus of documents

View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 9,922,032
App. No.
14/558,224
Granted
Mar 20, 2018
Kind
B2
Abstract

A system for building a knowledge base of co-occurring features extracted from a document corpus is disclosed. The method includes a plurality of feature extraction software modules that may extract different features from each document in the corpus. The system may include a knowledge base aggregator module that may keep count of the co-occurrences of features in the different documents of a corpus and determine appropriate co-occurrences to store in a knowledge base.

Claims (33)

1. A method comprising:

crawling, via an entity extraction computer, a corpus of a plurality of electronic documents;

extracting, via the entity extraction computer, a plurality of features from each of the crawled electronic documents in the corpus;

aggregating, via a knowledge base aggregator computer, instances of co-occurrence of two or more of the plurality of features across the crawled electronic documents to determine a count of the instances of co-occurrence across the plurality of electronic documents,

wherein the count of the instances of co-occurrence of two or more features are determined by adding up number of co-occurrences of the two or more features co-occurred in each of the plurality of electronic documents,

wherein the knowledge base aggregator computer includes a main memory hosting a feature co-occurrence in-memory database, wherein the entity extraction computer is distinct from the knowledge base aggregator computer, wherein the feature co-occurrence in-memory database comprises one or more entries of the plurality of features, wherein each entry of the one or more entries contains semantically-related features of the plurality of features, and wherein the co-occurrence is an instance of a feature of the plurality of features identified by an entry of the one or more entries in the corpus of the electronic documents in the co-occurrence database, and wherein the co-occurrence corresponds to determination of related entities upon the occurrence of the two or more of the plurality of features that are related to each other across the crawled electronic documents in no defined order; and

adding, via the knowledge base aggregator computer, the instances of the co-occurrence of the two or more features to the feature co-occurrence database based on the count of the instances of co-occurrence across the plurality of electronic documents exceeding a predetermined threshold.

2. The method of claim 1 wherein each of the plurality of features is selected from the group consisting of a person, a location, an organization name, a topic, an event, and a fact.

3. The method of claim 1 further comprising adding, via the knowledge base aggregator computer, the instances of the co-occurrence of the two or more features to the feature co-occurrence database along with metadata pertaining to the instance.

4. The method of claim 3 wherein the metadata is selected from the group consisting of feature type, document identifier, document corpus identifier, distance in text between co-occurring features, and a confidence score.

5. The method of claim 3 where in the metadata includes a confidence score indicative whether the two or more of the plurality of features associated with a respective instance of co-occurrence are unique.

6. The method of claim 5 wherein the confidence score is calculated using a plurality of parameters including a number of co-occurrences in a single electronic document, a number of co-occurrences in the corpus of electronic documents, a size of the corpus of electronic documents, a number of co-occurrences in different corpora of electronic documents, distance in text from co-occurring features, and an indicator of human verification.

7. A system comprising:

an entity extraction computer configured to crawl a corpus of a plurality of electronic documents and extract a plurality of features from each of the crawled electronic documents in the corpus;

a feature co-occurrence in-memory database comprises one or more entries of the plurality of features, wherein each entry of the one or more entries contains semantically related features of the plurality of features; and

a knowledge base aggregator computer including a main memory hosting the feature co-occurrence in-memory database, wherein the entity extraction computer is distinct from the knowledge base aggregator computer, wherein the knowledge base aggregator computer is configured to aggregate instances of co-occurrence two or more of the plurality of features across the crawled electronic documents to determine a count of the instances of co-occurrence across the plurality of electronic documents, wherein the count of the instances of co-occurrence of two or more features are determined by adding up number of co-occurrences of the two or more features co-occurred in each of the plurality of electronic documents, wherein the co-occurrence is an instance of a feature of the plurality of features identified by an entry of the one or more entries in the corpus of the electronic documents in the co-occurrence database, and wherein the co-occurrence corresponds to determination of related entities upon the occurrence of the two or more of the plurality of features that are related to each other across the crawled electronic documents in no defined order, and add the instances of the co-occurrence of the two or more features to a feature co-occurrence database based on the count of the instances of co-occurrence across the plurality of electronic documents exceeding a predetermined threshold.

8. The system of claim 7 wherein each of the plurality of features is selected from the group consisting of a person, a location, an organization name, a topic, an event, and a fact.

9. The system of claim 7 wherein the knowledge aggregator computer is further configured to add the instances of the co-occurrence of the two or more features to the feature co-occurrence database along with metadata pertaining to the instance.

10. The system of claim 9 wherein the metadata is selected from the group consisting of feature type, document identifier, document corpus identifier, distance in text between co-occurring features, and a confidence score.

11. The system of claim 9 where in the metadata includes a confidence score indicative whether the two or more of the plurality of features associated with a respective instance of co-occurrence are unique.

12. The system of claim 11 wherein the knowledge base aggregator computer is further configured to calculate the confidence score using a plurality of parameters including a number of co-occurrences in a single electronic document, a number of co-occurrences in the corpus of electronic documents, a size of the corpus of electronic documents, a number of co-occurrences in different corpora of electronic documents, distance in text from co-occurring features, and an indicator of human verification.

13. A non-transitory computer readable medium having stored thereon computer executable instructions comprising:

crawling, via an entity extraction computer, a corpus of a plurality of electronic documents;

extracting, via the entity extraction computer, a plurality of features from each of the crawled electronic documents in the corpus;

aggregating, via a base aggregator computer, instances of co-occurrence of two or more of the plurality of features across the crawled documents to determine a count of the instances of co-occurrence across the plurality of electronic documents,

wherein the count of the instances of co-occurrence of two or more features are determined by adding up number of co-occurrences of the two or more features co-occurred in each of the plurality of electronic documents,

wherein the base aggregator computer includes a main memory hosting a feature co-occurrence in-memory database, wherein the entity extraction computer is distinct from the knowledge base aggregator computer, wherein the feature co-occurrence in-memory database comprises one or more entries of the plurality of features, wherein each entry of the one or more entries contains semantically related features of the plurality of features, and wherein the co-occurrence is an instance of a feature of the plurality of features identified by an entry of the one or more entries in the corpus of the electronic documents in the co-occurrence database, and wherein the co-occurrence corresponds to determination of related entities upon the occurrence of the two or more of the plurality of features that are related to each other across the crawled electronic documents in no defined order; and

adding, via the base aggregator computer, the instance of the co-occurrence of the two or more features to the feature co-occurrence database based on the count of the instances of co-occurrence across the plurality of electronic documents exceeding a predetermined threshold.

14. The computer readable medium of claim 13 wherein each of the plurality of features is selected from the group consisting of a person, a location, an organization name, a topic, an event, and a fact.

15. The computer readable medium of claim 13 wherein the instructions further comprise adding, via the entity extraction computer, the instances of the co-occurrence of the two or more features to the feature co-occurrence database along with metadata pertaining to the instance.

16. The computer readable medium of claim 15 wherein the metadata is selected from the group consisting of feature type, document identifier, document corpus identifier, distance in text between co-occurring features, and a confidence score.

17. The computer readable medium of claim 15 where in the metadata includes a confidence score indicative whether the two or more of the plurality of features associated with a respective instance of co-occurrence are unique.

18. The computer readable medium of claim 17 wherein the instructions further comprise calculating, via the knowledge base aggregator computer, the confidence score using a plurality of parameters including a number of co-occurrences in a single electronic document, a number of co-occurrences in the corpus of electronic documents, a size of the corpus of electronic documents, a number of co-occurrences in different corpora of electronic documents, distance in text from co-occurring features, and an indicator of human verification.

Assignments (2)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Sep 30, 2021
From: QBASE, LLC
To: FINCH COMPUTING, LLC
Reel/Frame 057663/0426 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Dec 5, 2014
From: LIGHTNER, SCOTT; DAVE, RAKESH; BODDHU, SANJAY
To: QBASE, LLC
Reel/Frame 034390/0524 →