IP Library › Granted Patent US 11,651,159
Granted Patent B2
US 11,651,159 · App. 16/289,708 · Granted May 16, 2023

Semi-supervised system to mine document corpus on industry specific taxonomies

Inventors: Pietro Mazzoleni (New York, NY); Wesley M Gifford (Ridgefield, CT); Elham Khabiri (Briarcliff Manor, NY)
Assignee: International Business Machines Corporation
G06F40/30G06F16/9024G06F16/9035G06F16/9538G06F18/2193G06N5/022G06N20/00
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,651,159
App. No.
16/289,708
Granted
May 16, 2023
Kind
B2
Abstract

A method, computer system, and a computer program product for generating a custom corpus is provided. The present invention may include generating a domain graph. The present invention may also include gathering seed data based on the generated domain graph. The present invention may then include identifying domain related data based on the gathered seed data. The present invention may further include querying the domain related data. The present invention may also include creating word embeddings for the domain related data. The present invention may then include evaluating the domain related data.

Claims (46)

1. A method for generating a custom corpus, the method comprising:

generating a domain graph comprising nodes representing terms associated with an industry taxonomy;

gathering seed data based on the generated domain graph by querying a knowledgebase using the nodes in the domain graph;

identifying domain related data based on the gathered seed data, wherein identifying the domain related data further comprises establishing a ground truth for the terms such that a true meaning of a term is determined according to a context of a respective domain associated with the industry taxonomy and further comprises identifying relevant and irrelevant data to a domain associated with the industry taxonomy based on the ground truth;

querying the domain related data and identifying other terms and candidate data to add to the industry taxonomy based on the query and relevancy to respective domains;

creating word embeddings for the terms associated with the domain related data;

evaluating the domain related data in one or more loops based on a plurality of evaluation criteria, wherein the plurality of evaluation criteria comprises determining and evaluating word coverage, context, and accuracy of the created word embeddings, and wherein evaluating the domain related data in one or more loops further comprises comparing at each loop a new version of a word embedding with a previous version of the word embedding to determine whether quality of the word embedding is improved based on the plurality of evaluation criteria; and

generating the custom corpus comprising the domain related data in response to detecting no further improvements to the created embeddings based on the plurality of evaluation criteria.

2. The method of claim 1 , further comprising:

based on the evaluated domain related data, determining a completeness of the domain related data; and

in response to determining a completeness of the domain related data, generating a domain specific corpus of relevant data.

3. The method of claim 1 , wherein the seed documents are gathered by querying the knowledgebase, a database or a corpus.

4. The method of claim 1 , wherein the domain related data is a set of data gathered to represent an industry specific taxonomy.

5. The method of claim 1 , wherein the domain related data is identified using the machine learning algorithms comprising semi-supervised machine learning, supervised machine learning and unsupervised machine learning.

6. The method of claim 1 , wherein the domain related data is evaluated by analyzing the accuracy of a classification algorithm and a statistical analysis, wherein the evaluation results in a generated algorithm that is compared to a validation dataset.

7. A computer system for generating a custom corpus, comprising:

one or more processors, one or more computer-readable memories, one or more computer-readable tangible storage media, and program instructions stored on at least one of the one or more computer-readable tangible storage media for execution by at least one of the one or more processors via at least one of the one or more computer-readable memories, wherein the computer system is capable of performing a method comprising:

generating a domain graph comprising nodes representing terms associated with an industry taxonomy;

gathering seed data based on the generated domain graph by querying a knowledgebase using the nodes in the domain graph;

identifying domain related data based on the gathered seed data, wherein identifying the domain related data further comprises establishing a ground truth for the terms such that a true meaning of a term is determined according to a context of a respective domain associated with the industry taxonomy and further comprises identifying relevant and irrelevant data to a domain associated with the industry taxonomy based on the ground truth;

querying the domain related data and identifying other terms and candidate data to add to the industry taxonomy based on the query and relevancy to respective domains;

creating word embeddings for the terms associated with the domain related data;

evaluating the domain related data in one or more loops based on a plurality of evaluation criteria, wherein the plurality of evaluation criteria comprises determining and evaluating word coverage, context, and accuracy of the created word embeddings, and wherein evaluating the domain related data in one or more loops further comprises comparing at each loop a new version of a word embedding with a previous version of the word embedding to determine whether quality of the word embedding is improved based on the plurality of evaluation criteria; and

generating the custom corpus comprising the domain related data in response to detecting no further improvements to the created embeddings based on the plurality of evaluation criteria.

8. The computer system of claim 7 , further comprising:

based on the evaluated domain related data, determining a completeness of the domain related data; and

in response to determining a completeness of the domain related data, generating a domain specific corpus of relevant data.

9. The computer system of claim 7 , wherein the seed documents are gathered by querying the knowledgebase, a database or a corpus.

10. The computer system of claim 7 , wherein the domain related data is a set of data gathered to represent an industry specific taxonomy.

11. The computer system of claim 7 , wherein the domain related data is identified using the machine learning algorithms comprising semi-supervised machine learning, supervised machine learning and unsupervised machine learning.

12. The computer system of claim 7 , wherein the domain related data is evaluated by analyzing the accuracy of a classification algorithm and a statistical analysis, wherein the evaluation results in a generated algorithm that is compared to a validation dataset.

13. A computer program product for generating a custom corpus, comprising:

one or more computer-readable storage media and program instructions stored on at least one of the one or more computer-readable storage media, the program instructions executable by a processor to cause the processor to perform a method comprising:

generating a domain graph comprising nodes representing terms associated with an industry taxonomy;

gathering seed data based on the generated domain graph by querying a knowledgebase using the nodes in the domain graph;

identifying domain related data based on the gathered seed data, wherein identifying the domain related data further comprises establishing a ground truth for the terms such that a true meaning of a term is determined according to a context of a respective domain associated with the industry taxonomy and further comprises identifying relevant and irrelevant data to a domain associated with the industry taxonomy based on the ground truth;

querying the domain related data and identifying other terms and candidate data to add to the industry taxonomy based on the query and relevancy to respective domains;

creating word embeddings for the terms associated with the domain related data; and

evaluating the domain related data in one or more loops based on a plurality of evaluation criteria, wherein the plurality of evaluation criteria comprises determining and evaluating word coverage, context, and accuracy of the created word embeddings, and wherein evaluating the domain related data in one or more loops further comprises comparing at each loop a new version of a word embedding with a previous version of the word embedding to determine whether quality of the word embedding is improved based on the plurality of evaluation criteria; and

generating the custom corpus comprising the domain related data in response to detecting no further improvements to the created embeddings based on the plurality of evaluation criteria.

14. The computer program product of claim 13 , further comprising:

based on the evaluated domain related data, determining a completeness of the domain related data; and

in response to determining a completeness of the domain related data, generating a domain specific corpus of relevant data.

15. The computer program product of claim 13 , wherein the seed documents are gathered by querying the knowledgebase, a database or a corpus.

16. The computer program product of claim 13 , wherein the domain related data is a set of data gathered to represent an industry specific taxonomy.

17. The computer program product of claim 13 , wherein the domain related data is identified using the machine learning algorithms comprising semi-supervised machine learning, supervised machine learning and unsupervised machine learning.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Mar 1, 2019
From: MAZZOLENI, PIETRO; GIFFORD, WESLEY M; KHABIRI, ELHAM
To: INTERNATIONAL BUSINESS MACHINES CORPORATION
Reel/Frame 048474/0134 →
Continuity (1)
Related Publication 20200279171A1 · Sep 3, 2020