IP Library › Granted Patent US 11,887,010
Granted Patent B2
US 11,887,010 · App. 15/842,965 · Granted Jan 30, 2024

Data classification for data lake catalog

Inventors: Marcio T. Moura (Wellington, FL); Qiqing C. Ouyang (Yorktown Heights, NY); Jo A. Ramos (Grapevine, TX); Deepak Rangarao (Cupertino, CA)
Assignee: International Business Machines Corporation
G06N5/02G06F16/90332G06F16/90344G06N20/00G06Q10/067
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,887,010
App. No.
15/842,965
Granted
Jan 30, 2024
Kind
B2
Abstract

Data classification extracts from structured text business data inputs, via natural language understanding processing, training set data elements (training keywords, training concepts, training entities, and/or training taxonomy classifications). Embodiments identify associations within the structured training business data of business class categories with respective ones of extracted training set data elements, and build a logical relationship data classification training knowledge base ontology that connects business classes to respective associated ones of extracted training data elements as questions into knowledge base ontology question-business class associations.

Claims (26)

1. A computer-implemented method for a data classifier the method comprising executing on a computer processor:

extracting from a structured text business data input, via natural language understanding processing, training set data elements that are selected from the group consisting of training keywords, training concepts, training entities, and training taxonomy classifications;

identifying associations within the structured text business data of each of a plurality of business classes with respective ones of the extracted training set data elements, wherein the business classes comprise at least one of business terms, entity names, attribute names, and column names; and

building a logical relationship data classification training knowledge base ontology that connects ones of the business classes to respective associated ones of the extracted training data elements as questions, into a plurality of knowledge base ontology question-business class associations.

2. The method of claim 1 , further comprising:

integrating computer-readable program code into a computer system comprising a processor, a computer readable memory in circuit communication with the processor, and a computer readable storage medium in circuit communication with the processor; and

wherein the processor executes program code instructions stored on the computer-readable storage medium via the computer readable memory and thereby performs the extracting the training set data elements, the identifying the associations within the structured training business data, and the building the logical relationship data classification training knowledge base ontology.

3. The method of claim 2 , wherein the computer-readable program code is provided as a service in a cloud environment.

4. The method of claim 1 , further comprising:

extracting, via a natural language understanding process, a plurality of entity data elements from an entity business data input that defines an organizational attribute of a business entity, wherein the plurality of entity data elements are selected from the group consisting of entity keywords, entity concepts, constituent entities, and entity taxonomy classifications;

identifying candidates of the business classes that are likely associated with each of the entity data elements extracted for the business entity as functions of strength of match to the question-business class associations defining the knowledge base ontology;

determining confidence values for each of the business classes candidates that represent respective strengths of matching to respective ones of the knowledge base ontology question-business class associations; and

building, via a natural language classifier process, a classifier ontology for the entity that links the extracted entity data elements to respective ones of the business class candidates to which they have highest confidence values of association; and

wherein the entity business data input is selected from the group consisting of structured text data, semi-structured text data and unstructured text data.

5. The method of claim 4 , further comprising:

identifying, via a feedback auditing process, an error in a data classification linking a first of the extracted entity data elements as a first question to a first of the business class candidates within the question-business class associations of the built classifier;

revising the built classifier structure to delete a logical link of the first question to the first business class candidate, and to replace the deleted link with a new link of the first question to a second candidate of the business classes that has a next-highest confidence score of association to the extracted entity element of the first question; and

revising the knowledge base ontology to reflect the revision to the built classifier structure.

6. The method of claim 5 , wherein the extracting the training set data elements from the structured text business data input via the natural language understanding processing comprises:

deriving semantic information from input text content; and

using the natural language understanding processing to generate sentiment values for detected entities and keywords as a function of the derived semantic information.

7. The method of claim 6 , the building the logical relationship data classification training knowledge base ontology further comprising:

connecting business classes within the structured text business data input to respected identified associated ones of the extracted training data elements as questions into a plurality of question-business class associations;

generating predictions with respect to best ones of the business classes for matching to short text content selections within the extracted training data elements, wherein the short text content selections are less than a totality of the extracted training data elements; and

building the classifier ontology by selecting the business class candidates predicted as best ones for matching to the short text content selections as the ones of the business class candidates linked to the extracted entity data elements.

8. The method of claim 7 , wherein the short text content selections are selected from the group consisting of a sentence of a plurality of text words, and a phrase of a plurality of text words.

Assignments (2)
CORRECTIVE ASSIGNMENT TO CORRECT THE EXECUTION DATE PREVIOUSLY RECORDED AT REEL: 044416 FRAME: 0555. ASSIGNOR(S) HEREBY CONFIRMS THE ASSIGNMENT . Recorded Feb 1, 2018
From: MOURA, MARCIO T; OUYANG, QIQING C.; RAMOS, JO A.; RANGARAO, DEEPAK
To: INTERNATIONAL BUSINESS MACHINES CORPORATION
Reel/Frame 045215/0962 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Dec 18, 2017
From: MOURA, MARCIO T.; OUYANG, QIQING C.; RAMOS, JO A.; RANGARAO, DEEPAK
To: INTERNATIONAL BUSINESS MACHINES CORPORATION
Reel/Frame 044416/0555 →
Continuity (2)
Continuation 15823771 · Nov 28, 2017
Related Publication 20190164063A1 · May 30, 2019