IP Library Granted Patent US 11,669,682
Granted Patent B2
US 11,669,682 · App. 17/130,850 · Granted Jun 6, 2023

Bespoke transformation and quality assessment for term definition

Inventors: Gretel De Paepe (Watermaal-Bosvoorde, BE); Michael Tandecki (Vilvoorde, BE)
Assignee: Collibra Belgium BV
G06F40/242G06F18/214G06F40/284G06F40/40G06N3/04G06N20/00
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,669,682
App. No.
17/130,850
Granted
Jun 6, 2023
Kind
B2
Abstract

An enterprise data management system with definition quality assessment capabilities for automatically assessing the quality of definitions for terms stored in the enterprise data management system. The system can include a processor programmed to receive a term and a corresponding definition. The processor assess the quality of the definition, including for each of a plurality of quantifiable definition guidelines: deriving feature inputs based on the definition; feeding the feature inputs into a machine learning model corresponding to the definition guideline; and receiving a quality score for the definition guideline from the corresponding machine learning model. An overall quality score is calculated based on the quality score for each of the definition guidelines. The overall quality score and the quality score for each of the plurality of definition guidelines is displayed and if the overall quality score is less than a selected threshold score, a transformation of the definition is recommended.

Claims (61)

1. An enterprise data management system with definition quality assessment capabilities for automatically assessing the quality of definitions for terms stored in the enterprise data management system, the system comprising:

at least one memory device storing instructions for causing at least one processor to:

receive a term;

receive a definition corresponding to the term;

assess the quality of the definition, including for each of a plurality of quantifiable definition guidelines:

deriving at least one feature input based on at least the definition;

feeding the at least one feature input into a machine learning model corresponding to the definition guideline; and

receiving a quality score for the definition guideline from the corresponding machine learning model;

wherein one of the plurality of definition guidelines is a structure guideline and wherein evaluating the definition with respect to the structure guideline comprises deriving a part of speech feature input and a part of speech bag of words input;

wherein the machine learning model corresponding to the structure guideline includes a CNN and an LSTM, and wherein feeding at least one feature input into the machine learning model comprises:

feeding the part of speech feature input into the CNN and the LSTM;

concatenating an output of the CNN and an output of the LSTM with the part of speech bag of words input to create a concatenated input; and

feeding the concatenated input into a subsequent NN;

calculate an overall quality score based on the quality score for each of the plurality of definition guidelines;

display the overall quality score and the quality score for each of the plurality of definition guidelines; and

if the overall quality score is less than a selected threshold score, recommend a transformation of the definition.

2. The system of claim 1 , further comprising instructions to receive a transformed version of the definition and assess the quality of the transformed version of the definition.

3. The system of claim 1 , further comprising instructions to receive user feedback for one or more of the displayed quality scores and input the received user feedback into a retraining process associated with the machine learning model corresponding to each of the one or more quality scores.

4. The system of claim 1 , wherein evaluating at least one of the plurality of definition guidelines includes deriving at least one feature input based on the definition and the term.

5. The system of claim 1 , wherein evaluating at least one of the plurality of definition guidelines includes deriving a number of words feature input and a number of sentences feature input.

6. The system of claim 1 , wherein deriving the at least one feature input based on at least the definition includes calculating a feature metric.

7. The system of claim 1 , further comprising instructions to train each machine learning model corresponding to each of the definition guidelines with a set of definitions each labeled as to whether it satisfies the corresponding definition guideline.

8. An enterprise data management system with definition quality assessment capabilities for automatically assessing the quality of definitions for terms stored in the enterprise data management system, the system comprising:

at least one memory device storing instructions for causing at least one processor to:

receive a term;

receive a plurality of definitions corresponding to the term;

assess the quality of each definition, including:

for each of a plurality of quantifiable definition guidelines:

deriving at least one feature input based on the definition;

feeding the at least one feature input into a machine learning model corresponding to the definition guideline; and

receiving a quality score for the definition guideline from the corresponding machine learning model;

wherein one of the plurality of definition guidelines is a structure guideline and wherein evaluating the definition with respect to the structure guideline comprises deriving a part of speech feature input and a part of speech bag of words input;

wherein the machine learning model corresponding to the structure guideline includes a CNN and an LSTM, and wherein feeding at least one feature input into the machine learning model comprises:

feeding the part of speech feature input into the CNN and the LSTM;

concatenating an output of the CNN and an output of the LSTM with the part of speech bag of words input to create a concatenated input; and

feeding the concatenated input into a subsequent NN;

calculate an overall quality score for each definition based on the quality score for each of the plurality of definition guidelines;

display the overall quality score for each definition; and

select the definition having the highest overall quality score as the only definition for the term.

9. The system of claim 8 , wherein evaluating at least one of the plurality of definition guidelines includes deriving at least one feature input based on the definition and the term.

10. The system of claim 8 , wherein evaluating at least one of the plurality of definition guidelines includes deriving a number of words feature input and a number of sentences feature input.

11. The system of claim 8 , wherein deriving the at least one feature input based on at least the definition includes calculating a feature metric.

12. The system of claim 8 , further comprising instructions to train each machine learning model corresponding to each of the definition guidelines with a set of definitions each labeled as to whether it satisfies the corresponding definition guideline.

13. An enterprise data management system with definition quality assessment capabilities for automatically assessing the quality of definitions for terms stored in the enterprise data management system, the system comprising:

at least one memory device storing instructions for causing at least one processor to:

receive a term;

receive a definition corresponding to the term;

assess the quality of the definition, including for each of a plurality of quantifiable definition guidelines:

deriving at least one feature input based on at least the definition;

feeding the at least one feature input into a machine learning model corresponding to the definition guideline; and

receiving a quality score for the definition guideline from the corresponding machine learning model;

wherein one of the plurality of definition guidelines is a structure guideline and wherein evaluating the definition with respect to the structure guideline comprises deriving a part of speech feature input and a part of speech bag of words input;

wherein the machine learning model corresponding to the structure guideline includes a CNN and an LSTM, and wherein feeding at least one feature input into the machine learning model comprises:

feeding the part of speech feature input into the CNN and the LSTM;

concatenating an output of the CNN and an output of the LSTM with the part of speech bag of words input to create a concatenated input; and

feeding the concatenated input into a subsequent NN;

calculate an overall quality score based on the quality score for each of the plurality of definition guidelines; and

display the overall quality score and the quality score for each of the plurality of definition guidelines;

wherein each machine learning model corresponding to each of the definition guidelines is trained with a set of definitions each labeled as to whether it satisfies the corresponding definition guideline.

14. The system of claim 13 , further comprising instructions to receive a transformed version of the definition and assess the quality of the transformed version of the definition.

15. The system of claim 13 , wherein evaluating at least one of the plurality of definition guidelines includes deriving a number of words feature input and a number of sentences feature input.

Assignments (3)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Mar 15, 2023
From: COLLIBRA NV; CNV NEWCO B.V.; COLLIBRA B.V.
To: COLLIBRA BELGIUM BV
Reel/Frame 062989/0023 →
SECURITY INTEREST Recorded Jan 4, 2023
From: COLLIBRA BELGIUM BV
To: SILICON VALLEY BANK UK LIMITED
Reel/Frame 062269/0955 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Dec 31, 2020
From: DE PAEPE, GRETEL; TANDECKI, MICHAEL
To: COLLIBRA NV
Reel/Frame 054787/0924 →
Continuity (1)
Related Publication 20220198139A1 · Jun 23, 2022