IP Library › Granted Patent US 11,494,611
Granted Patent B2
US 11,494,611 · App. 16/527,546 · Granted Nov 8, 2022

Metadata-based scientific data characterization driven by a knowledge database at scale

Inventors: Renan Francisco Santos Souza (Rio de Janeiro, BR); Reinaldo Mozart da Gama e Silva (Rio de Janeiro, BR); Rodrigo da Silva Ferreira (Rio de Janeiro, BR); Emilio Ashton Vital Brazil (Rio de Janeiro, BR); Viviane Torres da Silva (Rio de Janeiro, BR)
Assignee: INTERNATIONAL BUSINESS MACHINES CORPORATION
G06N3/0445G06F16/144G06F16/2448G06F16/24564G06F16/24573
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,494,611
App. No.
16/527,546
Granted
Nov 8, 2022
Kind
B2
Abstract

A metadata-based scientific data characterization method, system, and computer program product include requesting a user input for a task to specify a rule for the task to determine a quality and a relationship of a data file in a data file database based on metadata associated with the data file, processing a user feedback of results using the rule run on the data file database and tracking the user feedback on the results in order to learn from the user feedback, and based on the learning, creating a modified rule to determine a quality and a relationship of a second data file.

Claims (42)

1. A computer-implemented metadata-based characterization method for a knowledge database including heterogeneous files, the method comprising:

based on a task, requesting a user input to specify a rule that characterizes a single raw data file of the heterogeneous files in the knowledge database according to an extension name and associated metadata such that both of a quality and a relationship of the single raw data file in the knowledge database is determined by the user with respect to the task, the relationship that is input by the user input being between the single raw data file and another data file in the knowledge database and being an intersection of a content of data between the single raw data file and another data file;

applying, via a processor including scalable processing in a parallel hardware making use of a domain-specific raw data parser running a machine-learning based service, the rule at scale against a plurality of data files in the knowledge database to identify other data files of the plurality of data files that have a same relationship and associated metadata specified in the rule by only processing the data files with metadata specified in the rule thereby speeding up processing by not processing entire knowledge database;

processing, via the processor running the machine-learning based service, a user feedback of a result including the other data files of the plurality of data files using the rule run on the plurality of data files of the knowledge database and tracking the user feedback on the result in order to learn, via the processor, from the user feedback to identify that other data files having the same relationship specified in the rule are a part of the result;

based on the user feedback and by iteratively using the machine-learning based service at each iteration, creating a modified rule via the processor, at each iteration, to determine a quality and a relationship of a second data file while also requesting an approval or an adjustment from the user of the modified rule, the modified rule replacing the rule when applying at scale against the knowledge database; and

recording the modified rule in the knowledge database after the approval or the adjustment from the user to recommend the modified rule for a different user that fits the task,

wherein the requesting interactively and iteratively systematizes the user input for the metadata associated to characterize a number of data files less than a second number of data files in the knowledge database,

further comprising determining a quality and a relationship of data files in the knowledge database by applying the modified rule, at scale, on the knowledge database,

wherein the extension name is correlated with a type of the data file.

2. The method of claim 1 , wherein the processing further processes a second user feedback based on the modified rule run at scale to iteratively create a third modified rule, and

wherein, based on the learning, the third modified rule is created to determine a quality and a relationship of a rest of the data files in the knowledge database.

3. The method of claim 1 , wherein, based on the learning, the modified rule is created to determine a quality and a relationship of a rest of the data files in the knowledge database.

4. The method of claim 1 , further comprising applying the modified rule for a rest of the data files in the data file database to determine a quality and a relationship of the data files in the knowledge database.

5. The method of claim 4 , wherein the processing further processes a second user feedback based on the modified rule run at scale to iteratively create a third modified rule.

6. The method of claim 1 , wherein the processing is performed by a cloud computing environment to offload computing requirements from a user device to the cloud computing environment, and

wherein the cloud computing environment comprises a cloud computing model of a service delivery comprising two or more clouds of a private cloud, a community cloud, and a public cloud that remain unique entities but are bound together by technology that enables data and application portability that results in load-balancing between the two or more clouds.

7. A computer program product for metadata-based characterization, the computer program product comprising a computer-readable storage medium having program instructions embodied therewith, the program instructions executable by a computer to cause the computer to perform:

based on a task, requesting a user input to specify a rule that characterizes a single raw data file of the heterogeneous files in the knowledge database according to an extension name and associated metadata such that both of a quality and a relationship of the single raw data file in the knowledge database is determined by the user with respect to the task, the relationship that is input by the user input being between the single raw data file and another data file in the knowledge database and being an intersection of a content of data between the single raw data file and another data file;

applying, via a processor including scalable processing in a parallel hardware making use of a domain-specific raw data parser running a machine-learning based service, the rule at scale against a plurality of data files in the knowledge database to identify other data files of the plurality of data files that have a same relationship and associated metadata specified in the rule by only processing the data files with metadata specified in the rule thereby speeding up processing by not processing entire knowledge database;

processing, via the processor running the machine-learning based service, a user feedback of a result including the other data files of the plurality of data files using the rule run on the plurality of data files of the knowledge database and tracking the user feedback on the result in order to learn, via the processor, from the user feedback to identify that other data files having the same relationship specified in the rule are a part of the result;

based on the user feedback and by iteratively using the machine-learning based service at each iteration, creating a modified rule via the processor, at each iteration, to determine a quality and a relationship of a second data file while also requesting an approval or an adjustment from the user of the modified rule, the modified rule replacing the rule when applying at scale against the knowledge database; and

recording the modified rule in the knowledge database after the approval or the adjustment from the user to recommend the modified rule for a different user that fits the task,

wherein the requesting interactively and iteratively systematizes the user input for the metadata associated to characterize a number of data files less than a second number of data files in the knowledge database,

further comprising determining a quality and a relationship of data files in the knowledge database by applying the modified rule, at scale, on the knowledge database,

wherein the extension name is correlated with a type of the data file.

8. The computer program product of claim 7 , wherein the processing further processes a second user feedback based on the modified rule run at scale to iteratively create a third modified rule, and

wherein, based on the learning, the third modified rule is created to determine a quality and a relationship of a rest of the data files in the knowledge database.

9. The computer program product of claim 7 , wherein, based on the learning, the modified rule is created to determine a quality and a relationship of a rest of the data files in the knowledge database.

10. The computer program product of claim 7 , further comprising applying the modified rule for a rest of the data files in the data file database to determine a quality and a relationship of the data files in the knowledge database.

11. A metadata-based characterization system, the system comprising:

a processor; and

a memory, the memory storing instructions to cause the processor to perform:

based on a task, requesting a user input to specify a rule that characterizes a single raw data file of the heterogeneous files in the knowledge database according to an extension name and associated metadata such that both of a quality and a relationship of the single raw data file in the knowledge database is determined by the user with respect to the task, the relationship that is input by the user input being between the single raw data file and another data file in the knowledge database and being an intersection of a content of data between the single raw data file and another data file;

applying, via a processor including scalable processing in a parallel hardware making use of a domain-specific raw data parser running a machine-learning based service, the rule at scale against a plurality of data files in the knowledge database to identify other data files of the plurality of data files that have a same relationship and associated metadata specified in the rule by only processing the data files with metadata specified in the rule thereby speeding up processing by not processing entire knowledge database;

processing, via the processor running the machine-learning based service, a user feedback of a result including the other data files of the plurality of data files using the rule run on the plurality of data files of the knowledge database and tracking the user feedback on the result in order to learn, via the processor, from the user feedback to identify that other data files having the same relationship specified in the rule are a part of the result;

based on the user feedback and by iteratively using the machine-learning based service at each iteration, creating a modified rule via the processor, at each iteration, to determine a quality and a relationship of a second data file while also requesting an approval or an adjustment from the user of the modified rule, the modified rule replacing the rule when applying at scale against the knowledge database; and

recording the modified rule in the knowledge database after the approval or the adjustment from the user to recommend the modified rule for a different user that fits the task,

wherein the requesting interactively and iteratively systematizes the user input for the metadata associated to characterize a number of data files less than a second number of data files in the knowledge database,

further comprising determining a quality and a relationship of data files in the knowledge database by applying the modified rule, at scale, on the knowledge database,

wherein the extension name is correlated with a type of the data file.

12. The system of claim 11 , wherein the processing further processes a second user feedback based on the modified rule run at scale to iteratively create a third modified rule, and

wherein, based on the learning, the third modified rule is created to determine a quality and a relationship of a rest of the data files in the knowledge database.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jul 31, 2019
From: SOUZA, RENAN FRANCISCO SANTOS; DA GAMA E SILVA, REINALDO MOZART; DA SILVA FERREIRA, RODRIGO; ASHTON VITAL BRAZIL, EMILIO; SILVA, VIVIANE TORRES DA
To: INTERNATIONAL BUSINESS MACHINES CORPORATION
Reel/Frame 049922/0896 →
Continuity (1)
Related Publication 20210034948A1 · Feb 4, 2021