IP Library Granted Patent US 12,282,463
Granted Patent B2
US 12,282,463 · App. 18/198,144 · Granted Apr 22, 2025

Inline data quality schema management system

Inventors: Vivekanand Apuri (Hyderabad, IN); Naresh Dolani (Mumbai, IN); Sasi Reka Velliangiri (Chennai, IN)
Assignee: Bank of America Corporation
G06F16/215G06F16/22
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,282,463
App. No.
18/198,144
Granted
Apr 22, 2025
Kind
B2
Abstract

Various aspects of the disclosure relate to automatically inferring data quality rules for relational and/or non-SQL datasets. The data quality rules may be added as additional metadata for a data schema to improve and add validations for data to ensure data consistency and/or to identify invalid data. Rules may be automatically inferred based on data of data elements in one or more datasets and an identified significance of data in data points to extract common characteristics of data to create, teach and train one or more data quality (DQ) models for data elements across all data stores of the enterprise network.

Claims (55)

1. A system comprising:

an application computing system comprising a data repository; and

a computing device comprising:

at least one processor; and

non-transitory memory storing computer-readable instructions that, when executed by the at least one processor, cause the computing device to:

process a first raw data set received from a data source to generate processed data;

remove irrelevant information based on a category associated with the processed data and personally identifiable information, wherein the irrelevant information is removed from the data set;

identify, by a neural network, data quality rules from the processed data, wherein the data quality rules correspond to schemas across disparate database technologies and are automatically inferred based on data of data elements in the processed data and an identified significance of data in data points associated with extraction of common characteristics of data;

generate, by the neural network and based on multiple destination applications, a data quality rules model based on the processed data, wherein the data quality rules model ensures consistent data quality and accuracy of data elements for data sets used across the multiple destination applications and wherein each data element type of a plurality of data element types is associated with corresponding data quality metadata comprising a unique data quality ruleset;

import, from the data source, a second raw data set;

add, as data quality metadata for a first data schema for a particular relational database management system and for a second data quality schema for a noSQL databases, the data quality rules; and

train, based on feedback received from the application computing system and by the neural network, the data quality rules model, wherein the feedback comprises data quality rule effectiveness information.

2. The system of claim 1 , wherein the instructions cause the computing device to process, within a data fabric, the first raw data set to generate the processed data set, wherein the processed data comprises the processed data and associated metadata.

3. The system of claim 2 , wherein the instructions cause the computing device to process, within a data fabric, the first raw data set to generate the processed data set comprises cataloging each data element of the first raw data set.

4. The system of claim 2 , wherein the instructions cause the computing device to process, within a data fabric, the first raw data set to generate the processed data set comprises masking non-public information for each data element of the first raw data set.

5. The system of claim 2 , wherein the instructions cause the computing device to process, within a data fabric, the first raw data set to generate the processed data set comprises transforming a first data element from a source data type to a target data type, wherein the target data type is associated with the data repository.

6. The system of claim 2 , wherein the instructions cause the computing device to:

process, within a data fabric, the first raw data set to generate the processed data set comprises applying at least one data quality model to the first raw data set; and

synthesize data quality rules, comprising multiple generated data value patterns, while weighing context of data values by identifying data groups based on relevance and semantics of the processed data.

7. The system of claim 1 , wherein the instructions cause the computing device to:

retrieve, from a data confidence data store, data consumption information associated with data usage reports from the application computing system; and

train, a data quality model based on the data consumption information, wherein the data consumption information corresponds to data accuracy information and data quality rule effectiveness information.

8. A method comprising:

processing, a first raw data set received from a data source, to generate processed data;

removing irrelevant information based on a category associated with the processed data and personally identifiable information, wherein the irrelevant information is removed from the data set;

identifying, by a neural network, data quality rules from the processed data, wherein the data quality rules correspond to schemas across disparate database technologies and are automatically inferred based on data of data elements in the processed data and an identified significance of data in data points associated with extraction of common characteristics of the processed data;

generating, by the neural network and based on multiple destination applications, a data quality rules model based on the processed data, wherein the data quality rules model ensures consistent data quality and accuracy of data elements for data sets used across the multiple destination applications and wherein each data element type of a plurality of data element types is associated with corresponding data quality metadata comprising a unique data quality ruleset;

importing, from the data source, a second raw data set;

adding, as data quality metadata for a data schema for each particular relational database management system and each noSQL database, the data quality rules; and

training, based on feedback received from an application computing system and by the neural network, the data quality rules model, wherein the feedback comprises data quality rule effectiveness information.

9. The method of claim 8 , further comprising generating the processed data set, wherein the processed data comprises the processed data and associated metadata.

10. The method of claim 9 , further comprising cataloging each data element of the first raw data set.

11. The method of claim 9 , further comprising masking non-public information for each data element of the first raw data set.

12. The method of claim 9 , further comprising transforming a first data element from a source data type to a target data type, wherein the target data type is associated with a data repository.

13. The method of claim 9 , further comprising applying at least one data quality model to the first raw data set.

14. The method of claim 9 , further comprising:

retrieving, from a data confidence data store, data consumption information associated with data usage reports from an application computing system; and

training, a data quality model based on the data consumption information.

15. A computing device comprising:

at least one processor; and

non-transitory memory storing computer-readable instructions that, when executed by the at least one processor, cause the computing device to:

process, a first raw data set received from a data source, to generate processed data;

remove irrelevant information based on a category associated with the processed data and personally identifiable information, wherein the irrelevant information is removed from the data set;

identify, by a neural network, data quality rules from the processed data, wherein the data quality rules correspond to schemas across disparate database technologies and are automatically inferred based on data of data elements in the processed data and an identified significance of data in data points associated with extraction of common characteristics of the processed data;

generate, by the neural network and based on multiple destination applications, a data quality rules model based on the processed data, wherein the data quality rules model ensures consistent data quality and accuracy of data elements for data sets used across the multiple destination applications and wherein each data element type of a plurality of data element types is associated with corresponding data quality metadata comprising a unique data quality ruleset;

import, from the data source, a second raw data set;

add, as data quality metadata for a data schema for each particular relational database management system and each noSQL database, the data quality rules; and

train, based on feedback received from an application computing system and by the neural network, the data quality rules model, wherein the feedback comprises data quality rule effectiveness information.

16. The computing device of claim 15 , wherein the instructions cause the computing device to process, within a data fabric, the first raw data set to generate the processed data set, wherein the processed data comprises the processed data and associated metadata.

17. The computing device of claim 16 , wherein the instructions cause the computing device to process, within a data fabric, the first raw data set to generate the processed data set comprises cataloging each data element of the first raw data set.

18. The computing device of claim 16 , wherein the instructions cause the computing device to process, within a data fabric, the first raw data set to generate the processed data set comprises masking non-public information for each data element of the first raw data set.

19. The computing device of claim 16 , wherein the instructions cause the computing device to process, within a data fabric, the first raw data set to generate the processed data set comprises transforming a first data element from a source data type to a target data type, wherein the target data type is associated with a data repository.

20. The computing device of claim 15 , wherein the instructions cause the computing device to:

retrieve, from a data confidence data store, data consumption information associated with data usage reports from the application computing system; and

train, a data quality model based on the data consumption information.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded May 16, 2023
From: APURI, VIVEKANAND; DOLANI, NARESH; VELLIANGIRI, SASI REKA
To: BANK OF AMERICA CORPORATION
Reel/Frame 063659/0709 →
Continuity (1)
Related Publication 20240386000A1 · Nov 21, 2024
References Cited (17)
US 5842202A · Kon · 1998 [cited by applicant]
US 8401987B2 · Agrawal et al. · 2013 [cited by applicant]
US 9836713B2 · Bagchi et al. · 2017 [cited by applicant]
US 10318500B2 · Dani et al. · 2019 [cited by applicant]
US 11327935B2 · Yamashita et al. · 2022 [cited by applicant]
US 20050108631A1 · Amorin et al. · 2005 [cited by applicant]
US 20060212381A1 · Rowe, III · 2006 [cited by examiner]
US 20080082834A1 · Mattsson · 2008 [cited by examiner]
US 20120330911A1 · Gruenheid et al. · 2012 [cited by applicant]
US 20160004742A1 · Mohan et al. · 2016 [cited by applicant]
US 20180039680A1 · Nelke et al. · 2018 [cited by applicant]
US 20180089561A1 · Oliner · 2018 [cited by examiner]
US 20180373579A1 · Rathore et al. · 2018 [cited by applicant]
US 20190205636A1 · Saraswat · 2019 [cited by examiner]
US 20200082279A1 · Arora · 2020 [cited by examiner]
US 20210081476A1 · Weinstein · 2021 [cited by examiner]
US 20230135962A1 · Lee · 2023 [cited by examiner]