IP Library Granted Patent US 10,795,752
Granted Patent B2
US 10,795,752 · App. 16/002,516 · Granted Oct 6, 2020

Data validation

Inventors: Chung-Sheng Li (San Jose, CA); Emmanuel Munguia Tapia (San Jose, CA); Mohammad Ghorbani (Foster City, CA); Jingyun Fan (Berkeley, CA); Priyankar Bhowal (Gurgaon, IN); David Clune (Knoxville, TN); Sumraat Singh (Faridabad, IN)
Assignee: ACCENTURE GLOBAL SOLUTIONS LIMITED
G06F11/079G06F11/0751G06F11/0772G06F11/0787G06F11/0793G06F40/279G06F40/30
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 10,795,752
App. No.
16/002,516
Granted
Oct 6, 2020
Kind
B2
Abstract

In an example, data, such as, a journal entry in a ledger, to be validated and associated supporting documents may be extracted. Further, an entity, indicative of a feature of the data may be extracted. Based on the extracted entity, one or more probable values for a field of the data may be determined. A probability score may be associated each of the probable values of the field. At least one of the probable values of the field may be compared with an actual value of the field of the data. Based on comparison, a notification indicative of a potential error in the data may generated. The data and historical data associated with the data may be processed, based on at least one of predefined rules and a machine learning technique, to detect an anomaly in the data, the anomaly being related to a contextual information associated with the data.

Claims (75)

1. A validation system comprising:

a processor;

a data extractor coupled to the processor to,

extract data to be validated and associated supporting documents, the data including a journal entry in a ledger stored in a database, wherein validation of the data includes reconciliation of the journal entry in the ledger to identify a potential error and an anomaly indicative of a deviation from a predefined behavior in the journal entry, and wherein the reconciliation is performed continuously in near-real time, in response to detection of a new journal entry in the database; and

extract an entity associated with the data stored in the database using a natural language processing technique, the entity being indicative of a feature of the data associated with the journal entry in the ledger;

a classifier coupled to the processor to:

determine a value for a field of the data corresponding to the journal entry, based on the extracted entity, the field of the data including information to be validated;

ascertain whether the extracted entity and the field have a one to one mapping, based on historical data associated with the extracted entity;

when the extracted entity and the field have the one to one mapping, select a corresponding value of the field as the value;

when the extracted entity and the field do not have the one to one mapping, obtain augmented extracted entity data;

process the augmented extracted entity data to determine one or more values of the field;

associate a probability score with each of the one or more values of the field of the data corresponding to the journal entry, the probability score being indicative of a likelihood of determination of the one or more values of the field of the data corresponding to the journal entry being correct;

compare the determined one or more values of the field of the data corresponding to the journal entry with an actual value of the field of the data present in the journal entry; and

based on the comparison, generate a notification indicative of the potential error and the anomaly in the data; and

an anomaly detector coupled to the processor, the anomaly detector to,

process the data and the historical data associated with the data, based on at least one of predefined rules and a machine learning technique; and

detect the anomaly in the data, based on the processing, the anomaly being related to contextual information indicative of the deviation from behavioral aspects associated with the data.

2. The system as claimed in claim 1 , wherein the system further comprises an augmentor coupled to the processor, the augmentor is to augment the extracted entity with augmentation data to determine features related to the extracted entity, the augmentation data including information pertaining to fields of the data from an external source.

3. The system of claim 2 , wherein the classifier is to process the extracted data using augmented extracted data to determine the one or more values.

4. The system as claimed in claim 1 , wherein the anomaly detector comprises:

a semantic matcher to determine whether a description of the journal entry is semantically similar to corresponding historical data to detect a semantic anomaly;

an outlier detector to determine whether the journal entry is in a standard statistical range to detect an outlier anomaly; and

a rule based anomaly detector to determine whether a journal entry property fits into a pre-defined acceptable range to detect a rule based anomaly.

5. The system of claim 1 , wherein the system comprises a data generator coupled to the processor to generate the data to be validated, based on a supporting document.

6. The system as claimed in claim 1 , wherein the system further includes a notification analyzer to,

receive the notification indicating that the data includes one of the potential error and the anomaly;

generate a hypothesis providing an explanation for the potential error and the anomaly; and

provide a remedial action for correcting the potential error and the anomaly.

7. A method comprising:

extracting data to be validated and associated supporting documents, the data including a journal entry in a ledger stored in a database, wherein validation of the data includes reconciliation of the journal entry in the ledger to identify a potential error and an anomaly indicative of a deviation from a predefined behavior in the journal entry and wherein the reconciliation is performed continuously in near-real time, in response to detection of a new journal entry in the database; and

extracting an entity associated with the data stored in the database using a natural language processing technique, the entity being indicative of a feature of the data associated with the journal entry in the ledger;

determining a value for a field of the data corresponding to the journal entry, based the extracted entity, the field of the data including information to be validated;

ascertaining whether the extracted entity and the field have a one to one mapping, based on historical data associated with the extracted entity;

when the extracted entity and the field have the one to one mapping, selecting a corresponding value of the field as the value;

when the extracted entity and the field do not have the one to one mapping, obtaining augmented extracted entity data;

processing the augmented extracted entity data to determine one or more values of the field;

associating a probability score with each of the one or more values of the field of the data corresponding to the journal entry, the probability score being indicative of a likelihood of determination of the one or more values of the field of the data corresponding to the journal entry being correct;

comparing the determined one or more values of the field of the data corresponding to the journal entry with an actual value of the field of the data present in the journal entry;

based on the comparison, generating a notification indicative of the potential error and the anomaly in the data;

processing the data and the historical data associated with the data, based on at least one of predefined rules and a machine learning technique; and

detecting the anomaly in the data, based on the processing, the anomaly being related to contextual information indicative of the deviation from behavioral aspects associated with the data.

8. The method as claimed in claim 7 , wherein the method further comprises augmenting the extracted entity with augmentation data to determine features related to the extracted entity, the augmentation data including information pertaining to fields of the data from an external source.

9. The method as claimed in claim 8 , wherein the one or more values for the field is determined based on the augmented extracted data.

10. The method as claimed in claim 7 , wherein processing the data and the historical data for detecting the anomaly comprises:

determining whether a description of the journal entry is semantically similar to corresponding historical data to detect a semantic anomaly;

determining whether the journal entry is in a standard statistical range to detect an outlier anomaly; and

determining whether a journal entry property fits into a pre-defined acceptable range to detect a rule based anomaly.

11. The method as claimed in claim 7 , wherein the method further comprises generating the data to be validated, based on a supporting document.

12. The method as claimed in claim 7 , wherein the method further comprises:

receiving the notification indicating that the data includes one of the potential error and the anomaly;

generating a hypothesis providing an explanation for the potential error and the anomaly; and

providing a remedial action for correcting the potential error and the anomaly.

13. A non-transitory computer readable medium including machine readable instructions that are executable by a processor to:

extract data to be validated and associated supporting documents, the data including a journal entry in a ledger stored in a database, wherein validation of the data includes reconciliation of the journal entry in the ledger to identify a potential error and an anomaly indicative of a deviation from a predefined behavior in the journal entry and wherein the reconciliation is performed continuously in near-real time, in response to, detection of a new journal entry in the database; and

extract an entity associated with the data stored in the database using a natural language processing technique, the entity being indicative of a feature of the data associated with the journal entry in the ledger;

determine a value for a field of the data corresponding to the journal entry, based the extracted entity, the field of the data including information to be validated;

ascertain whether the extracted entity and the field have a one to one mapping, based on historical data associated with the extracted entity;

when the extracted entity and the field have the one to one mapping, select a corresponding value of the field as the value;

when the extracted entity and the field do not have the one to one mapping, obtain augmented extracted entity data;

process the augmented extracted entity data to determine one or more values of the field;

associate a probability score with each of the one or more values of the field of the data corresponding to the journal entry, the probability score being indicative of a likelihood of determination of the one or more values being correct;

compare the determined one or more values of the field of the data corresponding to the journal entry with an actual value of the field of the data present in the journal entry;

based on the comparison, generate a notification indicative of the potential error and the anomaly in the data;

process the data and the historical data associated with the data, based on at least one of predefined rules and a machine learning technique; and

detect the anomaly in the data, based on the processing, the anomaly being related to contextual information indicative of deviation from behavioral aspects associated with the data.

14. The non-transitory computer readable medium as claimed in claim 13 , wherein the one or more values for the field is determined based on the augmented extracted data.

15. The non-transitory computer readable medium as claimed in claim 14 , wherein the processor is to:

receive the notification indicating that the data includes one of the potential error and the anomaly;

generate a hypothesis providing an explanation for the potential error and the anomaly; and

provide a remedial action for correcting the potential error and the anomaly.

16. The non-transitory computer readable medium as claimed in claim 13 , wherein the processor is to augment the extracted entity with augmentation data to determine features related to the extracted entity, the augmentation data including information pertaining to fields of the data from an external source.

17. The non-transitory computer readable medium as claimed in claim 13 , wherein the processor is to

determine whether a description of the journal entry is semantically similar to corresponding historical data to detect a semantic anomaly;

determine whether the journal entry is in a standard statistical range to detect an outlier anomaly; and

determine whether a journal entry property fits into a pre-defined acceptable range to detect a rule based anomaly.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jun 22, 2018
From: LI, CHUNG-SHENG; MUNGUIA TAPIA, EMMANUEL; GHORBANI, MOHAMMAD; FAN, JINGYUN; BHOWAL, PRIYANKAR; CLUNE, DAVID; SINGH, SUMRAAT
To: ACCENTURE GLOBAL SOLUTIONS LIMITED
Reel/Frame 046178/0787 →
Continuity (1)
Related Publication 20190377624A1 · Dec 12, 2019