Fault tolerant method for processing data with human intervention
The present disclosure is directed to methods and non-transitory program storage devices for identifying demographic information in an input data file even in the face of known errors that would otherwise prevent the method from operating. When a fault condition is detected, a fault handler may attempt to fix the faulty data, remove the faulty data from the input data file being processed, or provide the faulty data at a user interface so a human user can intervene. The method and storage devices may continue processing the data regardless of whether human input has been received because the system can bypass or remove the data, thereby keeping the method fault tolerant.
1 . A computer-implemented method, comprising:
detecting, by one or more computing devices, a data type of one or more of a plurality of fields of demographic information associated with a data file from a third-party;
assigning, by the one or more computing devices, a naive label to the data type of the data file, wherein the data file comprises a mislabeled data type of the one or more of the plurality of fields of demographic information;
generating, by the one or more computing devices, a request for human intervention based on an identified exception condition associated with the mislabeled data type, further comprising:
providing, by the one or more computing devices via a user interface, a notification to notify a user device that the request for human intervention has been generated,
prior to receiving a user input at the user interface from the user device, bypassing, by the one or more computing devices, the mislabeled data type of the one or more of the plurality of fields of demographic information without interrupting an ongoing processing of the data file,
in response to the notification and receiving the user input, assigning, by the one or more computing devices, an active label to the mislabeled data type based on the received user input at the user interface, and
storing, by the one or more computing devices, one or more of the received user input and the assigned active label in a memory;
normalizing, by the one or more computing devices, naive labeled data and active labeled data; and
outputting, by the one or more computing devices, the normalized data using a format.
2 . The computer-implemented method of claim 1 , wherein:
the third-party is a healthcare provider;
the data file is a medical roster; and
the format is pre-specified by the healthcare provider.
3 . The computer-implemented method of claim 1 , the detecting of the data type is based on analyzing at least one of a semantic content, a data shape, a style, a name, or a phrase; and further comprising:
determining, by the one or more computing devices, the data type as at least one of an address or a phone number, based on analyzing at least one of a column type, a column title, a neighboring data type, or the active label stored in the memory.
4 . The computer-implemented method of claim 1 , further comprising suggesting, by the one or more computing devices, a possible data type for human confirmation based on determining a probability that the data type was identified correctly.
5 . The computer-implemented method of claim 1 , wherein the mislabeled data type with the active label is used as training data for at least one machine learning model.
6 . The computer-implemented method of claim 1 , wherein the one or more of the received user input and the assigned active label are used to determine an additional data type without being used as training data for a machine learning model.
7 . The computer-implemented method of claim 1 , further comprising validating, by the one or more computing devices, the normalized data.
8 . A non-transitory program storage device having instructions stored thereon that, when executed by at least one computing device, causes the at least one computing device to perform operations comprising:
detecting a data type of one or more of a plurality of fields of demographic information associated with a data file from a third-party;
assigning a naive label to the data type of the data file, wherein the data file comprises a mislabeled data type of the one or more of the plurality of fields of demographic information;
generating a request for human intervention based on an identified exception condition associated with the mislabeled data type, further comprising:
providing, via a user interface, a notification to notify a user device that the request for human intervention has been generated,
prior to receiving a user input at the user interface from the user device, bypassing the mislabeled data type of the one or more of the plurality of fields of demographic information without interrupting an ongoing processing of the data file,
in response to the notification and receiving the user input, assigning an active label to the mislabeled data type based on the received user input at the user interface, and
storing one or more of the received user input and the assigned active label in a memory;
normalizing naive labeled data and active labeled data; and
outputting the normalized data using a format.
9 . The non-transitory program storage device of claim 8 , wherein:
the third-party is a healthcare provider;
the data file is a medical roster; and
the format is pre-specified by the healthcare provider.
10 . The non-transitory program storage device of claim 8 , the detecting of the data type is based on analyzing at least one of a semantic content, a data shape, a style, a name, or a phrase; and further comprising:
determining the data type as at least one of an address or a phone number, based on analyzing at least one of a column type, a column title, a neighboring data type, or the active label stored in the memory.
11 . The non-transitory program storage device of claim 8 , wherein the operations further comprise suggesting a possible data type for human confirmation based on determining a probability that the data type was identified correctly.
12 . The non-transitory program storage device of claim 8 , wherein the mislabeled data type with the active label is used as training data for at least one machine learning model.
13 . The non-transitory program storage device of claim 8 , wherein the one or more of the received user input and the assigned active label are used to determine an additional data type without being used as training data for a machine learning model.
14 . The non-transitory program storage device of claim 8 , wherein the operations further comprise validating the normalized data.
15 . A system, comprising:
a memory; and
at least one processor coupled to the memory and configured to perform operations comprising:
detecting a data type of one or more of a plurality of fields of demographic information associated with a data file from a third-party;
assigning a naive label to the data type of the data file, wherein the data file comprises a mislabeled data type of the one or more of the plurality of fields of demographic information;
generating a request for human intervention based on an identified exception condition associated with the mislabeled data type, further comprising:
providing, via a user interface, a notification to notify a user device that the request for human intervention has been generated,
prior to receiving a user input at the user interface from the user device, bypassing the mislabeled data type of the one or more of the plurality of fields of demographic information without interrupting an ongoing processing of the data file,
in response to the notification and receiving the user input, assigning an active label to the mislabeled data type based on the received user input at the user interface, and
storing one or more of the received user input and the assigned active label in the memory;
normalizing naive labeled data and active labeled data; and
outputting the normalized data using a format.
16 . The system of claim 15 , the detecting of the data type is based on analyzing at least one of a semantic content, a data shape, a style, a name, or a phrase; and further comprising:
determining the data type as at least one of an address or a phone number, based on analyzing at least one of a column type, a column title, a neighboring data type, or the active label stored in the memory.
17 . The system of claim 15 , wherein the operations further comprise suggesting a possible data type for human confirmation based on determining a probability that the data type was identified correctly.
18 . The system of claim 15 , wherein the mislabeled data type with the active label is used as training data for at least one machine learning model.
19 . The system of claim 15 , wherein the one or more of the received user input and the assigned active label are used to determine an additional data type without being used as training data for a machine learning model.
20 . The system of claim 15 , wherein the operations further comprise validating the normalized data.