Optical character recognition filtering
An optical character recognition (OCR) filter described herein filters non-textual files in scanned customer data from OCR and pattern analysis of text generated thereof for sensitive customer data. The OCR filter is trained on files labelled using feature values for features generated from OCR applied to the corresponding files. Moreover, the OCR filter stores internal representations of the files during training to avoid leaking potential sensitive customer data contained therein. Once trained, performance of the OCR filter in filtering files comprising image data without text is evaluated according to false positive rates and false negative rates by comparing classifications of the OCR filter to classifications according to feature values for features generated from OCR. Evaluation of the OCR filter ensures continued model performance and informs model updates.
1 . An apparatus comprising:
a processor; and
a non-transitory computer-readable medium having instructions stored thereon that are executable by the processor to cause the apparatus to,
determine labels for a plurality of training files comprising image data indicating whether image data for each of the plurality of training files comprises text based, at least in part, on corresponding feature values for optical recognition features;
input the plurality of training files into a first component of a machine learning model;
store representations of the plurality of training files output by the first component of the machine learning model in secure storage for training;
discard the plurality of training files; and
train a second component of the machine learning model on the stored representations of the plurality of training files and corresponding labels.
2 . The apparatus of claim 1 , wherein the machine learning model comprises a convolutional neural network (CNN), wherein the first component comprises a first set of one or more internal layers of the CNN, wherein the second component comprises a second set of one or more internal layers of the CNN.
3 . The apparatus of claim 2 , wherein the first set of one or more internal layers comprises at least a convolutional layer, wherein the second set of one or more internal layers comprises at least a flattening layer and an activation layer.
4 . The apparatus of claim 1 , wherein the instructions executable by the processor to cause the apparatus to determine labels for the plurality of training files comprise instructions to determine whether the feature values for optical recognition features are below corresponding threshold values.
5 . The apparatus of claim 1 , wherein the optical recognition features comprise at least one of a number of bounding boxes, a number of words, a number of characters, a pixel width, and a pixel height of corresponding image data.
6 . The apparatus of claim 1 , wherein the computer-readable medium further has stored thereon instructions executable by the processor to cause the apparatus to evaluate the machine learning model on additional files for at least one of false positive rate and false negative rate.
7 . The apparatus of claim 1 , wherein the instructions to train the second component of the machine learning model comprise instructions executable by the processor to cause the apparatus to train the second component of the machine learning model to filter representations of files comprising non-textual image data.
8 . A non-transitory machine-readable medium having program code stored thereon, the program code comprising instructions to:
determine labels for a plurality of training files comprising image data indicating whether image data for each of the plurality of training files comprises text based, at least in part, on corresponding feature values for optical recognition features;
input the plurality of training files into a first component of a machine learning model;
store representations of the plurality of training files output by the first component of the machine learning model in secure storage for training;
discard the plurality of training files; and
train a second component of the machine learning model on the stored representations of the plurality of training files and corresponding labels.
9 . The non-transitory machine-readable medium of claim 8 , wherein the machine learning model comprises a convolutional neural network (CNN), wherein the first component comprises a first set of one or more internal layers of the CNN, wherein the second component comprises a second set of one or more internal layers of the CNN.
10 . The non-transitory machine-readable medium of claim 9 , wherein the first set of one or more internal layers comprises at least a convolutional layer, wherein the second set of one or more internal layers comprises at least a flattening layer and an activation layer.
11 . The non-transitory machine-readable medium of claim 8 , wherein the instructions to determine labels for the plurality of training files comprise instructions to determine whether the feature values for optical recognition features are below corresponding threshold values.
12 . The non-transitory machine-readable medium of claim 8 , wherein the optical recognition features comprise at least one of a number of bounding boxes, a number of words, a number of characters, a pixel width, and a pixel height of corresponding image data.
13 . The non-transitory machine-readable medium of claim 8 , wherein the program code further comprises instructions to evaluate the machine learning model on additional files for at least one of false positive rate and false negative rate.
14 . The non-transitory machine-readable medium of claim 8 , wherein the instructions to train the second component of the machine learning model comprise instructions to train the second component of the machine learning model to filter representations of files comprising non-textual image data.
15 . A method comprising:
determining labels for a plurality of training files comprising image data indicating whether image data for each of the plurality of training files comprises text based, at least in part, on corresponding feature values for optical recognition features;
inputting the plurality of training files into a first component of a machine learning model;
storing representations of the plurality of training files output by the first component of the machine learning model in secure storage for training;
discarding the plurality of training files; and
training a second component of the machine learning model on the stored representations of the plurality of training files and corresponding labels.
16 . The method of claim 15 , wherein the machine learning model comprises a convolutional neural network (CNN), wherein the first component comprises a first set of one or more internal layers of the CNN, wherein the second component comprises a second set of one or more internal layers of the CNN.
17 . The method of claim 16 , wherein the first set of one or more internal layers comprises at least a convolutional layer, wherein the second set of one or more internal layers comprises at least a flattening layer and an activation layer.
18 . The method of claim 15 , wherein determining labels for the plurality of training files comprises determining whether the feature values for optical recognition features are below corresponding threshold values.
19 . The method of claim 15 , wherein the optical recognition features comprise at least one of a number of bounding boxes, a number of words, a number of characters, a pixel width, and a pixel height of corresponding image data.
20 . The method of claim 15 , further comprising evaluating the machine learning model on additional files for at least one of false positive rate and false negative rate.