FOCUSING UNSTRUCTURED DATA AND GENERATING FOCUSED DATA DETERMINATIONS FROM AN UNSTRUCTURED DATA SET
Embodiments provide for improvements in generating focused data from an unstructured data set. The focused data generated from the unstructured data set may provide data insight(s) into analysis of the unstructured data set, and/or provide for improved capabilities for a user to efficiently navigate through relevant data portions of such data utilizing a user interface, even when such relevant data portions are not immediately distinguishable without further processing of the unstructured data set. Some embodiments receive an unstructured data set, extract an identified relevant subset utilizing at least one high-level extractor model, extract low-level relevant data from the identified relevant subset utilizing at least one low-level extractor model, generate fraud probability data by applying at least the low-level relevant data and the identified relevant subset to a fraud processing model, and output at least the fraud probability data, identified relevant subset, and/or low-level relevant data, and/or derivations therefrom.
1 . A computer-implemented method comprising:
receiving, by one or more processors, an unstructured data set;
extracting, by the processors and using a high-level extractor model, an identified relevant subset from the unstructured data set based at least in part on the unstructured data set;
extracting, by the processors and using a low-level extractor model, low-level relevant data from the identified relevant subset of the unstructured data set;
generating, by the processors and using a fraud processing model, fraud probability data based at least in part on the low-level relevant data and the identified relevant subset; and
outputting, by the one or more processors, the fraud probability data.
2 . The computer-implemented method of claim 1 , wherein outputting the fraud probability data comprises:
causing rendering of a user interface comprising the fraud probability data.
3 . The computer-implemented method of claim 2 , wherein the identified relevant subset comprises a renderable page, and wherein the user interface further comprises at least a first renderable page comprising a visually distinguished data portion based at least in part on the low-level relevant data.
4 . The computer-implemented method of claim 2 , wherein the user interface further includes the identified relevant subset.
5 . The computer-implemented method of claim 4 , wherein the user interface further displays a highlighted portion corresponding to the low-level relevant data.
6 . The computer-implemented method of claim 1 , wherein the high-level extractor model comprises a machine learning model that is specially trained to classify each portion of the unstructured data set as a selected classification from a plurality of candidate classifications.
7 . The computer-implemented method of claim 1 , wherein the high-level extractor model comprises at least one machine learning model that is specially trained for classification of a plurality of candidate classifications.
8 . The computer-implemented method of claim 1 , wherein at least one high-level extractor model comprises at least one of a text processing model or an image processing model.
9 . The computer-implemented method of claim 1 , wherein at least one low-level extractor model comprises a text processing model or an image processing model.
10 . The computer-implemented method of claim 1 further comprising:
identifying, using a page relevancy model, relevant text from the identified relevant subset based at least in part on the identified relevant subset,
wherein generating the fraud probability data further base at least in part on the relevant text.
11 . The computer-implemented method of claim 1 , further comprising:
generating, using a page relevancy model, page rating data corresponding to the identified relevant subset; and
outputting the page rating data.
12 . The computer-implemented method of claim 11 , wherein outputting the page rating data comprises:
causing rendering of a user interface comprising the identified relevant subset and a portion of the page rating data corresponding to each data portion of the identified relevant subset.
13 . The computer-implemented method of claim 1 further comprising:
extracting, using a keyword extraction model, an initial keyword set from the identified relevant subset, wherein the keyword extraction model generates a keyword relevance score for each keyword of the initial keyword set;
identifying a irrelevant keyword based at least in part on the keyword relevance score for each keyword of the initial keyword set and a keyword relevance threshold, wherein the irrelevant keyword is identified based at least in part on trusted description data corresponding to the keyword;
generating an updated keyword set by at least removing the irrelevant keyword from the initial keyword set;
generating a filtered keyword set by at least applying a dictionary filter model to the updated keyword set, wherein the dictionary filter model is based at least in part on a central truth source; and
outputting at least one keyword from the filtered keyword set.
14 . The computer-implemented method of claim 13 further comprising:
removing at least one unknown keyword from the updated keyword set.
15 . The computer-implemented method of claim 13 , wherein outputting the filtered keyword set comprises:
causing rendering of a user interface comprising at least one keyword of the filtered keyword set in at least one data portion of the identified relevant subset.
16 . The computer-implemented method of claim 1 further comprising:
training a first model based at least in part on a first data set, wherein the first data set is associated with a first model domain;
integrating the first model into a second model;
training an initial portion of the second model based at least in part on the first data set;
freezing the first model integrated into the second model,
training a remaining portion of the second model based at least in part on a second data set, wherein the second data set is associated with a second model domain; and
unfreezing the first model integrated into the second model,
wherein the second model is stored as the fraud processing model.
17 . The computer-implemented method of claim 16 further comprising:
increasing a learning rate of the second model while the first model integrated into the second model is frozen and during training the remaining portion of the second model; and
decreasing the learning rate of the second model after unfreezing the first model integrated into the second model.
18 . The computer-implemented method of claim 16 further comprising:
after unfreezing the first model integrated into the second model, fine-tuning the second model based at least in part on the second data set.
19 . A computing apparatus comprising a processor and memory including program code, the memory and the program code configured to, when executed by the processor, cause the computing apparatus to:
receive an unstructured data set;
extract, using a high-level extractor model, an identified relevant subset from the unstructured data set based at least in part on the unstructured data set;
extract, using a low-level extractor model, low-level relevant data from the identified relevant subset of the unstructured data set;
generate, using a fraud processing model, fraud probability data based at least in part on the low-level relevant data and the identified relevant subset; and
output the fraud probability data.
20 . A computer program product comprising a non-transitory computer-readable storage medium, the non-transitory computer-readable storage medium including instructions that, when executed by a computing apparatus, cause the computing apparatus to:
receive an unstructured data set;
extract, using a high-level extractor model, an identified relevant subset from the unstructured data set based at least in part on the unstructured data set;
extract, using a low-level extractor model, low-level relevant data from the identified relevant subset of the unstructured data set;
generate, using a fraud processing model, fraud probability data based at least in part on the low-level relevant data and the identified relevant subset; and
output the fraud probability data.