Remediation of unstructured data using artificial intelligence
The systems and methods disclosed herein obtain (e.g., via a user interface) a collection of unstructured data, where each document includes a content set. Using a first AI model set, multiple summaries are generated by categorizing each document into clusters based on vector comparisons of content sets and summarizing the content for each cluster. A second AI model set (same as or different from the first AI model set) identifies duplicate content within the unstructured data by generating similarity values between pairs of summaries and determining if the similarity values meet a predefined threshold. A report is generated (e.g., on the user interface) indicating the duplicate content sets and/or the collection of unstructured data.
1 . A system comprising:
at least one hardware processor; and
at least one non-transitory memory storing instructions, which, when executed by the at least one hardware processor, cause the system to:
obtain a dataset that comprises a plurality of content sets;
generate multiple data representations defining the plurality of content sets by:
linking each content set of the plurality of content sets into one or more clusters, and
for each particular cluster, generate a respective data representation of the multiple data representations by summarizing respective content sets of the particular cluster;
identify, using a first artificial intelligence (AI) model set, at least one particular content set using a similarity value between the multiple data representations,
wherein the similarity value of the multiple data representations satisfies a constraint set;
generate, using a second AI model set, a reconfiguration command set configured to remove the at least one particular content set from the plurality of content sets; and
cause execution of the reconfiguration command set on the dataset to modify a portion of the dataset to remove the at least one particular content set from the plurality of content sets.
2 . The system of claim 1 , wherein the first AI model set is further configured to identify at least one content conflict between one or more pairs of data representations within the multiple data representations by:
linking a first data representation of the multiple data representations to (a) a topic and (b) a first information set,
linking a second data representation of the multiple data representations to (a) the topic and (b) a second information set, and
identify the content conflict between the first and second information sets by:
comparing the first and second information sets to determine a degree of similarity between the first and second information sets, and
determining that the degree of similarity fails one or more criterion.
3 . The system of claim 1 , wherein the system is further caused to:
cause transmission of a report indicating one or more of (a) the identified at least one particular content set or (b) the reconfiguration command set.
4 . The system of claim 1 , wherein the constraint set includes a threshold generated using a separate AI model that evaluates a degree of satisfaction of the threshold against a set of performance metrics.
5 . The system of claim 1 , wherein the at least one particular content set includes at least one duplicate content set.
6 . A non-transitory computer-readable storage medium comprising instructions stored thereon, wherein the instructions when executed by at least one data processor of a system, cause the system to:
obtain a dataset that comprises a plurality of content sets;
determine multiple data representations defining the plurality of content sets by:
linking each content set of the plurality of content sets into one or more clusters, and
for each particular cluster, generate a respective data representation of the multiple data representations by summarizing respective content sets of the particular cluster;
identify, using a first artificial intelligence (AI) model set, at least one particular content set using a similarity value between the multiple data representations,
wherein the similarity value of the multiple data representations satisfies a constraint set;
generate, using a second AI model set, a reconfiguration command set configured to remove the at least one particular content set from the plurality of content sets; and
cause execution of the reconfiguration command set on the dataset to modify a portion of the dataset to remove the at least one particular content set from the plurality of content sets.
7 . The non-transitory computer-readable storage medium of claim 6 , wherein the reconfiguration command set is generated based on a predefined rule set.
8 . The non-transitory computer-readable storage medium of claim 6 , wherein one or more AI models in the first AI model set is an AI agent.
9 . The non-transitory computer-readable storage medium of claim 6 , wherein the instructions further cause the system to:
detect a data linkage between portions of the dataset based on the one or more clusters; and
generate a representation of the data linkage in a form of one or more of: a knowledge graph, a tree structure, or a table.
10 . The non-transitory computer-readable storage medium of claim 6 , wherein the instructions further cause the system to:
map one or more content sets within the dataset to one or more organizational references associated with an organization.
11 . The non-transitory computer-readable storage medium of claim 6 , wherein the instructions further cause the system to:
receive feedback via an interface; and
modify one or more reconfiguration commands within the reconfiguration command set in accordance with the feedback.
12 . The non-transitory computer-readable storage medium of claim 6 , wherein the constraint set is associated with a majority vote of multiple AI models within the second AI model set.
13 . The non-transitory computer-readable storage medium of claim 6 ,
wherein the at least one particular content set comprises at least one content conflict between a first data representation and a second data representation, and
wherein the reconfiguration command set includes updating the first data representation or the second data representation.
14 . A computer-implemented method comprising:
obtaining a dataset that comprises a plurality of content sets;
determining multiple data representations defining the plurality of content sets;
identifying, using a first artificial intelligence (AI) model set, at least one particular content set using a similarity value between the multiple data representations,
wherein the similarity value of the multiple data representations satisfies a constraint set;
generating, using a second AI model set, a reconfiguration command set configured to remove the at least one particular content set from the plurality of content sets; and
causing execution of the reconfiguration command set on the dataset to modify a portion of the dataset to remove the at least one particular content set from the plurality of content sets.
15 . The computer-implemented method of claim 14 , wherein at least one of the first AI model set or the second AI model set includes an agent.
16 . The computer-implemented method of claim 14 , wherein the at least one particular content set comprises at least one of a content conflict or a duplicate content.
17 . The computer-implemented method of claim 14 , further comprising:
generating a compliance report indicating a compliance status of the dataset with predefined criteria.
18 . The computer-implemented method of claim 14 , further comprising:
maintaining a history of one or more data modifications to the dataset,
wherein each data modification is determined based on the modified portion of the dataset.
19 . The computer-implemented method of claim 18 , further comprising:
identifying one or more events associated with the at least one particular content set using one or more correlations between values of different data variables in the dataset.
20 . The computer-implemented method of claim 14 ,
wherein the modified portion of the dataset includes enriched data, and
wherein the enriched data replaces one or more missing values within the dataset based on reference data sources.