Methods and systems for synchronizing communication records in computer networks based on detecting patterns in categories of metadata
Methods and systems are described herein for synchronizing communication records in computer networks. For example, the methods and systems may determine whether or not a first communication relates to a second and generate a recommendation that the communications relate to a single communication. In particular, the methods and systems described herein describe synchronizing communication records in computer networks based on detecting patterns in categories of metadata. For example, the methods and systems retrieve specific types of metadata and compare this metadata between communications in order to synchronize and/or deduplicate them.
1 . A system for synchronizing communication records in computer networks based on detecting patterns in categories of metadata, the system comprising:
one or more processors; and
a non-transitory, computer-readable medium storing instructions that, when executed by the one or more processors, cause operations comprising:
training a machine learning model using a training data set;
generating a first set of patterns for a first set of metadata associated with a first set of communications and a second set of patterns for a second set of metadata associated with a second set of communications, wherein each pattern in the first set of patterns is associated with a corresponding communication in the first set of communications and wherein each pattern in the second set of patterns is associated with the corresponding communication in the second set of communications;
inputting the first set of patterns and the second set of patterns into the machine learning model, wherein the machine learning model determines a likelihood that a first pattern of the first set of patterns matches a second pattern from the second set of patterns, and wherein the first pattern is not identical to the second pattern;
based on receiving, from the machine learning model, matching patterns from the first set of patterns and the second set of patterns, identifying a first communication of the first set of communications that matches a second communication of the second set of communications;
generating an indication of a match based on determining that the first communication and the second communication correspond to a single communication;
deduplicating the first communication and the second communication based on the indication of the match; and
resolving the first communication and the second communication into communication counterparts.
2 . A method for synchronizing communication records in computer networks based on detecting patterns in categories of metadata, the method comprising:
generating a first set of patterns for a first set of metadata associated with a first set of communications and a second set of patterns for a second set of metadata associated with a second set of communications, wherein each pattern in the first set of patterns is associated with a corresponding communication in the first set of communications and wherein each pattern in the second set of patterns is associated with the corresponding communication in the second set of communications;
training a machine learning model using a training data set;
inputting the first set of patterns and the second set of patterns into the machine learning model, wherein the machine learning model determines a likelihood that a first pattern of the first set of patterns matches a second pattern from the second set of patterns, and wherein the first pattern is not identical to the second pattern;
based on receiving, from the machine learning model, matching patterns from the first set of patterns and the second set of patterns, identifying a first communication of the first set of communications that matches a second communication of the second set of communications;
generating an indication of a match based on determining that the first communication and the second communication correspond to a single communication;
deduplicating the first communication and the second communication based on the indication of the match; and
resolving the first communication and the second communication into communication counterparts.
3 . The method of claim 2 , wherein the first set of metadata comprises a respective set of field categories for each communication of the first set of communications and wherein a first value of a first field category of the respective set of field categories comprises text strings of numerical data, a second value of a second field category of the respective set of field categories comprises recurring text strings of alphanumeric text strings, and a third value of a third field category of the respective set of field categories comprises text strings of fifteen to twenty alphanumeric characters.
4 . The method of claim 2 , further comprising generating for display the single communication in a list of aggregated communications.
5 . The method of claim 2 , further comprising:
receiving a user input setting a predetermined time period; and
filtering the first set of metadata and the second set of metadata based on the predetermined time period.
6 . The method of claim 2 , wherein the first set of metadata and the second set of metadata are stored on a cloud-based big data framework.
7 . The method of claim 2 , wherein the second set of metadata is retrieved using a cloud-based managed cluster platform comprising clusters and nodes including a master node that manages a cluster by running software components to coordinate a distribution of data and tasks among other nodes for processing.
8 . The method of claim 2 , wherein the second set of metadata is retrieved using a cloud-based managed cluster platform comprising clusters and nodes including a core node that comprises software components that run tasks and store data in a Hadoop Distributed File System for a cluster.
9 . The method of claim 2 , wherein the second set of metadata is retrieved using a cloud-based managed cluster platform comprising clusters and nodes including a task node that runs tasks and does not store data in a Hadoop Distributed File System for a cluster.
10 . The method of claim 2 , wherein the first communication corresponds to a posted communication, and wherein the second communication corresponds to an authorization communication.
11 . The method of claim 2 , wherein resolving the first communication and the second communication into the communication counterparts comprises resolving the first communication and the second communication into the communication counterparts using a database join function.
12 . One or more non-transitory computer-readable media storing instructions that, when executed by one or more processors, cause operations comprising:
generating a first set of patterns for a first set of metadata associated with a first set of communications and a second set of patterns for a second set of metadata associated with a second set of communications, wherein each pattern in the first set of patterns is associated with a corresponding communication in the first set of communications and wherein each pattern in the second set of patterns is associated with the corresponding communication in the second set of communications;
training a machine learning model using a training data set;
inputting the first set of patterns and the second set of patterns into the machine learning model, wherein the machine learning model determines a likelihood that a first pattern of the first set of patterns matches a second pattern from the second set of patterns, and wherein the first pattern is not identical to the second pattern;
based on receiving, from the machine learning model, matching patterns from the first set of patterns and the second set of patterns, identifying a first communication of the first set of communications that matches a second communication of the second set of communications;
generating an indication of a match based on determining that the first communication and the second communication correspond to a single communication;
deduplicating the first communication and the second communication based on the indication of the match; and
resolving the first communication and the second communication into communication counterparts.
13 . The one or more non-transitory computer-readable media of claim 12 , wherein the first set of metadata comprises a respective set of field categories for each communication of the first set of communications and wherein a first value of a first field category of the respective set of field categories comprises text strings of numerical data, a second value of a second field category of the respective set of field categories comprises recurring text strings of alphanumeric text strings, and a third value of a third field category of the respective set of field categories comprises text strings of fifteen to twenty alphanumeric characters.
14 . The one or more non-transitory computer-readable media of claim 12 , further comprising generating for display the single communication in a list of aggregated communications.
15 . The one or more non-transitory computer-readable media of claim 12 , further comprising:
receiving a user input setting a predetermined time period; and
filtering the first set of metadata and the second set of metadata based on the predetermined time period.
16 . The one or more non-transitory computer-readable media of claim 12 , wherein the first set of metadata and the second set of metadata are stored on a cloud-based big data framework.
17 . The one or more non-transitory computer-readable media of claim 12 , wherein the second set of metadata is retrieved using a cloud-based managed cluster platform comprising clusters and nodes including a master node that manages a cluster by running software components to coordinate a distribution of data and tasks among other nodes for processing.
18 . The one or more non-transitory computer-readable media of claim 12 , wherein the second set of metadata is retrieved using a cloud-based managed cluster platform comprising clusters and nodes including a core node that comprises software components that run tasks and store data in a Hadoop Distributed File System for a cluster.
19 . The one or more non-transitory computer-readable media of claim 12 , wherein the second set of metadata is retrieved using a cloud-based managed cluster platform comprising clusters and nodes including a task node that runs tasks and does not store data in a Hadoop Distributed File System for a cluster.
20 . The one or more non-transitory computer-readable media of claim 12 , wherein resolving the first communication and the second communication into the communication counterparts comprises resolving the first communication and the second communication into the communication counterparts using a database join function.