SYSTEM AND METHOD FOR CROSS-CLOUD TOPIC MATCHING
A system and method for cross-cloud topic matching. The method comprises: receiving unstructured data as a collection of unstructured data portions; analyzing each of the unstructured data portions to identify at least one tag in each unstructured data portion; determining a topic for each unstructured data portion based on the identified at least one tag; analyzing the determined topics to identify at least one match between the topics; and generating at least one searchable term respective of the at least one match.
1 . A method for cross-cloud topic matching, comprising:
receiving unstructured data as a collection of unstructured data portions;
analyzing each of the unstructured data portions to identify at least one tag in each unstructured data portion;
determining a topic for each unstructured data portion based on the identified at least one tag;
analyzing the determined topics to identify at least one match between the topics; and
generating at least one searchable term respective of the at least one match.
2 . The method of claim 1 , wherein generating at least one searchable term respective of the at least one match further comprises:
correlating the determined topics; and
selecting a most descriptive term based on the correlating.
3 . The method of claim 1 , wherein analyzing the determined topics to identify at least one match between the topics further comprises:
determining if any portions of the topics are related; and
upon determining that portions of the topics are related, identifying a match between the related portions.
4 . The method of claim 1 , wherein analyzing each unstructured data portion to identify at least one tag in each portion further comprises:
determining, for each unstructured data portion, at least one textual term within the unstructured data portion;
comparing each of the at least one textual term with a plurality of predetermined textual terms, wherein a tag is assigned to each of the predetermined textual terms; and
upon determining that a textual term of the at least one textual term matches one of the predetermined textual terms, identifying the tag assigned to the matching predetermined textual term.
5 . The method of claim 4 , further comprising:
upon determining that none of the at least one textual term matches one of the predetermined textual terms, generating a tag based for each of the at least one textual term.
6 . The method of claim 4 , wherein analyzing each unstructured data portion to identify at least one tag in each portion further comprises:
filtering insignificant textual terms from the at least one textual term.
7 . The method of claim 1 , wherein determining a topic for each unstructured data portion based on the identified at least one tag further comprises:
comparing the at least one tag of each unstructured data portion to a plurality of combinations of tags to determine at least one context; and
determining the topic based on the context.
8 . The method of claim 7 , wherein the at least one context is further based on a source of the unstructured data.
9 . The method of claim 1 , wherein the unstructured data is any of: a document, a message, an image, a video clip, and a calendar event description.
10 . A non-transitory computer readable medium having stored thereon instructions for causing one or more processing units to execute the method according to claim 1 .
11 . A system for cross-cloud topic matching, comprising:
a processing unit; and
a memory, the memory containing instructions that, when executed by the processing unit, configure the system to:
receive unstructured data including at least one unstructured data portion;
analyze each unstructured data portion to identify at least one tag in each unstructured data portion;
determine a topic for each unstructured data portion based on the identified at least one tag;
analyze the determined topics to identify at least one match between the topics; and
generate at least one searchable term respective of the at least one match.
12 . The system of claim 11 , wherein the system is further configured to:
correlate the determined topics; and
select a most descriptive term based on the correlating.
13 . The system of claim 11 , wherein the system is further configured to:
determine if any portions of the topics are related; and
upon determining that portions of the topics are related, identify a match between the related portions.
14 . The system of claim 11 , wherein the system is further configured to:
determine, for each unstructured data portion, at least one textual term within the unstructured data portion;
compare each of the at least one textual term with a plurality of predetermined textual terms, wherein a tag is assigned to each of the predetermined textual terms; and
upon determining that a textual term of the at least one textual term matches one of the predetermined textual terms, identify the tag assigned to the matching predetermined textual term.
15 . The system of claim 14 , wherein the system is further configured to:
upon determining that none of the at least one textual term matches one of the predetermined textual terms, generate a tag based for each of the at least one textual term.
16 . The system of claim 14 , wherein the system is further configured to:
filter insignificant textual terms from the at least one textual term.
17 . The system of claim 11 , wherein the system is further configured to:
compare the at least one tag of each unstructured data portion to a plurality of combinations of tags to determine at least one context; and
determine the topic based on the context.
18 . The system of claim 17 , wherein the at least one context is further based on a source of the unstructured data.
19 . The system of claim 11 , wherein the unstructured data is any of: a document, a message, an image, a video clip, and a calendar event description.
20 . An agent for cross-cloud topic matching, comprising:
a network interface for receiving and sending unstructured data, the unstructured data including at least one portion of unstructured data;
an analyzing unit for identifying at least one tag respective of each portion of the unstructured data;
a topic determination unit for generating at least one topic respective of each portion of unstructured data; and
a term generator for generating at least one searchable term based on matches between the topics.