SYSTEM AND METHOD FOR DETERMINATION OF CAUSALITY BASED ON BIG DATA ANALYSIS
A method and system for determining causality based on big data analysis are provided. The method comprises extracting a plurality of unstructured data elements from a plurality of unstructured big data sources; generating at least one signature for each of the plurality of unstructured data elements; identifying at least one common pattern within the signatures of the plurality of unstructured data elements; matching the at least one common pattern to at least one hypothesis by comparing at least one signature of the common pattern to at least one hypothesis; and determining the causality of the at least one common pattern based on the at least one hypothesis matching the at least one common pattern.
1 . A method for determining causality based on big data analysis, comprising:
extracting a plurality of unstructured data elements from a plurality of unstructured big data sources;
generating at least one signature for each of the plurality of unstructured data elements;
identifying at least one common pattern within the signatures of the plurality of unstructured data elements;
matching the at least one common pattern to at least one hypothesis by comparing at least one signature of the common pattern to at least one hypothesis; and
determining the causality of the at least one common pattern based on the at least one hypothesis matching the at least one common pattern.
2 . The method of claim 1 , further comprising:
generating the at least one signature for the at least one hypothesis.
3 . The method of claim 1 , wherein identifying the at least one common pattern further comprises:
clustering the signatures identified to have common patterns; and
correlating the generated clusters to identify associations between their respective identified common patterns.
4 . The method of claim 1 , wherein determining the causality further comprises:
correlating the at least one hypothesis matching the at least one common pattern.
5 . The method of claim 1 , wherein determining the causality further comprises:
computing a matching score based on the comparisons of the signatures;
determining a hypothesis having the highest matching score to be the causality.
6 . The method of claim 1 , wherein the at least one hypothesis is textual content representing a series of natural events, wherein the at least one hypothesis is extracted from big data sources.
7 . The method of claim 1 , wherein the unstructured data element is at least one of: multimedia content, a document, metadata, a collection of health records, analog data, a file, unstructured text, and a web page.
8 . The method of claim 7 , wherein the multimedia content element is at least one of: an image, graphics, a video stream, a video clip, an audio stream, an audio clip, a video frame, a photograph, images of signals, combinations thereof, and portions thereof.
9 . A non-transitory computer readable medium having stored thereon instructions for causing one or more processing units to execute the method according to claim 1 .
10 . A system for determination of a causality based on big data analysis, comprising:
an interface to a network for connecting to a plurality of big data sources;
a processor;
a memory connected to the processor, the memory contains instructions that when executed by the processor, configure the system to:
extract a plurality of unstructured data elements from a plurality of unstructured big data sources;
generate at least one signature for each of the plurality of unstructured data elements;
identify at least one common pattern within the signatures of the plurality of unstructured data elements;
match the at least one common pattern to at least one hypothesis by comparing at least one signature of the common pattern to at least one hypothesis; and
determine the causality of the at least one common pattern based on the at least one hypothesis matching the at least one common pattern.
11 . The system of claim 10 , wherein the system is further configured to:
generate the at least one signature for the at least one hypothesis.
12 . The system of claim 10 , wherein the system is further configured to:
cluster the signatures identified to have common patterns; and
correlate the generated clusters to identify associations between their respective identified common patterns.
13 . The system of claim 10 , wherein the system is further configured to:
correlate the at least one hypothesis matching the at least one common pattern.
14 . The system of claim 10 , wherein the system is further configured to:
compute a matching score based on the comparisons of the signatures;
determine a hypothesis having the highest matching score to be the causality.
15 . The system of claim 10 , wherein the at least one hypothesis is textual content representing a series of natural events, wherein the at least one hypothesis is extracted from big data sources.
16 . The system of claim 10 , wherein the unstructured data element is at least one of: multimedia content, a document, metadata, a collection of health records, analog data, a file, unstructured text, and a web page.
17 . The system of claim 16 , wherein the multimedia content element is at least one of: an image, graphics, a video stream, a video clip, an audio stream, an audio clip, a video frame, a photograph, images of signals, combinations thereof, and portions thereof.
18 . The system of claim 10 , further comprising a database for storing the determined causality.
19 . A method for determining a probability of a hypothesis based on big data analysis, comprising:
receiving a request to check the probability of a hypothesis a hypothesizes;
generating at least one signature to the hypotheses;
crawling through a plurality of relevant big data sources to detect unstructured data elements;
generating at least one signature for each detected unstructured data element; and
determining the probability of the hypothesis respective of the generated signatures.
20 . The method of claim 19 , further comprising:
matching the at least one signature of the unstructured data element to the at least one generated for of the hypothesis; and
computing the probability based on the matching results.
21 . A non-transitory computer readable medium having stored thereon instructions for causing one or more processing units to execute the method according to claim 13 .