Systems and methods for identifying and analyzing risk events from data sources
Conventional methods of analyzing social media content involves performing sentimental analysis to understand related sentiment and effects of events on communities. However, such analysis may not be completely accurate and are prone to errors. Present disclosure provides system and method that identify and analyze risk events from data collected from various sources. Key phrases obtained from sources is received, pre-processed, and clustered accordingly. The clustering is performed based on frequency of incoming words. The clustered dataset obtained is classified into one or more categories based on a polarity score. Dataset of specific category (e.g., negative category dataset) is analysed to identify events and topics which are then grouped using an associated label to obtain grouped entities. Each entity is then ranked and assigned a risk score for identifying high-risk events which are then analyzed using simulation and optimization technique(s) and an explainability text for the analyzed risk events is generated.
1 . A processor implemented method, comprising:
receiving, via one or more hardware processors, one or more key phrases extracted from at least one source, wherein receiving the one or more key phrases comprises collecting data and content from social sensor sources including social media sites, news media outlets, satellite imagery repositories via an Application Programming Interface (API), wherein key and value pairs are defined in the API;
pre-processing, via the one or more hardware processors, the one or more key phrases to obtain a set of pre-processed key phrases, wherein the pre-processing includes a real-time event extraction and operatively communicating through the API, wherein an interface is facilitated to enter a predefined monitoring and a search criteria of the one or more key phrases, wherein the step of pre-processing includes creating a data frame, creating a configuration file, wherein the configuration file is based on the key and value pairs defined as a file name and a experiment name, wherein the set of pre-processed key phrases is cleansed before using by a system based on the key and value pairs defined including a text column name, a target variable, a polarity method, wherein the target variable and the polarity method is dependent on each other, wherein the target variable is used in presence of a labelled data and when the labelled data is absent then the polarity method is implemented, wherein data is labelled using the polarity method by assigning a polarity score, wherein based on the polarity score the data is further labelled automatically and the labelled data is used to train a supervised machine learning model;
clustering, via the one or more hardware processors, the set of pre-processed key phrases based on frequency of one or more incoming words comprised in the one or more key phrases to obtain a clustered dataset;
classifying, via the one or more hardware processors, the clustered dataset based on a polarity score into one or more categories;
identifying, via the one or more hardware processors, one or more events and one or more topics classified as a negative category amongst the one or more categories, wherein the one or more events are identified by analyzing extracted textual content and images occurring within a defined time window of a day using a burst-event detection approach based on a feature-pivot clustering technique;
identifying an active period of an incident event on a social-media platform by identifying a bursty feature over the defined time window using a feature-based Gaussian distribution, and grouping the identified bursty features based on similarity and determining one or more hot periods corresponding to the bursty feature, wherein a transfer-based framework is configured to identify the incident event and perform unsupervised pre-training and the machine-learning model is invoked for execution;
grouping, via the one or more hardware processors, the one or more events and the one or more topics using an associated label and a frequency of an incoming content stream from the at least one source to obtain a set of grouped entities;
generating, an alert by analyzing a largest group of data items having negative polarity, wherein the alerts are generated based on frequency of social media complaints;
ranking and assigning, via the one or more hardware processors, a risk score to each entity from the set of grouped entities based on one or more event ranking rules;
identifying, via the one or more hardware processors, one or more high-risk events from the set of grouped entities based on the assigned risk score;
analysing, via the one or more hardware processors, the one or more high-risk events by using one or more simulation and optimization techniques; and
generating, via the one or more hardware processors, an explainability text pertaining to the one or more high-risk events upon analysing the one or more high risk events, wherein the explainability text provides a reason or justification of the identified one or more high risk events, wherein the generated explainability is either local to a geography, applicable globally or time based.
2 . The processor implemented method of claim 1 , wherein the one or more topics are identified using one or more statistical techniques.
3 . The processor implemented method of claim 1 , wherein one or more events comprised in the set of grouped entities are ranked based on at least one of an associated criticality value and an associated importance value.
4 . The processor implemented method of claim 1 , wherein the risk score is assigned based on at least one of a frequency and a severity associated with each entity from the set of grouped entities.
5 . A system, comprising:
a memory storing instructions;
one or more communication interfaces; and
one or more hardware processors coupled to the memory via the one or more communication interfaces, wherein the one or more hardware processors are configured by the instructions to:
receive one or more key phrases extracted from at least one source, wherein receiving the one or more key phrases comprises collecting data and content from social sensor sources including social media sites, news media outlets, satellite imagery repositories via an Application Programming Interface (API), wherein key and value pairs are defined in the API;
pre-process the one or more key phrases to obtain a set of pre-processed key phrases, wherein the pre-processing includes real-time event extraction and operatively communicating through the API, wherein an interface is facilitated to enter a predefined monitoring and a search criteria of the one or more key phrases, wherein the step of preprocessing includes creating a data frame, creating a configuration file, wherein the configuration file is based on the key and value pairs defined as a file name and a experiment name, wherein the set of pre-processed key phrases is cleansed before using by a system based on the key and value pairs defined including a text column name, a target variable, a polarity method, wherein the target variable and the polarity method is dependent on each other, wherein the target variable is used in presence of a labelled data and when the labelled data is absent then the polarity method is implemented, wherein data is labelled using the polarity method by assigning a polarity score, wherein based on the polarity score the data is further labelled automatically and the labelled data is used to train a supervised machine learning model;
cluster the set of pre-processed key phrases based on frequency of one or more incoming words comprised in the one or more key phrases to obtain a clustered dataset;
classify the clustered dataset based on a polarity score into one or more categories;
identify one or more events and one or more topics classified as a negative category amongst the one or more categories, wherein the one or more events are identified by analyzing extracted textual content and images occurring within a defined time window of a day using a burst-event detection approach based on a feature-pivot clustering technique;
identify an active period of an incident event on a social-media platform by identifying a bursty feature over the defined time window using a feature-based Gaussian distribution, grouping the identified bursty features based on similarity and determining one or more hot periods corresponding to the bursty feature, wherein a transfer-based framework is configured to identify the incident event and perform unsupervised pre-training and the machine-learning model is invoked for execution;
group the one or more events and the one or more topics using an associated label and a frequency of an incoming content stream from the at least one source to obtain a set of grouped entities;
generate, an alert by analyzing a largest group of data items having negative polarity, wherein the alerts are generated based on frequency of social media complaints;
rank and assign a risk score to each entity from the set of grouped entities based on one or more event ranking rules;
identify one or more high-risk events from the set of grouped entities based on the assigned risk score;
analyze the one or more high-risk events by using one or more simulation and optimization techniques; and
generate, via the one or more hardware processors, an explainability text pertaining to the one or more high-risk events upon analysing the one or more high risk events, wherein the explainability text provides a reason or justification of the identified one or more high risk events, wherein the generated explainability is either local to a geography, applicable globally or time based.
6 . The system of claim 5 , wherein the one or more topics are identified using one or more statistical techniques.
7 . The system of claim 5 , wherein one or more events comprised in the set of grouped entities are ranked based on at least one of an associated criticality value and an associated importance value.
8 . The system of claim 5 , wherein the risk score is assigned based on at least one of a frequency and a severity associated with each entity from the set of grouped entities.
9 . One or more non-transitory machine-readable information storage mediums comprising one or more instructions which when executed by one or more hardware processors cause:
receiving one or more key phrases extracted from at least one source, wherein receiving the one or more key phrases comprises collecting data and content from social sensor sources including social media sites, news media outlets, satellite imagery repositories via an Application Programming Interface (API), wherein key and value pairs are defined in the API;
pre-processing, via the one or more hardware processors, the one or more key phrases to obtain a set of pre-processed key phrases, wherein the pre-processing includes a real-time event extraction and operatively communicating through the API, wherein an interface is facilitated to enter a predefined monitoring and a search criteria of the one or more key phrases, wherein the step of pre-processing includes creating a data frame, creating a configuration file, wherein the configuration file is based on the key and value pairs defined as a file name and a experiment name, wherein the set of pre-processed key phrases is cleansed before using by a system based on the key and value pairs defined including a text column name, a target variable, a polarity method, wherein the target variable and the polarity method is dependent on each other, wherein the target variable is used in presence of a labelled data and when the labelled data is absent then the polarity method is implemented, wherein data is labelled using the polarity method by assigning a polarity score, wherein based on the polarity score the data is further labelled automatically and the labelled data is used to train a supervised machine learning model;
clustering the set of pre-processed key phrases based on frequency of one or more incoming words comprised in the one or more key phrases to obtain a clustered dataset;
classifying the clustered dataset based on a polarity score into one or more categories;
identifying one or more events and one or more topics classified as a negative category amongst the one or more categories, wherein the one or more events are identified by analyzing extracted textual content and images occurring within a defined time window of a day using a burst-event detection approach based on a feature-pivot clustering technique;
identifying an active period of an incident event on a social-media platform comprises identifying a bursty feature over the defined time window using a feature-based Gaussian distribution, grouping the identified bursty features based on similarity and determining one or more hot periods corresponding to the burst, wherein a transfer-based framework is configured to identify the incident event and perform unsupervised pre-training and the machine-learning model is invoked for execution;
grouping the one or more events and the one or more topics using an associated label and a frequency of an incoming content stream from the at least one source to obtain a set of grouped entities;
generating, an alert by analyzing a largest group of data items having negative polarity, wherein the are alerts generated based on frequency of social media complaints;
ranking and assigning a risk score to each entity from the set of grouped entities based on one or more event ranking rules;
identifying one or more high-risk events from the set of grouped entities based on the assigned risk score; and
analysing the one or more high-risk events by using one or more simulation and optimization techniques; and
generating, via the one or more hardware processors, an explainability text pertaining to the one or more high-risk events upon analysing the one or more high risk events, wherein the explainability text provides a reason or justification of the identified one or more high risk events, wherein the generated explainability is either local to a geography, applicable globally or time based.
10 . The one or more non-transitory machine-readable information storage mediums of claim 9 , wherein the one or more topics are identified using one or more statistical techniques.
11 . The one or more non-transitory machine-readable information storage mediums of claim 9 , wherein one or more events comprised in the set of grouped entities are ranked based on at least one of an associated criticality value and an associated importance value.
12 . The one or more non-transitory machine-readable information storage mediums of claim 9 , wherein the risk score is assigned based on at least one of a frequency and a severity associated with each entity from the set of grouped entities.
13 . The processor implemented method of claim 1 , wherein the key and value pairs of the API comprises a consumer key, a consumer secret, an access token, access secret, days ago, a geo code, a longitude, a latitude, a radius, max tweets, keyword search and the experiment name, wherein the search criteria includes latitude, longitude, key words and time parameter, social news and media sites, wherein the pre-processing includes removal of special characters, uniform resource locators, wherein the one or more key phrases are tokenized to remove stop words, wherein the one or more events are identified using an entity, topic clustering and incidence peak detection, wherein risk is assessed as combination of hazard and vulnerability, wherein vulnerability is damages and values assessment of elements at risk, wherein a report for the incident event includes a start time of the event, a detection time, and a set of corresponding comments related to an event topic.
14 . The processor implemented method of claim 1 , wherein the transformer-based framework is configured to perform the unsupervised pre-training using a Bidirectional Encoder Representations from Transformers (BERT)-like framework, and the transformer-based framework employing scaled dot-product self-attention to identify correlations between input segments, thereby enabling real-time event detection, wherein the self-attention operation is used for downstream tasks including an incident event sentiment classification and prediction.