Event detection system
An event detection system is configured to access a repository that contains a collection of media content. The media content may for example include images, videos, audio clips, and the like, wherein the media content comprises features that include: tags (e.g., hashtags or other similar mechanisms to label and sort content); captions that comprises one or more words or phrases; continuous numerical values; geolocation data (e.g., geo-hash, check-in data, coordinates); as well as temporal data (e.g., timestamps).
1 . A system comprising:
a memory; and
at least one hardware processor coupled to the memory and comprising instructions that causes the system to perform operations comprising:
accessing a repository that comprises a collection of media content that comprises metadata that comprises geolocation data, temporal data, and content feature data;
extracting the metadata from the collection of media content;
generating a graph comprising a first axis representing location values, a second axis representing temporal values, and a third axis representing feature values;
plotting representations of the media content from among the collection of media content on the graph based on the extracted metadata;
receiving, from a client device, a user input that specifies a location and time period filter;
filtering the plotted representations based on the location and time period filter to identify a subset of the media content;
determining a total number of occurrences of content features within the subset of the media content;
selecting a content feature from among the content features based on the total number of occurrences representing a most common content feature;
identifying one or more clusters among the filtered representations based on a clustering parameter that comprises a geographical threshold and a temporal threshold, the one or more clusters comprising groupings of media content having geolocation data and temporal data within the geographical threshold and the temporal threshold of the clustering parameter;
designating the selected content feature to the one or more clusters; and
causing display of a visualization of the one or more clusters that comprises a table at the client device, the visualization including an indication of the location and the time period specified by the user input, and the designated content feature associated with the selected content feature.
2 . The system of claim 1 , wherein the content features comprise text strings that include captions associated with the collection of media content.
3 . The system of claim 1 , wherein subset of the collection of media content includes a first media content, the graph comprises a first axis that represents location values, a second axis that represents temporal values, and a third axis that represents feature values, and the plotting the representation of the first media content upon the graph based on the metadata includes:
designating a content feature of the first media content to a position along the third axis; and
plotting a representation of the first media content upon the graph based on the metadata of the first media content and the position of the content feature along the third axis.
4 . The system of claim 1 , wherein the grouping the subset of collection of media content based on the metadata includes:
receiving a grouping parameter that comprises a temporal threshold and a geological threshold; and
grouping the subset of the collection of media content based on the grouping parameter.
5 . The system of claim 1 , wherein the subset of the collection of media content includes a first media content, the representation of the first media content is a first representation, and the operations further comprise:
plotting a second representation of a second media content upon the graph based on the metadata of the second media content;
determining that the second representation of the second media content and the first representation of the first media content are within a threshold distance on the graph; and
detecting a similarity between the first media content and the second media content based on the determining that the second representation of the second media content and the first representation of the first media content are within the threshold distance on the graph.
6 . A method comprising:
accessing a repository that comprises a collection of media content that comprises metadata that comprises geolocation data, temporal data, and content feature data;
extracting the metadata from the collection of media content;
generating a graph comprising a first axis representing location values, a second axis representing temporal values, and a third axis representing feature values;
plotting representations of the media content from among the collection of media content on the graph based on the extracted metadata;
receiving, from a client device, a user input that specifies a location and time period filter;
filtering the plotted representations based on the location and time period filter to identify a subset of the media content;
determining a total number of occurrences of content features within the subset of the media content;
selecting a content feature from among the content features based on the total number of occurrences representing a most common content feature;
identifying one or more clusters among the filtered representations based on a clustering parameter that comprises a geographical threshold and a temporal threshold, the one or more clusters comprising groupings of media content having geolocation data and temporal data within the geographical threshold and the temporal threshold of the clustering parameter;
designating the selected content feature to the one or more clusters; and
causing display of a visualization of the one or more clusters that comprises a table at the client device, the visualization including an indication of the location and the time period specified by the user input, and the designated content feature associated with the selected content feature.
7 . The method of claim 6 , wherein the content features comprise text strings that include captions associated with the collection of media content.
8 . The method of claim 6 , wherein subset of the collection of media content includes a first media content, the graph comprises a first axis that represents location values, a second axis that represents temporal values, and a third axis that represents feature values, and the plotting the representation of the first media content upon the graph based on the metadata includes:
designating a content feature of the first media content to a position along the third axis; and
plotting a representation of the first media content upon the graph based on the metadata of the first media content and the position of the content feature along the third axis.
9 . The method of claim 6 , wherein the grouping the subset of collection of media content based on the metadata includes:
receiving a grouping parameter that comprises a temporal threshold and a geological threshold; and
grouping the subset of the collection of media content based on the grouping parameter.
10 . The method of claim 6 , wherein the subset of the collection of media content includes a first media content, the representation of the first media content is a first representation, and the operations further comprise:
plotting a second representation of a second media content upon the graph based on the metadata of the second media content;
determining that the second representation of the second media content and the first representation of the first media content are within a threshold distance on the graph; and
detecting a similarity between the first media content and the second media content based on the determining that the second representation of the second media content and the first representation of the first media content are within the threshold distance on the graph.
11 . A non-transitory machine-readable storage medium comprising instructions that, when executed by one or more processors of a machine, cause the machine to perform operations comprising:
accessing a repository that comprises a collection of media content that comprises metadata that comprises geolocation data, temporal data, and content feature data;
extracting the metadata from the collection of media content;
generating a graph comprising a first axis representing location values, a second axis representing temporal values, and a third axis representing feature values;
plotting representations of the media content from among the collection of media content on the graph based on the extracted metadata;
receiving, from a client device, a user input that specifies a location and time period filter;
filtering the plotted representations based on the location and time period filter to identify a subset of the media content;
determining a total number of occurrences of content features within the subset of the media content;
selecting a content feature from among the content features based on the total number of occurrences representing a most common content feature;
identifying one or more clusters among the filtered representations based on a clustering parameter that comprises a geographical threshold and a temporal threshold, the one or more clusters comprising groupings of media content having geolocation data and temporal data within the geographical threshold and the temporal threshold of the clustering parameter;
designating the selected content feature to the one or more clusters; and
causing display of a visualization of the one or more clusters that comprises a table at the client device, the visualization including an indication of the location and the time period specified by the user input, and the designated content feature associated with the selected content feature.
12 . The non-transitory machine-readable storage medium of claim 11 , wherein the content features comprise text strings that include captions associated with the collection of media content.
13 . The non-transitory machine-readable storage medium of claim 11 , wherein subset of the collection of media content includes a first media content, the graph comprises a first axis that represents location values, a second axis that represents temporal values, and a third axis that represents feature values, and the plotting the representation of the first media content upon the graph based on the metadata includes:
designating a content feature of the first media content to a position along the third axis; and
plotting a representation of the first media content upon the graph based on the metadata of the first media content and the position of the content feature along the third axis.
14 . The non-transitory machine-readable storage medium of claim 11 , wherein the grouping the subset of collection of media content based on the metadata includes:
receiving a grouping parameter that comprises a temporal threshold and a geological threshold; and
grouping the subset of the collection of media content based on the grouping parameter.
15 . The non-transitory machine-readable storage medium of claim 11 , wherein the subset of the collection of media content includes a first media content, the representation of the first media content is a first representation, and the operations further comprise:
plotting a second representation of a second media content upon the graph based on the metadata of the second media content;
determining that the second representation of the second media content and the first representation of the first media content are within a threshold distance on the graph; and
detecting a similarity between the first media content and the second media content based on the determining that the second representation of the second media content and the first representation of the first media content are within the threshold distance on the graph.