VIDEO UNDERSTANDING PLATFORM
In one embodiment, a method includes accessing a video-content object, determining a first feature vector representing the video-content object using a first recognition module of a first type based on an object in the video-content object, and determining a second feature vector representing the video-content object using a second recognition module of a second type based on the first feature vector. The first type is different from the second type. The method also includes determining a context of the video-content object based on the second feature vector.
1 . A method comprising:
by one or more computing devices, accessing a video-content object;
by one or more computing devices, determining a first feature vector representing the video-content object using a first recognition module of a first type based on an object in the video-content object;
by one or more computing devices, determining a second feature vector representing the video-content object using a second recognition module of a second type based on the first feature vector, wherein the first type is different from the second type; and
by one or more computing devices, determining a context of the video-content object based on the second feature vector.
2 . The method of claim 1 , wherein:
the first recognition module is an audio-recognition module;
the first feature vector represents a predicted transcript of the video-content object, wherein the transcript comprises text; and
the second recognition module is a text-recognition module.
3 . The method of claim 1 , wherein:
the first recognition module is a video-recognition module and the second recognition module is a text-recognition module;
the first recognition module is a video-recognition module and the second recognition module is an audio-recognition module;
the first recognition module is a text-recognition module and the second recognition module is a video-recognition module;
the first recognition module is a text-recognition module and the second recognition module is an audio-recognition module;
the first recognition module is an audio-recognition module and the second recognition module is a video-recognition module; or
the first recognition module is an audio-recognition module and the second recognition module is a text-recognition module.
4 . The method of claim 1 , wherein:
the video-content object corresponds to a node in a social graph of a social-networking system;
the social graph comprises a plurality of nodes and edges connecting the nodes; and
the context of the video-content object is determined based on social-graph information based at least in part on one or more nodes or edges connected to the node corresponding to the video-content object, in addition to the second feature vector.
5 . The method of claim 1 , wherein:
the video-content object comprises frames and audio and is associated with text; and
the object in the video-content object is one of:
one or more of the frames;
one or more portions of the audio; or
at least some of the text.
6 . The method of claim 1 , wherein:
the first recognition module is a video-recognition module;
the first feature vector represents an intermediate output prediction; and
the second recognition module is an audio-recognition module.
7 . The method of claim 1 , further comprising:
by one or more computing devices, determining a third feature vector representing the video-content object using a third recognition module of a third type based on at least one of the first feature vector and the second feature vector, wherein the third type is different from the first and second types; and
by one or more computing devices, determining a context of the video-content object based on the third feature vector.
8 . The method of claim 1 , wherein determining the first feature vector comprises:
extracting at least one feature from each frame of a first set of frames of the video-content object to generate a first set of feature vectors; and
polling two or more of the first set of feature vectors to generate the first feature vector.
9 . One or more computer-readable non-transitory storage media embodying software that is operable when executed to:
access a video-content object;
determine a first feature vector representing the video-content object using a first recognition module of a first type based on an object in the video-content object;
determine a second feature vector representing the video-content object using a second recognition module of a second type based on the first feature vector, wherein the first type is different from the second type; and
determine a context of the video-content object based on the second feature vector.
10 . The media of claim 9 , wherein:
the first recognition module is an audio-recognition module;
the first feature vector represents a predicted transcript of the video-content object, wherein the transcript comprises text; and
the second recognition module is a text-recognition module.
11 . The media of claim 9 , wherein:
the first recognition module is a video-recognition module and the second recognition module is a text-recognition module;
the first recognition module is a video-recognition module and the second recognition module is an audio-recognition module;
the first recognition module is a text-recognition module and the second recognition module is a video-recognition module;
the first recognition module is a text-recognition module and the second recognition module is an audio-recognition module;
the first recognition module is an audio-recognition module and the second recognition module is a video-recognition module; or
the first recognition module is an audio-recognition module and the second recognition module is a text-recognition module.
12 . The media of claim 9 , wherein:
the video-content object corresponds to a node in a social graph of a social-networking system;
the social graph comprises a plurality of nodes and edges connecting the nodes; and
the context of the video-content object is determined based on social-graph information based at least in part on one or more nodes or edges connected to the node corresponding to the video-content object, in addition to the second feature vector.
13 . The media of claim 9 , wherein:
the video-content object comprises frames and audio and is associated with text; and
the object in the video-content object is one of:
one or more of the frames;
one or more portions of the audio; or
at least some of the text.
14 . The media of claim 9 , wherein:
the first recognition module is a video-recognition module;
the first feature vector represents an intermediate output prediction; and
the second recognition module is an audio-recognition module.
15 . The media of claim 9 , wherein the software is further operable when executed to:
determine a third feature vector representing the video-content object using a third recognition module of a third type based on at least one of the first feature vector and the second feature vector, wherein the third type is different from the first and second types; and
determine a context of the video-content object based on the third feature vector.
16 . The media of claim 9 , wherein the software is operable to determine the first feature vector by:
extracting at least one feature from each frame of a first set of frames of the video-content object to generate a first set of feature vectors; and
polling two or more of the first set of feature vectors to generate the first feature vector.
17 . A system comprising:
one or more processors; and
a memory coupled to the processors and comprising instructions operable when executed by the processors to cause the processors to:
access a video-content object;
determine a first feature vector representing the video-content object using a first recognition module of a first type based on an object in the video-content object;
determine a second feature vector representing the video-content object using a second recognition module of a second type based on the first feature vector, wherein the first type is different from the second type; and
determine a context of the video-content object based on the second feature vector.
18 . The system of claim 17 , wherein:
the first recognition module is an audio-recognition module;
the first feature vector represents a predicted transcript of the video-content object, wherein the transcript comprises text; and
the second recognition module is a text-recognition module.
19 . The system of claim 17 , wherein:
the first recognition module is a video-recognition module and the second recognition module is a text-recognition module;
the first recognition module is a video-recognition module and the second recognition module is an audio-recognition module;
the first recognition module is a text-recognition module and the second recognition module is a video-recognition module;
the first recognition module is a text-recognition module and the second recognition module is an audio-recognition module;
the first recognition module is an audio-recognition module and the second recognition module is a video-recognition module; or
the first recognition module is an audio-recognition module and the second recognition module is a text-recognition module.
20 . The method of claim 17 , wherein:
the video-content object corresponds to a node in a social graph of a social-networking system;
the social graph comprises a plurality of nodes and edges connecting the nodes; and
the context of the video-content object is determined based on social-graph information based at least in part on one or more nodes or edges connected to the node corresponding to the video-content object, in addition to the second feature vector.