IP Library Patent Application 16114059
Patent Application
App. No. 16/114,059

VIDEO UNDERSTANDING PLATFORM

Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US None
App. No.
16/114,059
Abstract

In one embodiment, a method includes accessing a video-content object, determining a first feature vector representing the video-content object using a first recognition module of a first type based on an object in the video-content object, and determining a second feature vector representing the video-content object using a second recognition module of a second type based on the first feature vector. The first type is different from the second type. The method also includes determining a context of the video-content object based on the second feature vector.

Claims (94)

1 . A method comprising:

by one or more computing devices, accessing a video-content object;

by one or more computing devices, determining a first feature vector representing the video-content object using a first recognition module of a first type based on an object in the video-content object;

by one or more computing devices, determining a second feature vector representing the video-content object using a second recognition module of a second type based on the first feature vector, wherein the first type is different from the second type; and

by one or more computing devices, determining a context of the video-content object based on the second feature vector.

2 . The method of claim 1 , wherein:

the first recognition module is an audio-recognition module;

the first feature vector represents a predicted transcript of the video-content object, wherein the transcript comprises text; and

the second recognition module is a text-recognition module.

3 . The method of claim 1 , wherein:

the first recognition module is a video-recognition module and the second recognition module is a text-recognition module;

the first recognition module is a video-recognition module and the second recognition module is an audio-recognition module;

the first recognition module is a text-recognition module and the second recognition module is a video-recognition module;

the first recognition module is a text-recognition module and the second recognition module is an audio-recognition module;

the first recognition module is an audio-recognition module and the second recognition module is a video-recognition module; or

the first recognition module is an audio-recognition module and the second recognition module is a text-recognition module.

4 . The method of claim 1 , wherein:

the video-content object corresponds to a node in a social graph of a social-networking system;

the social graph comprises a plurality of nodes and edges connecting the nodes; and

the context of the video-content object is determined based on social-graph information based at least in part on one or more nodes or edges connected to the node corresponding to the video-content object, in addition to the second feature vector.

5 . The method of claim 1 , wherein:

the video-content object comprises frames and audio and is associated with text; and

the object in the video-content object is one of:

one or more of the frames;

one or more portions of the audio; or

at least some of the text.

6 . The method of claim 1 , wherein:

the first recognition module is a video-recognition module;

the first feature vector represents an intermediate output prediction; and

the second recognition module is an audio-recognition module.

7 . The method of claim 1 , further comprising:

by one or more computing devices, determining a third feature vector representing the video-content object using a third recognition module of a third type based on at least one of the first feature vector and the second feature vector, wherein the third type is different from the first and second types; and

by one or more computing devices, determining a context of the video-content object based on the third feature vector.

8 . The method of claim 1 , wherein determining the first feature vector comprises:

extracting at least one feature from each frame of a first set of frames of the video-content object to generate a first set of feature vectors; and

polling two or more of the first set of feature vectors to generate the first feature vector.

9 . One or more computer-readable non-transitory storage media embodying software that is operable when executed to:

access a video-content object;

determine a first feature vector representing the video-content object using a first recognition module of a first type based on an object in the video-content object;

determine a second feature vector representing the video-content object using a second recognition module of a second type based on the first feature vector, wherein the first type is different from the second type; and

determine a context of the video-content object based on the second feature vector.

10 . The media of claim 9 , wherein:

the first recognition module is an audio-recognition module;

the first feature vector represents a predicted transcript of the video-content object, wherein the transcript comprises text; and

the second recognition module is a text-recognition module.

11 . The media of claim 9 , wherein:

the first recognition module is a video-recognition module and the second recognition module is a text-recognition module;

the first recognition module is a video-recognition module and the second recognition module is an audio-recognition module;

the first recognition module is a text-recognition module and the second recognition module is a video-recognition module;

the first recognition module is a text-recognition module and the second recognition module is an audio-recognition module;

the first recognition module is an audio-recognition module and the second recognition module is a video-recognition module; or

the first recognition module is an audio-recognition module and the second recognition module is a text-recognition module.

12 . The media of claim 9 , wherein:

the video-content object corresponds to a node in a social graph of a social-networking system;

the social graph comprises a plurality of nodes and edges connecting the nodes; and

the context of the video-content object is determined based on social-graph information based at least in part on one or more nodes or edges connected to the node corresponding to the video-content object, in addition to the second feature vector.

13 . The media of claim 9 , wherein:

the video-content object comprises frames and audio and is associated with text; and

the object in the video-content object is one of:

one or more of the frames;

one or more portions of the audio; or

at least some of the text.

14 . The media of claim 9 , wherein:

the first recognition module is a video-recognition module;

the first feature vector represents an intermediate output prediction; and

the second recognition module is an audio-recognition module.

15 . The media of claim 9 , wherein the software is further operable when executed to:

determine a third feature vector representing the video-content object using a third recognition module of a third type based on at least one of the first feature vector and the second feature vector, wherein the third type is different from the first and second types; and

determine a context of the video-content object based on the third feature vector.

16 . The media of claim 9 , wherein the software is operable to determine the first feature vector by:

extracting at least one feature from each frame of a first set of frames of the video-content object to generate a first set of feature vectors; and

polling two or more of the first set of feature vectors to generate the first feature vector.

17 . A system comprising:

one or more processors; and

a memory coupled to the processors and comprising instructions operable when executed by the processors to cause the processors to:

access a video-content object;

determine a first feature vector representing the video-content object using a first recognition module of a first type based on an object in the video-content object;

determine a second feature vector representing the video-content object using a second recognition module of a second type based on the first feature vector, wherein the first type is different from the second type; and

determine a context of the video-content object based on the second feature vector.

18 . The system of claim 17 , wherein:

the first recognition module is an audio-recognition module;

the first feature vector represents a predicted transcript of the video-content object, wherein the transcript comprises text; and

the second recognition module is a text-recognition module.

19 . The system of claim 17 , wherein:

the first recognition module is a video-recognition module and the second recognition module is a text-recognition module;

the first recognition module is a video-recognition module and the second recognition module is an audio-recognition module;

the first recognition module is a text-recognition module and the second recognition module is a video-recognition module;

the first recognition module is a text-recognition module and the second recognition module is an audio-recognition module;

the first recognition module is an audio-recognition module and the second recognition module is a video-recognition module; or

the first recognition module is an audio-recognition module and the second recognition module is a text-recognition module.

20 . The method of claim 17 , wherein:

the video-content object corresponds to a node in a social graph of a social-networking system;

the social graph comprises a plurality of nodes and edges connecting the nodes; and

the context of the video-content object is determined based on social-graph information based at least in part on one or more nodes or edges connected to the node corresponding to the video-content object, in addition to the second feature vector.

Assignments (2)
CHANGE OF NAME Recorded Jan 3, 2022
From: FACEBOOK, INC.
To: META PLATFORMS, INC.
Reel/Frame 058605/0840 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Dec 27, 2021
From: PALURI, BALMANOHAR; DUMOULIN, BENOIT F.; DENG, MERLYN; PHILIP, REENA; GARCIA, DARIO GARCIA
To: FACEBOOK, INC.
Reel/Frame 058485/0343 →