IP Library › Granted Patent US 10,984,246
Granted Patent B2
US 10,984,246 · App. 16/352,605 · Granted Apr 20, 2021

Gating model for video analysis

Inventors: Sharadh Ramaswamy (Newark, CA); Sourish Chaudhuri (San Francisco, CA); Joseph Roth (San Francisco, CA)
Assignee: Google LLC
G06K9/00718G06F40/169G06K9/00744G06K9/6256G06N3/08
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 10,984,246
App. No.
16/352,605
Granted
Apr 20, 2021
Kind
B2
Abstract

Implementations described herein relate to methods, devices, and computer-readable media to perform gating for video analysis. In some implementations, a computer-implemented method includes obtaining a video comprising a plurality of frames and corresponding audio. The method further includes performing sampling to select a subset of the plurality of frames based on a target frame rate and extracting a respective audio spectrogram for each frame in the subset of the plurality of frames. The method further includes reducing resolution of the subset of the plurality of frames. The method further includes applying a machine-learning based gating model to the subset of the plurality of frames and corresponding audio spectrograms and obtaining, as output of the gating model, an indication of whether to analyze the video to add one or more video annotations.

Claims (62)

1. A computer-implemented method comprising:

obtaining a video comprising a plurality of frames and corresponding audio;

performing sampling to select a subset of the plurality of frames based on a target frame rate that is less than or equal to a frame rate of the video;

extracting a respective audio spectrogram for each frame in the subset of the plurality of frames;

reducing resolution of the subset of the plurality of frames;

after reducing the resolution, dividing the video into a plurality of segments, each segment including multiple frames;

applying a machine-learning based gating model to the subset of the plurality of frames and corresponding audio spectrograms, wherein applying the gating model is performed iteratively over the plurality of segments in sequence; and

obtaining, as output of the gating model, an indication of whether to analyze the video to add one or more video annotations, wherein the indication is generated at each iteration and wherein if the indication at a particular iteration is that the video is to be analyzed, application of the gating model is terminated such that one or more of the plurality of segments are excluded.

2. The computer-implemented method of claim 1 , wherein each segment of the plurality of segments overlaps with another segment of the plurality of segments.

3. The computer-implemented method of claim 1 , wherein the gating model is trained to determine whether a particular feature is present in input videos provided to the gating model.

4. The computer-implemented method of claim 3 , wherein the particular feature includes at least one of a human face, a type of object, a type of movement, or a type of audio.

5. The computer-implemented method of claim 1 , wherein applying the gating model comprises:

applying a first model that determines a likelihood that a particular feature is present; and

applying a second model that receives as input the likelihood that the particular feature is present and generates the indication of whether to analyze the video.

6. The computer-implemented method of claim 5 , wherein the first model includes:

a first convolutional neural network that includes a plurality of layers, trained to analyze video;

a second convolutional neural network that includes a plurality of layers, trained to analyze audio; and

a fusion network that includes a plurality of layers, that receives output of the first convolutional neural network and the second convolutional neural network as inputs, and provides the likelihood that the particular feature is present to the second model.

7. The computer-implemented method of claim 5 , wherein the second model is implemented using one or more of heuristics, a recurrent neural network, or a Markov chain analysis technique.

8. The computer-implemented method of claim 5 , further comprising providing an additional input to the second model, wherein the additional input includes one or more of:

identification of a portion of a particular frame of the subset of the plurality of frames in which the particular feature is detected to be present,

a duration of time in which the particular feature appears in the subset of the plurality of frames, or

heuristics regarding early termination,

and wherein the second model utilizes the additional input to generate the indication.

9. The computer-implemented method of claim 1 , further comprising, when the indication is to analyze the video, programmatically analyzing the video to add the one or more video annotations, wherein the video annotations comprise one or more labels that are indicative of presence in the video of one or more of a face, a particular type of object, a particular type of movement, or a particular type of audio.

10. A computing device comprising:

a processor; and

a memory, with instructions stored thereon that, when executed by the processor cause the processor to perform operations comprising:

obtaining a video comprising a plurality of frames and corresponding audio;

performing sampling to select a subset of the plurality of frames based on a target frame rate that is less than or equal to a frame rate of the video;

extracting a respective audio spectrogram for each frame in the subset of the plurality of frames;

reducing resolution of the subset of the plurality of frames;

after reducing the resolution, dividing the video into a plurality of segments, each segment including multiple frames;

applying a machine-learning based gating model to the subset of the plurality of frames and corresponding audio spectrograms, wherein applying the gating model is performed iteratively over the plurality of segments in sequence; and

obtaining, as output of the gating model, an indication of whether to analyze the video to add one or more video annotations, wherein the indication is generated at each iteration and wherein if the indication at a particular iteration is that the video is to be analyzed, application of the gating model is terminated such that one or more of the plurality of segments are excluded.

11. The computing device of claim 10 , wherein the operation of dividing the video into the plurality of segments is performed such that each segment of the plurality of segments overlaps with another segment of the plurality of segments.

12. The computing device of claim 10 , wherein the gating model is trained to determine whether a particular feature is present in input videos provided to the gating model.

13. The computing device of claim 10 , wherein the particular feature includes at least one of a human face, a type of object, a type of movement, or a type of audio.

14. The computing device of claim 10 , wherein applying the gating model comprises:

applying a first model that determines a likelihood that a particular feature is present; and

applying a second model that receives as input the likelihood that the particular feature is present and generates the indication of whether to analyze the video.

15. The computing device of claim 14 , wherein the first model includes:

a first convolutional neural network that includes a plurality of layers, trained to analyze video;

a second convolutional neural network that includes a plurality of layers, trained to analyze audio; and

a fusion network that includes a plurality of layers, that receives output of the first convolutional neural network and the second convolutional neural network as inputs, and provides the likelihood that the particular feature is present to the second model.

16. A non-transitory computer-readable medium with instructions stored thereon that, when executed by a computer, cause the computer to perform comprising:

obtaining a video comprising a plurality of frames and corresponding audio;

performing sampling to select a subset of the plurality of frames based on a target frame rate that is less than or equal to a frame rate of the video;

extracting a respective audio spectrogram for each frame in the subset of the plurality of frames;

reducing resolution of the subset of the plurality of frames;

after reducing the resolution, dividing the video into a plurality of segments, each segment including multiple frames;

applying a machine-learning based gating model to the subset of the plurality of frames and corresponding audio spectrograms, wherein applying the gating model is performed iteratively over the plurality of segments in sequence; and

obtaining, as output of the gating model, an indication of whether to analyze the video to add one or more video annotations, wherein the indication is generated at each iteration and wherein if the indication at a particular iteration is that the video is to be analyzed, application of the gating model is terminated such that one or more of the plurality of segments are excluded.

17. The non-transitory computer-readable medium of claim 16 , wherein the operation of dividing the video into the plurality of segments is performed such that each segment of the plurality of segments overlaps with another segment of the plurality of segments.

18. The non-transitory computer-readable medium of claim 16 , wherein each segment of the plurality of segments overlaps with another segment of the plurality of segments.

19. The non-transitory computer-readable medium of claim 16 , wherein the operation of applying the gating model comprises:

applying a first model that determines a likelihood that a particular feature is present; and

applying a second model that receives as input the likelihood that the particular feature is present and generates the indication of whether to analyze the video.

20. The non-transitory computer-readable medium of claim 19 , wherein the first model includes:

a first convolutional neural network that includes a plurality of layers, trained to analyze video;

a second convolutional neural network that includes a plurality of layers, trained to analyze audio; and

a fusion network that includes a plurality of layers, that receives output of the first convolutional neural network and the second convolutional neural network as inputs, and provides the likelihood that the particular feature is present to the second model.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Mar 13, 2019
From: RAMASWAMY, SHARADH; CHAUDHURI, SOURISH; ROTH, JOSEPH
To: GOOGLE LLC
Reel/Frame 048590/0762 →
Continuity (1)
Related Publication 20200293783A1 · Sep 17, 2020
Cited By (1)
US 12,307,756