IP Library Granted Patent US 12,080,067
Granted Patent B2
US 12,080,067 · App. 17/461,755 · Granted Sep 3, 2024

Classifying a video stream using a self-attention-based machine-learning model

Inventors: Gediminas Bertasius (Boston, MA); Heng Wang (Mountain View, CA); Lorenzo Torresani (Norwich, VT)
Assignee: Meta Platforms, Inc.
G06V20/41G06F18/21G06N20/00G06V10/56G06V10/751G06V20/48G06V10/759
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,080,067
App. No.
17/461,755
Granted
Sep 3, 2024
Kind
B2
Abstract

In one embodiment, a method includes accessing a stream of F video frames, where each of the F video frames includes N patches that are non-overlapping, generating an initial embedding vector for each of the N×F patches in the F video frames, generating a classification embedding by processing the generated N×F initial embedding vectors using a self-attention-based machine-learning model that computes a temporal attention and a spatial attention for each of the N×F patches, and determining a class of the stream of video frames based on the generated classification embedding.

Claims (42)

1. A method comprising, by a computing device:

accessing a stream of F video frames, wherein the F video frames comprises N patches that are non-overlapping;

generating, for N×F patches in the F video frames, an initial embedding vector;

generating a classification embedding by processing generated N×F initial embedding vectors using a self-attention-based machine-learning model that computes a temporal attention and a spatial attention for the N×F patches, wherein the temporal attention and the spatial attention of a particular patch in a particular video frame of the F video frames are computed by comparing information associated with the particular patch with information associated with (1) patches in other video frames of the F video frames and (2) other patches in the particular video frame; and

determining a class of the stream of video frames based on the generated classification embedding.

2. The method of claim 1 , wherein the generated N×F initial embedding vectors are provided to the self-attention-based machine-learning model as input, wherein the self-attention-based machine-learning model comprises L serial encoding blocks.

3. The method of claim 2 , wherein the computing device generates embedding vectors corresponding to the N×F patches at blocks of the L serial encoding blocks based on embedding vectors generated at a preceding block of the L serial encoding blocks.

4. The method of claim 3 , wherein generating an embedding vector corresponding to the particular patch in the particular video frame at block l of the L serial encoding blocks comprises:

computing a temporal attention based on comparisons between the embedding vector corresponding to the particular patch generated at block l- 1 of the L serial encoding blocks and other embedding vectors corresponding to the patches in the other video frames generated at block l- 1 ;

computing a spatial attention based on comparisons between the embedding vector corresponding to the particular patch generated at block l- 1 and additional embedding vectors corresponding to the other patches in the particular video frame generated at block l- 1 ; and

generating an embedding vector corresponding to the particular patch based on the computed temporal attention and the computed spatial attention.

5. The method of claim 4 , wherein the patches in the other video frames comprises patches at a spatial location of the particular patch in their corresponding video frames.

6. The method of claim 4 , wherein generating the embedding vector corresponding to the particular patch based on the computed temporal attention and the computed spatial attention comprises processing the computed temporal attention and the computed spatial attention using a multilayer perceptron (MLP).

7. The method of claim 3 , wherein the classification embedding is generated by taking a layer normalization on one or more of embedding vectors generated at a last block of the L serial encoding blocks.

8. The method of claim 3 , wherein the L serial encoding blocks are multi-headed.

9. The method of claim 1 , wherein determining a class of the stream of video frames comprises processing the classification embedding using a multilayer perceptron (MLP).

10. The method of claim 1 , wherein generating an initial embedding vector corresponding to a patch comprises:

creating a color embedding vector by multiplying color information of the patch with a color information embedding matrix; and

adding a positional embedding vector to the color embedding vector.

11. The method of claim 10 , wherein the color information embedding matrix is trained during a training procedure, and wherein the positional embedding vector is trained during the training procedure.

12. The method of claim 1 , wherein the self-attention-based machine-learning model is a Transformer network.

13. One or more computer-readable non-transitory storage media embodying software that is operable when executed to:

access a stream of F video frames, wherein the F video frames comprises N patches that are non-overlapping;

generate, for N×F patches in the F video frames, an initial embedding vector;

generate a classification embedding by processing generated N×F initial embedding vectors using a self-attention-based machine-learning model that computes a temporal attention and a spatial attention for the N×F patches, wherein the temporal attention and the spatial attention of a particular patch in a particular video frame of the F video frames are computed by comparing information associated with the particular patch with information associated with (1) patches in other video frames of the F video frames and (2) other patches in the particular video frame; and

determine a class of the stream of video frames based on the generated classification embedding.

14. The media of claim 13 , wherein the generated N×F initial embedding vectors are provided to the self-attention-based machine-learning model as input, wherein the self-attention-based machine-learning model comprises L serial encoding blocks.

15. The media of claim 14 , wherein the media embodying the software is further operable when executed to generate embedding vectors corresponding to the N×F patches at blocks of the L serial encoding blocks based on embedding vectors generated at a preceding block of the L serial encoding blocks.

16. The media of claim 15 , wherein generating an embedding vector corresponding to the particular patch in the particular video frame at block l of the L serial encoding blocks comprises:

computing a temporal attention based on comparisons between the embedding vector corresponding to the particular patch generated at block l- 1 of the L serial encoding blocks and other embedding vectors corresponding to the patches in the other video frames generated at block l- 1 ;

computing a spatial attention based on comparisons between the embedding vector corresponding to the particular patch generated at block l- 1 and additional embedding vectors corresponding to the other patches in the particular video frame generated at block l- 1 ; and

generating an embedding vector corresponding to the particular patch based on the computed temporal attention and the computed spatial attention.

17. The media of claim 16 , wherein the patches in the other video frames comprises patches at a spatial location of the particular patch in their corresponding video frames.

18. The media of claim 16 , wherein generating the embedding vector corresponding to the particular patch based on the computed temporal attention and the computed spatial attention comprises processing the computed temporal attention and the computed spatial attention using a multilayer perceptron (MLP).

19. The media of claim 15 , wherein the classification embedding is generated by taking a layer normalization on one or more of embedding vectors generated at a last block of the L serial encoding blocks.

20. A system comprising:

one or more processors; and

a non-transitory memory coupled to the one or more processors comprising instructions executable by the one or more processors, the one or more processors operable when executing the instructions to:

access a stream of F video frames, wherein the F video frames comprises N patches that are non-overlapping;

generate, for N×F patches in the F video frames, an initial embedding vector;

generate a classification embedding by processing generated N×F initial embedding vectors using a self-attention-based machine-learning model that computes a temporal attention and a spatial attention for the N×F patches, wherein the temporal attention and the spatial attention of a particular patch in a particular video frame of the F video frames are computed by comparing information associated with the particular patch with information associated with (1) patches in other video frames of the F video frames and (2) other patches in the particular video frame; and

determine a class of the stream of video frames based on the generated classification embedding.

Assignments (2)
CHANGE OF NAME Recorded Dec 20, 2021
From: FACEBOOK, INC.
To: META PLATFORMS, INC.
Reel/Frame 058553/0802 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Oct 13, 2021
From: BERTASIUS, GEDIMINAS; WANG, HENG; TORRESANI, LORENZO
To: FACEBOOK, INC.
Reel/Frame 057783/0160 →
Continuity (2)
Provisional Application 63147137 · Feb 8, 2021
Related Publication 20220253633A1 · Aug 11, 2022