IP Library Granted Patent US 10,984,245
Granted Patent B1
US 10,984,245 · App. 16/286,377 · Granted Apr 20, 2021

Convolutional neural network based on groupwise convolution for efficient video analysis

View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 10,984,245
App. No.
16/286,377
Granted
Apr 20, 2021
Kind
B1
Abstract

In one embodiment, a method includes receiving a request for information associated with a video, determining the information associated with the video by processing the video using a machine-learning model which is based on a convolutional neural network comprising a plurality of layers, wherein at least one of the plurality of layers comprises one or more building blocks, wherein at least one of the one or more building blocks comprises a first filter configured to perform a three-dimensional (3D) pointwise convolutional operation and a second filter configured to perform a three-dimensional (3D) groupwise convolutional operation, and outputting the information associated with the video in response to the request.

Claims (46)

1. A method comprising, by one or more computing systems:

receiving a request for information associated with a video;

determining, by processing the video using a machine-learning model, the information associated with the video, wherein the machine-learning model is based on a convolutional neural network comprising a plurality of layers, wherein at least one of the plurality of layers comprises one or more building blocks, wherein at least one of the one or more building blocks comprises:

a first filter configured to perform a three-dimensional (3D) pointwise convolutional operation on an input to the first filter;

a second filter configured to perform a three-dimensional (3D) groupwise convolutional operation on an input to the second filter, wherein the input to the second filter comprises an output from the first filter; and

a third filter configured to perform a three-dimensional (3D) pointwise convolutional operation on an input to the third filter, wherein the input to the third filter comprises an output from the second filter; and

outputting, in response to the request, the information associated with the video.

2. The method of claim 1 , wherein the information associated with the video comprises one or more of:

a category associated with the video;

a detection result associated with the video; or

a segmentation result associated with the video.

3. The method of claim 1 , wherein the video is associated with one or more channels.

4. The method of claim 3 , wherein the 3D groupwise convolutional operation is associated with a process comprising:

determining one or more groups for the one or more channels, wherein each group comprises one or more of the one or more channels; and

applying a convolutional operation to each of the one or more groups separately.

5. The method of claim 3 , wherein the 3D groupwise convolutional operation comprises a 3D depthwise convolutional operation.

6. The method of claim 5 , wherein the 3D depthwise convolutional operation is associated with one or more input channels and one or more output channels, and wherein a number of the one or more input channels equals a number of the one or more output channels.

7. The method of claim 5 , wherein the 3D depthwise convolutional operation is associated with a process comprising:

determining one or more groups for the one or more channels, wherein each group comprises one channel of the one or more channels; and

applying a convolutional operation to each of the one or more groups separately.

8. The method of claim 1 , wherein the convolutional neural network further comprises a plurality of paddings, kernels, and stridings.

9. The method of claim 1 , wherein the convolutional neural network is based on a 3D network architecture.

10. The method of claim 1 , further comprising:

training the machine-learning model based on a plurality of training videos.

11. The method of claim 1 , wherein the request is associated with a requirement of a trade-off between accuracy and computational cost.

12. The method of claim 11 , further comprising:

determining a number of the plurality of layers based on the requirement of the trade-off between accuracy and computational cost.

13. The method of claim 11 , wherein the at least one building block comprises a plurality of first filters and a plurality of second filters.

14. The method of claim 13 , further comprising:

determining a number of the plurality of first filters based on the requirement of the trade-off between accuracy and computational cost.

15. The method of claim 13 , further comprising:

determining a number of the plurality of second filters based on the requirement of the trade-off between accuracy and computational cost.

16. One or more computer-readable non-transitory storage media embodying software that is operable when executed to:

receive a request for information associated with a video;

determine, by processing the video using a machine-learning model, the information associated with the video, wherein the machine-learning model is based on a convolutional neural network comprising a plurality of layers, wherein at least one of the plurality of layers comprises one or more building blocks, wherein at least one of the one or more building blocks comprises:

a first filter configured to perform a three-dimensional (3D) pointwise convolutional operation on an input to the first filter;

a second filter configured to perform a three-dimensional (3D) groupwise convolutional operation on an input to the second filter, wherein the input to the second filter comprises an output from the first filter; and

a third filter configured to perform a three-dimensional (3D) pointwise convolutional operation on an input to the third filter, wherein the input to the third filter comprises an output from the second filter; and

output, in response to the request, the information associated with the video.

17. A system comprising: one or more processors; and a non-transitory memory coupled to the processors comprising instructions executable by the processors, the processors operable when executing the instructions to:

receive a request for information associated with a video;

determine, by processing the video using a machine-learning model, the information associated with the video, wherein the machine-learning model is based on a convolutional neural network comprising a plurality of layers, wherein at least one of the plurality of layers comprises one or more building blocks, wherein at least one of the one or more building blocks comprises:

a first filter configured to perform a three-dimensional (3D) pointwise convolutional operation on an input to the first filter;

a second filter configured to perform a three-dimensional (3D) groupwise convolutional operation on an input to the second filter, wherein the input to the second filter comprises an output from the first filter; and

a third filter configured to perform a three-dimensional (3D) pointwise convolutional operation on an input to the third filter, wherein the input to the third filter comprises an output from the second filter; and

output, in response to the request, the information associated with the video.

Assignments (2)
CHANGE OF NAME Recorded Dec 20, 2021
From: FACEBOOK, INC.
To: META PLATFORMS, INC.
Reel/Frame 058553/0802 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Mar 14, 2019
From: TRAN, DU LE HONG; HE, KAIMING; WANG, HENG; FEISZLI, MATTHEW DAN; TORRESANI, LORENZO
To: FACEBOOK, INC.
Reel/Frame 048603/0751 →