IP Library › Granted Patent US 12,206,869
Granted Patent B2
US 12,206,869 · App. 17/527,659 · Granted Jan 21, 2025

Skip convolutions for efficient video processing

Inventors: Amirhossein Habibian (Amsterdam, NL); Davide Abati (Amsterdam, NL); Babak Ehteshami Bejnordi (Amsterdam, NL)
Assignee: QUALCOMM INCORPORATED
H04N19/197G06N3/02H04N19/172H04N19/184
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,206,869
App. No.
17/527,659
Granted
Jan 21, 2025
Kind
B2
Abstract

A method for video processing via an artificial neural network includes receiving a video stream as an input at the artificial neural network. A residual is computed based on a difference between a first feature of a current frame of the video stream and a second feature of a previous frame of the video stream. One or more portions of the current frame of the video stream are processed based on the residual. Additionally, processing is skipped for one or more portions of the current frame of the video based on the residual.

Claims (54)

1. A method for video processing with an artificial neural network (ANN), comprising:

receiving a video stream as an input at the artificial neural network;

computing a residual based on a difference between a first feature of a current frame of the video stream and a second feature of a previous frame of the video stream; and

determining whether to perform processing, associated with a convolutional layer of the artificial neural network, of one or more portions of the current frame of the video stream based on of the residual.

2. The method of claim 1 , in which the one or more portions of the current frame includes only salient regions of the current frame.

3. The method of claim 2 , further comprising:

determining to perform processing, associated with the convolutional layer of the artificial neural network, of one or more portions of the current frame; and

applying, based on the determination to perform processing, a convolution kernel to only the salient regions of the current frame.

4. The method of claim 2 , further comprising determining the one or more salient regions based on whether the residual is greater than a predetermined threshold value.

5. The method of claim 1 , further comprising not performing processing of at least one portion of the current frame of the video stream based on the residual.

6. The method of claim 5 , in which a first output corresponding to the at least one portion of the current frame is set equal to a second output corresponding to at least one portion of the previous frame.

7. The method of claim 1 , further comprising:

comparing the residual to a predefined threshold value; and

determining whether to apply a mask to the first feature based on the comparing.

8. The method of claim 1 , further comprising learning a gating function for gating the convolutional layer, the gating function configured to apply a mask to one or more portions of the current frame based on the residual.

9. The method of claim 8 , further comprising generating a saliency map based on the gating function.

10. The method of claim 1 , further comprising adaptively adjusting an amount of computation performed in processing the video stream based on an amount of information observed per frame.

11. An apparatus for video processing with an artificial neural network (ANN), comprising:

memory; and

at least one processor coupled to the memory, the at least one processor being configured:

to receive a video stream as an input at the artificial neural network;

to compute a residual based on a difference between a first feature of a current frame of the video stream and a second feature of a previous frame of the video stream;

and

to determine whether to perform processing, associated with a convolutional layer of the artificial neural network, of one or more portions of the current frame of the video stream based on of the residual.

12. The apparatus of claim 11 , in which the one or more portions of the current frame includes only salient regions of the current frame.

13. The apparatus of claim 12 , in which the at least one processor is further configured; to determine to perform processing, associated with the convolutional layer of the artificial neural network, of one or more portions of the current frame; and to apply, based on the determination to perform processing, a convolution kernel to only the salient regions of the current frame.

14. The apparatus of claim 12 , in which the at least one processor is further configured to determine the one or more salient regions based on whether the residual is greater than a predetermined threshold value.

15. The apparatus of claim 11 , in which the at least one processor is further configured to not perform refrain from processing of at least one portion of the current frame of the video stream based on the residual.

16. The apparatus of claim 15 , in which a first output corresponding to the at least one portion of the current frame is set equal to a second output corresponding to at least one portion of the previous frame.

17. The apparatus of claim 11 , in which the at least one processor is further configured:

to compare the residual to a predefined threshold value; and

to determine whether to apply a mask to the first feature based on the comparing.

18. The apparatus of claim 11 , in which the at least one processor is further configured to learn a gating function for gating the convolutional layer, the gating function configured to apply a mask to one or more portions of the current frame based on the residual.

19. The apparatus of claim 18 , in which the at least one processor is further configured to generate a saliency map based on the gating function.

20. The apparatus of claim 11 , in which the at least one processor is further configured to adaptively adjust an amount of computation based on an amount of information observed per frame.

21. An apparatus for video processing with an artificial neural network (ANN), comprising:

means for receiving a video stream as an input at the artificial neural network;

means for computing a residual based on a difference between a first feature of a current frame of the video stream and a second feature of a previous frame of the video stream; and

means for determining whether to perform processing, associated with a convolutional layer of the artificial neural network, of one or more portions of the current frame of the video stream based on of the residual.

22. The apparatus of claim 21 , in which the one or more portions of the current frame includes only salient regions of the current frame.

23. The apparatus of claim 22 , further comprising:

means for determining to perform processing, associated with the convolutional layer of the artificial neural network, of one or more portions of the current frame; and

means for applying, based on the determination to perform processing, a convolution kernel to only the salient regions of the current frame.

24. The apparatus of claim 22 , further comprising means for determining the one or more salient regions based on whether the residual is greater than a predetermined threshold value.

25. The apparatus of claim 21 , further comprising means for learning a gating function for gating the convolutional layer, the gating function configured to apply a mask to one or more portions of the current frame based on the residual.

26. A non-transitory computer readable medium having encoded thereon program code for video processing with an artificial neural network (ANN), the program code being executed by a processor and comprising:

program code to receive a video stream as an input at the artificial neural network;

program code to compute a residual based on a difference between a first feature of a current frame of the video stream and a second feature of a previous frame of the video stream; and

program code to determine whether to perform processing, associated with a convolutional layer of the artificial neural network, of one or more portions of the current frame of the video stream based on of the residual.

27. The non-transitory computer readable medium of claim 26 , in which the one or more portions of the current frame includes only salient regions of the current frame.

28. The non-transitory computer readable medium of claim 27 , in which the at least one processor is further configured to determine to perform processing, associated with the convolutional layer of the artificial neural network, of one or more portions of the current frame; and

to apply, based on the determination to perform processing, a convolution kernel to only the salient regions of the current frame.

29. The non-transitory computer readable medium of claim 28 , in which the at least one processor is further configured to determine the one or more salient regions based on whether the residual is greater than a predetermined threshold value.

30. The non-transitory computer readable medium of claim 26 , in which the at least one processor is further configured to learn a gating function for gating the convolutional layer, the gating function configured to apply a mask to one or more portions of the current frame based on the residual.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Dec 2, 2021
From: HABIBIAN, AMIRHOSSEIN; ABATI, DAVIDE; EHTESHAMI BEJNORDI, BABAK
To: QUALCOMM INCORPORATED
Reel/Frame 058265/0588 →
Continuity (2)
Provisional Application 63114348 · Nov 16, 2020
Related Publication 20220159278A1 · May 19, 2022
References Cited (10)
US 6608910B1 · Srinivasa · 2003 [cited by examiner]
US 20110051992A1 · Cobb · 2011 [cited by examiner]
US 20200152224A1 · Goldhor · 2020 [cited by examiner]
US 20210223424A1 · Valensi · 2021 [cited by examiner]
US 20220051466A1 · Doyle · 2022 [cited by examiner]
US 20220400270A1 · Meardi · 2022 [cited by examiner]
Cavigelli L., et al., “CBinfer: Change-Based Inference for Convolutional Neural Networks on Video Data”, arxiv.org, Cornell University Library, 201 OLIN Library Cornell University Ithaca, NY 14853, Jun. 2017. [cited by applicant]
International Search Report and Written Opinion—PCT/US2021/059581—ISA/EPO—Mar. 7, 2022. [cited by applicant]
Verelst T., et al., “Dynamic Convolutions: Exploiting Spatial Sparsity for Faster Inference”, 2020 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), IEEE, Jun. 13, 2020 (Jun. 13, 2020), pp. 2317-232… [cited by applicant]
Ying W., et al., “An Edge 3D CNN Accelerator for Low-Power Activity Recognition”, IEEE Transactions on Computer-Aided Design of Integrated Circuits and Systems, IEEE, USA, vol. 40, No. 5, May 21, 2021, pp. 918-930, [Ret… [cited by applicant]
Cited By (1)
US 12,573,013