IP Library › Granted Patent US 12,333,811
Granted Patent B2
US 12,333,811 · App. 17/769,246 · Granted Jun 17, 2025

Permutation invariant convolution (PIC) for recognizing long-range activities, generating global representation of input streams and classifying activity based on global representation

Inventors: Noureldien Mahmoud Elsayed Hussein (Amsterdam, NL); Efstratios Gavves (Amsterdam, NL); Arnold Wilhelmus Maria Smeulders (Amsterdam, NL)
Assignee: QUALCOMM Incorporated
G06V20/46G06V10/82G06V20/41G06V20/49G06V20/582G06V20/44G06V30/19173
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,333,811
App. No.
17/769,246
Granted
Jun 17, 2025
Kind
B2
Abstract

A method for recognizing long-range activities in videos includes segmenting an input video stream to generate multiple frame sets. For each of the frame sets, a frame with a highest likelihood of including one or more actions of a set of predefined actions is identified regardless of its order in the frame set. A global representation of the input stream is generated based on pooled representations of the identified frames. A long-range activity in the video stream is classified based on the global representation.

Claims (56)

1. A method, comprising:

segmenting an input stream to generate a plurality of frame sets, each frame set including a plurality of frames;

identifying, by a permutation invariant convolutional layer of a neural network, for each frame set from the plurality of frame sets, a frame with a highest likelihood of including one or more actions of a set of predefined actions;

generating a global representation of the input stream from pooled representations of the identified frames; and

classifying a long-range activity based on the global representation.

2. The method of claim 1 , in which identifying the frame comprises generating a similarity matrix from a dot product of features of the frame set and a first kernel.

3. The method of claim 2 , further comprising max-pooling similarities in the similarity matrix to identify the frame with the highest likelihood.

4. The method of claim 3 , further comprising:

generating an attention vector from the similarity matrix; and

generating the global representation based on a dot product of the attention vector and a second kernel.

5. The method of claim 4 , in which the first kernel and the second kernel are linked.

6. The method of claim 1 , in which the global representation is based on a dot product of the pooled representations of the identified frames.

7. The method of claim 1 , in which generating the global representation is performed by the permutation invariant convolutional layer of the neural network.

8. The method of claim 7 , in which the neural network comprises a plurality of cascading permutation invariant convolutional layers.

9. An apparatus, comprising:

a memory; and

at least one processor coupled to the memory, the at least one processor configured to:

segment an input stream to generate a plurality of frame sets, each frame set including a plurality of frames;

identify, by a permutation invariant convolutional layer of a neural network, for each frame set from the plurality of frame sets, a frame with a highest likelihood of including a one or more actions of a set of predefined actions;

generate a global representation of the input stream from pooled representations of the identified frames; and

classify a long-range activity based on the global representation.

10. The apparatus of claim 9 , in which the at least one processor is configured to identify the frame by generating a similarity matrix from a dot product of features of the frame set and a first kernel.

11. The apparatus of claim 10 , in which the at least one processor is further configured to max-pool similarities in the similarity matrix to identify the frame with the highest likelihood.

12. The apparatus of claim 11 , in which the at least one processor is further configured to:

generating an attention vector from the similarity matrix; and

generating the global representation based on a dot product of the attention vector and a second kernel.

13. The apparatus of claim 12 , in which the first kernel and the second kernel are linked.

14. The apparatus of claim 9 , in which the at least on processor is further configured to generate the global representation based on a dot product of the pooled representations of the identified frames.

15. The apparatus of claim 9 , in which the at least one processor is further configured to generate the global representation by the permutation invariant convolutional layer of the neural network.

16. The apparatus of claim 15 , in which the neural network comprises a plurality of cascading permutation invariant convolutional layers.

17. An apparatus, comprising:

means for segmenting an input stream to generate a plurality of frame sets, each frame set including a plurality of frames;

means for identifying, by a permutation invariant convolutional layer of a neural network, for each frame set from the plurality of frame sets, a frame with a highest likelihood of including one or more actions of a set of predefined actions;

means for generating a global representation of the input stream from pooled representations of the identified frames; and

means for classifying a long-range activity based on the global representation.

18. The apparatus of claim 17 , further comprising means for generating a similarity matrix from a dot product of features of the frame set and a first kernel.

19. The apparatus of claim 18 , further comprising means for max-pooling similarities in the similarity matrix to identify the frame with the highest likelihood.

20. The apparatus of claim 19 , further comprising:

means for generating an attention vector from the similarity matrix; and

means for generating the global representation based on a dot product of the attention vector and a second kernel.

21. The apparatus of claim 20 , in which the first kernel and the second kernel are linked.

22. The apparatus of claim 17 , in which the global representation is based on a dot product of the pooled representations of the identified frames.

23. A non-transitory computer-readable medium having program code recorded, the program code executed by a processor and comprising:

program code to segment an input stream to generate a plurality of frame sets, each frame set including a plurality of frames;

program code to identify, by a permutation invariant convolutional layer of a neural network, for each frame set from the plurality of frame sets, a frame with a highest likelihood of including one or more actions of a set of predefined actions;

program code to generate a global representation of the input stream from pooled representations of the identified frames; and

program code to classify a long-range activity based on the global representation.

24. The non-transitory computer-readable medium of claim 23 , further comprising program code to identify the frame by generating a similarity matrix from a dot product of features of the frame set and a first kernel.

25. The non-transitory computer-readable medium of claim 24 , further comprising program code to max-pool similarities in the similarity matrix to identify the frame with the highest likelihood.

26. The non-transitory computer-readable medium of claim 25 , further comprising program code:

program code to generate an attention vector from the similarity matrix; and

program code to generate the global representation based on a dot product of the attention vector and a second kernel.

27. The non-transitory computer-readable medium of claim 26 , in which the first kernel and the second kernel are linked.

28. The non-transitory computer-readable medium of claim 23 , further comprising program code to generate the global representation based on a dot product of the pooled representations of the identified frames.

29. The non-transitory computer-readable medium of claim 23 , further comprising program code to generate the global representation at the permutation invariant convolutional layer of the neural network.

30. The non-transitory computer-readable medium of claim 29 , in which the neural network comprises a plurality of cascading permutation invariant convolutional layers.

Assignments (2)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jul 14, 2022
From: HUSSEIN, NOURELDIEN MAHMOUD ELSAYED; GAVVES, EFSTRATIOS; SMEULDERS, ARNOLD WILHELMUS MARIA
To: UNIVERSITEIT VAN AMSTERDAM
Reel/Frame 060510/0517 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jul 14, 2022
From: UNIVERSITEIT VAN AMSTERDAM
To: QUALCOMM TECHNOLOGIES, INC.
Reel/Frame 060510/0524 →
Priority Claims (1)
GR 20190100517 · Nov 15, 2019 · national
Continuity (1)
Related Publication 20240135708A1 · Apr 25, 2024
References Cited (18)
US 12087043B2 · Mittal · 2024 [cited by examiner]
US 20190108399A1 · Escorcia et al. · 2019 [cited by applicant]
US 20200160061A1 · Deng · 2020 [cited by examiner]
US 20240094807A1 · Zhang · 2024 [cited by examiner]
CN 109389055A · 2019 [cited by applicant]
CN 109753884A · 2019 [cited by applicant]
CN 109961005A · 2019 [cited by applicant]
Non Patent Literature (NPL) published to AJ Piergiovanni et. al., on Jun. 18, 2018 in IEEE/CVF Conference on Computer Vision and Pattern Recognition Workshop. [cited by examiner]
Hussein N., et al., “PIC: Permutation Invariant Convolution for Recognizing Long-Range Activities”, arxiv.org. Cornell University Library, 201 Olin Library Cornell University Ithaca, NY 14853, Mar. 18, 2020 (Mar. 18, 20… [cited by applicant]
International Search Report and Written Opinion—PCT/US2020/060595—ISA/EPO—Apr. 14, 2021. [cited by applicant]
Limin W., et al., “Latent Hierarchical Model of Temporal Structure for Complex Activity Classification”, IEEE Transactions on Image Processing, IEEE Service Center, Piscataway, NJ, US, vol. 23. No. 2, Feb. 1, 2014 (Feb.… [cited by applicant]
Ng J.Y-H., et al., “Beyond Short Snippets: Deep Networks for Video Classification”, 2015 IEEE Conference on Computer Vision and Pattern Recognition (CVPR), IEEE, Jun. 7, 2015 (Jun. 7, 2015), pp. 4694-4702, XP032793927, … [cited by applicant]
Piergiovanni AJ., et al., “Fine-Grained Activity Recognition in Baseball Videos”, 2018 IEEE/CVF Conference on Computer Vision and Pattern Recognition Workshops (CVPRW), IEEE, Jun. 18, 2018 (Jun. 18, 2018), pp. 1821-1821… [cited by applicant]
Wang K., et al., “3D Human Activity Recognition with Reconfigurable Convolutional Neural Networks”, Multimedia, ACM. 2 Penn Plaza, Suite 701 New York NY 10121-0701 USA, Nov. 3, 2014 (Nov. 3, 2014), pp. 97-106, XP0580586… [cited by applicant]
Limin W., et al., “Latent Hierarchical Model of Temporal Structure for Complex Activity Classification”, IEEE Transactions on Image Processing, IEEE Service Center, Piscataway, NJ, US, vol. 23. No. 2, Jan. 7, 2015, 28 P… [cited by applicant]
Ng J.Y-H., et al., “Beyond Short Snippets: Deep Networks for Video Classification”, 2015 IEEE Conference on Computer Vision and Pattern Recognition (CVPR), IEEE, Jun. 7, 2015, pp. 4694-4702. [cited by applicant]
Piergiovanni AJ., et al., “Fine-Grained Activity Recognition in Baseball Videos”, 2018 IEEE/CVF Conference on Computer Vision and Pattern Recognition Workshops (CVPRW), IEEE, Jun. 18, 2018, pp. 1853-1861. [cited by applicant]
Wang K., et al., “3D Human Activity Recognition with Reconfigurable Convolutional Neural Networks”, Multimedia, ACM. 2 Penn Plaza, Suite 701 New York NY 10121-0701 USA, Nov. 3, 2014, 11 Pages. [cited by applicant]