IP Library › Granted Patent US 12,620,227
Granted Patent B2
US 12,620,227 · App. 18/360,741 · Granted May 5, 2026

Common action localization

Inventors: Juntae Lee (Seoul, KR); Mihir Jain (Zürich, CH); Sungrack Yun (Gyeonggi-do, KR)
Assignee: QUALCOMM Incorporated
G06V20/48G06F16/7328G06F16/735G06F16/75G06V10/764G06V20/41G06V20/46G06V10/82
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,620,227
App. No.
18/360,741
Granted
May 5, 2026
Kind
B2
Abstract

Aspects of the disclosure are directed to an apparatus configured to perform common-action localization. In certain aspects, the apparatus may receive a query video comprising a plurality of frames, wherein a first query proposal is determined based on a subset of frames of the plurality of frames, the first query proposal indicative of an action depicted on the subset of frames. In certain aspects, the apparatus may determine a first attendance for a first support video of a plurality of support videos. In certain aspects, the apparatus may determine a second attendance for a second support video of the plurality of support videos after computing the first attendance.

Claims (67)

1 . An apparatus for performing common-action localization, comprising:

one or more memories, individually or in combination, having instructions; and

one or more processors, individually or in combination, configured to execute the instructions and cause the apparatus to:

receive a query video comprising a plurality of frames, wherein a first query proposal is determined based on a subset of frames of the plurality of frames, the first query proposal indicative of an action depicted on the subset of frames;

compute relevance of the first query proposal to each support video of a plurality of support videos, individually one at a time, wherein the computing comprises:

determining a first attendance indicative of a first probability that the action is found in a first support video of the plurality of support videos by generating and up-sampling, via at least one neural network, at least one feature map associated with the first support video; and

determining a second attendance indicative of a second probability that the action is found in a second support video of the plurality of support videos by generating and up-sampling, via the at least one neural network, at least one feature map associated with the second support video, wherein the second attendance is determined after the first attendance is determined; and

output a classification of the subset of frames based at least in part on the first attendance and the second attendance.

2 . The apparatus of claim 1 , wherein the second attendance is determined independent of the first attendance.

3 . The apparatus of claim 1 , wherein the first attendance is further indicative of whether the first support video comprises one or more frames matching the first query proposal, and wherein the second attendance is further indicative of whether the second support video comprises one or more frames matching the first query proposal.

4 . The apparatus of claim 3 , wherein the one or more processors are further configured to cause the apparatus to:

apply a one-dimensional temporal convolution to the one or more frames from each of the first support video and the second support video matching the first query proposal.

5 . The apparatus of claim 3 , wherein the one or more processors are further configured to cause the apparatus to:

generate a third support video consisting of the one or more frames from each of the first support video and the second support video matching the first query proposal.

6 . The apparatus of claim 1 , wherein the first query proposal is one of multiple query proposals determined based on the plurality of frames, and wherein the one or more processors are further configured to cause the apparatus to:

determine the first query proposal of the multiple query proposals has a highest probability that the action the first query proposal is indicative of is found among the plurality of support videos, relative to one or more other remaining query proposals of the multiple query proposals that are indicative of one or more other actions.

7 . The apparatus of claim 1 , wherein the one or more processors are further configured to cause the apparatus to:

classify each of the first support video and the second support video based on pseudo-action classes.

8 . The apparatus of claim 7 , wherein the pseudo-action classes are mapped to a k-means cluster.

9 . A method for performing common-action localization, comprising:

receiving, by one or more processors, a query video comprising a plurality of frames, wherein a first query proposal is determined based on a subset of frames of the plurality of frames, the first query proposal indicative of an action depicted on the subset of frames;

computing, by one or more processors, relevance of the first query proposal to each support video of a plurality of support videos, individually one at a time, wherein the computing comprises:

determining a first attendance indicative of a first probability that the action is found in a first support video of the plurality of support videos by generating and up-sampling, via at least one neural network, at least one feature map associated with the first support video; and

determining a second attendance indicative of a second probability that the action is found in a second support video of the plurality of support videos by generating and up-sampling, via the at least one neural network, at least one feature map associated with the second support video, wherein the second attendance is determined after the first attendance is determined; and

outputting, by the one or more processors, a classification of the subset of frames based at least in part on the first attendance and the second attendance.

10 . The method of claim 9 , wherein the second attendance is determined independent of the first attendance.

11 . The method of claim 9 , wherein the first attendance is further indicative of whether the first support video comprises one or more frames matching the first query proposal, and wherein the second attendance is further indicative of whether the second support video comprises one or more frames matching the first query proposal.

12 . The method of claim 11 , wherein the method further comprises:

applying a one-dimensional temporal convolution to the one or more frames from each of the first support video and the second support video matching the first query proposal.

13 . The method of claim 11 , wherein the method further comprises:

generating a third support video consisting of the one or more frames from each of the first support video and the second support video matching the first query proposal.

14 . The method of claim 9 , wherein the first query proposal is one of multiple query proposals determined based on the plurality of frames, and wherein the method further comprises:

determining the first query proposal of the multiple query proposals has a highest probability that the action the first query proposal is indicative of is found among the plurality of support videos, relative to one or more other remaining query proposals of the multiple query proposals that are indicative of one or more other actions.

15 . The method of claim 9 , wherein the method further comprises:

classifying each of the first support video and the second support video based on pseudo-action classes.

16 . The method of claim 15 , wherein the pseudo-action classes are mapped to a k-means cluster.

17 . An apparatus for performing common-action localization, comprising:

means for receiving a query video comprising a plurality of frames, wherein a first query proposal is determined based on a subset of frames of the plurality of frames, the first query proposal indicative of an action depicted on the subset of frames;

means for computing relevance of the first query proposal to each support video of a plurality of support videos, individually one at a time, wherein the computing comprises:

determining a first attendance indicative of a first probability that the action is found in a first support video of a plurality of support videos by generating and up-sampling, via at least one neural network, at least one feature map associated with the first support video; and

determining a second attendance indicative of a second probability that the action is found in a second support video of the plurality of support videos by generating and up-sampling, via the at least one neural network, at least one feature map associated with the second support video, wherein the second attendance is determined after the first attendance is determined; and

means for outputting a classification of the subset of frames based at least in part on the first attendance and the second attendance.

18 . The apparatus of claim 17 , wherein the second attendance is determined independent of the first attendance.

19 . The apparatus of claim 17 , wherein the first attendance is further indicative of whether the first support video comprises one or more frames matching the first query proposal, and wherein the second attendance is further indicative of whether the second support video comprises one or more frames matching the first query proposal.

20 . The apparatus of claim 19 , wherein the apparatus further comprises:

means for applying a one-dimensional temporal convolution to the one or more frames from each of the first support video and the second support video matching the first query proposal.

21 . The apparatus of claim 19 , wherein the apparatus further comprises:

means for generating a third support video consisting of the one or more frames from each of the first support video and the second support video matching the first query proposal.

22 . The apparatus of claim 17 , wherein the first query proposal is one of multiple query proposals determined based on the plurality of frames, and wherein the apparatus further comprises:

means for determining the first query proposal of the multiple query proposals has a highest probability that the action the first query proposal is indicative of is found among the plurality of support videos, relative to one or more other remaining query proposals of the multiple query proposals that are indicative of one or more other actions.

23 . The apparatus of claim 17 , wherein the apparatus further comprises:

means for classifying each of the first support video and the second support video based on pseudo-action classes.

24 . The apparatus of claim 23 , wherein the pseudo-action classes are mapped to a k-means cluster.

25 . A non-transitory computer-readable medium comprising computer executable code, the code when executed by one or more processors causes the one or more processors, individually or in combination, to perform operations comprising:

receiving a query video comprising a plurality of frames, wherein a first query proposal is determined based on a subset of frames of the plurality of frames, the first query proposal indicative of an action depicted on the subset of frames;

computing relevance of the first query proposal to each support video of a plurality of support videos, individually one at a time, wherein the computing comprises:

determining a first attendance indicative of a first probability that the action is found in a first support video of the plurality of support videos by generating and up-sampling, via at least one neural network, at least one feature map associated with the first support video; and

determining a second attendance indicative of a second probability that the action is found in a second support video of the plurality of support videos by generating and up-sampling, via the at least one neural network, at least one feature map associated with the second support video, wherein the second attendance is determined after the first attendance is determined; and

outputting a classification of the subset of frames based at least in part on the first attendance and the second attendance.

26 . The non-transitory computer-readable medium of claim 25 , wherein the second attendance is determined independent of the first attendance.

27 . The non-transitory computer-readable medium of claim 25 , wherein the first attendance is further indicative of whether the first support video comprises one or more frames matching the first query proposal, and wherein the second attendance is further indicative of whether the second support video comprises one or more frames matching the first query proposal.

28 . The non-transitory computer-readable medium of claim 27 , wherein the operations further comprise:

applying a one-dimensional temporal convolution to the one or more frames from each of the first support video and the second support video matching the first query proposal.

29 . The non-transitory computer-readable medium of claim 27 , wherein the operations further comprise:

generating a third support video consisting of the one or more frames from each of the first support video and the second support video matching the first query proposal.

30 . The non-transitory computer-readable medium of claim 25 , wherein the first query proposal is one of multiple query proposals determined based on the plurality of frames, and wherein the operations further comprise:

determining the first query proposal of the multiple query proposals has a highest probability that the action the first query proposal is indicative of is found among the plurality of support videos, relative to one or more other remaining query proposals of the multiple query proposals that are indicative of one or more other actions.

Assignments (2)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Apr 25, 2025
From: LEE, JUNTAE; JAIN, MIHIR; YUN, SUNGRACK
To: QUALCOMM INCORPORATED
Reel/Frame 070952/0475 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Sep 1, 2023
From: LEE, JUNTAE; JAIN, MIHIR; YUN, SUNGRACK
To: QUALCOMM INCORPORATED
Reel/Frame 064777/0217 →
Continuity (2)
Provisional Application 63450924 · Mar 8, 2023
Related Publication 20240303987A1 · Sep 12, 2024
References Cited (34)
US 9361523B1 · Chen · 2016 [cited by examiner]
US 9672280B2 · Liu · 2017 [cited by examiner]
US 10210252B2 · Pereira · 2019 [cited by examiner]
US 10839223B1 · Jiang · 2020 [cited by examiner]
US 11281718B2 · Pereira · 2022 [cited by examiner]
US 11734287B2 · Sharifi · 2023 [cited by examiner]
US 12210564B2 · Majkowska · 2025 [cited by examiner]
US 20070117627A1 · Raman · 2007 [cited by examiner]
US 20150293996A1 · Liu · 2015 [cited by examiner]
US 20150363660A1 · Vidal · 2015 [cited by examiner]
US 20170068423A1 · Napolitano · 2017 [cited by examiner]
US 20170192980A1 · Pereira · 2017 [cited by examiner]
US 20170337271A1 · Lee · 2017 [cited by examiner]
US 20200004781A1 · Pereira · 2020 [cited by examiner]
US 20200159765A1 · Manin · 2020 [cited by examiner]
US 20220156514A1 · Gavrilyuk · 2022 [cited by examiner]
US 20220188321A1 · Sharifi · 2022 [cited by examiner]
US 20240134506A1 · Napolitano · 2024 [cited by examiner]
CN 115527152A · 2022 [cited by applicant]
EP 2955645B1 · 2017 [cited by examiner]
WO WO2021169209A1 · 2021 [cited by examiner]
WO WO2021221209A1 · 2021 [cited by examiner]
WO WO2023000779A1 · 2023 [cited by examiner]
Yang et al., “Localizing the common action among a few videos.” In European conference on computer vision, pp. 505-521. Cham: Springer International Publishing, 2020. (Year: 2020). [cited by examiner]
Gong et al., “Learning temporal co-attention models for unsupervised video action localization.” In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, pp. 9819-9828. 2020. (Year: 2020). [cited by examiner]
Liu et al., “Learning global pose features in graph convolutional networks for 3d human pose estimation.” In Proceedings of the Asian conference on computer vision. 2020. (Year: 2020). [cited by examiner]
Kim et al., “Efficient action recognition via dynamic knowledge propagation.” In Proceedings of the IEEE/CVF International Conference on Computer Vision, p. 13719-13728. 2021. (Year: 2021). [cited by examiner]
EP-2955645-B1 (machine translation) (Year: 2017). [cited by examiner]
WO-2023000779-A1 (machine translation) (Year: 2023). [cited by examiner]
International Search Report and Written Opinion—PCT/US2023/086100—ISA/EPO—Apr. 23, 2024. [cited by applicant]
Jain M., et al., “ActionBytes: Learning from Trimmed Videos to Localize Actions”, 2020 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), Jun. 13, 2020, pp. 1168-1177, XP033803437, abstract, figures … [cited by applicant]
Tan S., et al., “Learning Similarity: Feature-Aligning Network for Few-shot Action Recognition”, 2019 International Joint Conference on Neural Networks (IJCNN), IEEE, Jul. 14, 2019, 7 Pages, XP033621539, abstract, figur… [cited by applicant]
Yang H., et al., “One-shot Action Localization by Learning Sequence Matching Network”, 2018 IEEE/CVF Conference on Computer Vision and Pattern Recognition, Jun. 18, 2018, pp. 1450-1459, XP033476108, abstract, figures 1,… [cited by applicant]
Yang P., et al., “Few-Shot Transformation of Common Actions into Time and Space”, 2021 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), Jun. 20, 2021, pp. 16023-16035, XP034009773. [cited by applicant]