Common action localization
Aspects of the disclosure are directed to an apparatus configured to perform common-action localization. In certain aspects, the apparatus may receive a query video comprising a plurality of frames, wherein a first query proposal is determined based on a subset of frames of the plurality of frames, the first query proposal indicative of an action depicted on the subset of frames. In certain aspects, the apparatus may determine a first attendance for a first support video of a plurality of support videos. In certain aspects, the apparatus may determine a second attendance for a second support video of the plurality of support videos after computing the first attendance.
1 . An apparatus for performing common-action localization, comprising:
one or more memories, individually or in combination, having instructions; and
one or more processors, individually or in combination, configured to execute the instructions and cause the apparatus to:
receive a query video comprising a plurality of frames, wherein a first query proposal is determined based on a subset of frames of the plurality of frames, the first query proposal indicative of an action depicted on the subset of frames;
compute relevance of the first query proposal to each support video of a plurality of support videos, individually one at a time, wherein the computing comprises:
determining a first attendance indicative of a first probability that the action is found in a first support video of the plurality of support videos by generating and up-sampling, via at least one neural network, at least one feature map associated with the first support video; and
determining a second attendance indicative of a second probability that the action is found in a second support video of the plurality of support videos by generating and up-sampling, via the at least one neural network, at least one feature map associated with the second support video, wherein the second attendance is determined after the first attendance is determined; and
output a classification of the subset of frames based at least in part on the first attendance and the second attendance.
2 . The apparatus of claim 1 , wherein the second attendance is determined independent of the first attendance.
3 . The apparatus of claim 1 , wherein the first attendance is further indicative of whether the first support video comprises one or more frames matching the first query proposal, and wherein the second attendance is further indicative of whether the second support video comprises one or more frames matching the first query proposal.
4 . The apparatus of claim 3 , wherein the one or more processors are further configured to cause the apparatus to:
apply a one-dimensional temporal convolution to the one or more frames from each of the first support video and the second support video matching the first query proposal.
5 . The apparatus of claim 3 , wherein the one or more processors are further configured to cause the apparatus to:
generate a third support video consisting of the one or more frames from each of the first support video and the second support video matching the first query proposal.
6 . The apparatus of claim 1 , wherein the first query proposal is one of multiple query proposals determined based on the plurality of frames, and wherein the one or more processors are further configured to cause the apparatus to:
determine the first query proposal of the multiple query proposals has a highest probability that the action the first query proposal is indicative of is found among the plurality of support videos, relative to one or more other remaining query proposals of the multiple query proposals that are indicative of one or more other actions.
7 . The apparatus of claim 1 , wherein the one or more processors are further configured to cause the apparatus to:
classify each of the first support video and the second support video based on pseudo-action classes.
8 . The apparatus of claim 7 , wherein the pseudo-action classes are mapped to a k-means cluster.
9 . A method for performing common-action localization, comprising:
receiving, by one or more processors, a query video comprising a plurality of frames, wherein a first query proposal is determined based on a subset of frames of the plurality of frames, the first query proposal indicative of an action depicted on the subset of frames;
computing, by one or more processors, relevance of the first query proposal to each support video of a plurality of support videos, individually one at a time, wherein the computing comprises:
determining a first attendance indicative of a first probability that the action is found in a first support video of the plurality of support videos by generating and up-sampling, via at least one neural network, at least one feature map associated with the first support video; and
determining a second attendance indicative of a second probability that the action is found in a second support video of the plurality of support videos by generating and up-sampling, via the at least one neural network, at least one feature map associated with the second support video, wherein the second attendance is determined after the first attendance is determined; and
outputting, by the one or more processors, a classification of the subset of frames based at least in part on the first attendance and the second attendance.
10 . The method of claim 9 , wherein the second attendance is determined independent of the first attendance.
11 . The method of claim 9 , wherein the first attendance is further indicative of whether the first support video comprises one or more frames matching the first query proposal, and wherein the second attendance is further indicative of whether the second support video comprises one or more frames matching the first query proposal.
12 . The method of claim 11 , wherein the method further comprises:
applying a one-dimensional temporal convolution to the one or more frames from each of the first support video and the second support video matching the first query proposal.
13 . The method of claim 11 , wherein the method further comprises:
generating a third support video consisting of the one or more frames from each of the first support video and the second support video matching the first query proposal.
14 . The method of claim 9 , wherein the first query proposal is one of multiple query proposals determined based on the plurality of frames, and wherein the method further comprises:
determining the first query proposal of the multiple query proposals has a highest probability that the action the first query proposal is indicative of is found among the plurality of support videos, relative to one or more other remaining query proposals of the multiple query proposals that are indicative of one or more other actions.
15 . The method of claim 9 , wherein the method further comprises:
classifying each of the first support video and the second support video based on pseudo-action classes.
16 . The method of claim 15 , wherein the pseudo-action classes are mapped to a k-means cluster.
17 . An apparatus for performing common-action localization, comprising:
means for receiving a query video comprising a plurality of frames, wherein a first query proposal is determined based on a subset of frames of the plurality of frames, the first query proposal indicative of an action depicted on the subset of frames;
means for computing relevance of the first query proposal to each support video of a plurality of support videos, individually one at a time, wherein the computing comprises:
determining a first attendance indicative of a first probability that the action is found in a first support video of a plurality of support videos by generating and up-sampling, via at least one neural network, at least one feature map associated with the first support video; and
determining a second attendance indicative of a second probability that the action is found in a second support video of the plurality of support videos by generating and up-sampling, via the at least one neural network, at least one feature map associated with the second support video, wherein the second attendance is determined after the first attendance is determined; and
means for outputting a classification of the subset of frames based at least in part on the first attendance and the second attendance.
18 . The apparatus of claim 17 , wherein the second attendance is determined independent of the first attendance.
19 . The apparatus of claim 17 , wherein the first attendance is further indicative of whether the first support video comprises one or more frames matching the first query proposal, and wherein the second attendance is further indicative of whether the second support video comprises one or more frames matching the first query proposal.
20 . The apparatus of claim 19 , wherein the apparatus further comprises:
means for applying a one-dimensional temporal convolution to the one or more frames from each of the first support video and the second support video matching the first query proposal.
21 . The apparatus of claim 19 , wherein the apparatus further comprises:
means for generating a third support video consisting of the one or more frames from each of the first support video and the second support video matching the first query proposal.
22 . The apparatus of claim 17 , wherein the first query proposal is one of multiple query proposals determined based on the plurality of frames, and wherein the apparatus further comprises:
means for determining the first query proposal of the multiple query proposals has a highest probability that the action the first query proposal is indicative of is found among the plurality of support videos, relative to one or more other remaining query proposals of the multiple query proposals that are indicative of one or more other actions.
23 . The apparatus of claim 17 , wherein the apparatus further comprises:
means for classifying each of the first support video and the second support video based on pseudo-action classes.
24 . The apparatus of claim 23 , wherein the pseudo-action classes are mapped to a k-means cluster.
25 . A non-transitory computer-readable medium comprising computer executable code, the code when executed by one or more processors causes the one or more processors, individually or in combination, to perform operations comprising:
receiving a query video comprising a plurality of frames, wherein a first query proposal is determined based on a subset of frames of the plurality of frames, the first query proposal indicative of an action depicted on the subset of frames;
computing relevance of the first query proposal to each support video of a plurality of support videos, individually one at a time, wherein the computing comprises:
determining a first attendance indicative of a first probability that the action is found in a first support video of the plurality of support videos by generating and up-sampling, via at least one neural network, at least one feature map associated with the first support video; and
determining a second attendance indicative of a second probability that the action is found in a second support video of the plurality of support videos by generating and up-sampling, via the at least one neural network, at least one feature map associated with the second support video, wherein the second attendance is determined after the first attendance is determined; and
outputting a classification of the subset of frames based at least in part on the first attendance and the second attendance.
26 . The non-transitory computer-readable medium of claim 25 , wherein the second attendance is determined independent of the first attendance.
27 . The non-transitory computer-readable medium of claim 25 , wherein the first attendance is further indicative of whether the first support video comprises one or more frames matching the first query proposal, and wherein the second attendance is further indicative of whether the second support video comprises one or more frames matching the first query proposal.
28 . The non-transitory computer-readable medium of claim 27 , wherein the operations further comprise:
applying a one-dimensional temporal convolution to the one or more frames from each of the first support video and the second support video matching the first query proposal.
29 . The non-transitory computer-readable medium of claim 27 , wherein the operations further comprise:
generating a third support video consisting of the one or more frames from each of the first support video and the second support video matching the first query proposal.
30 . The non-transitory computer-readable medium of claim 25 , wherein the first query proposal is one of multiple query proposals determined based on the plurality of frames, and wherein the operations further comprise:
determining the first query proposal of the multiple query proposals has a highest probability that the action the first query proposal is indicative of is found among the plurality of support videos, relative to one or more other remaining query proposals of the multiple query proposals that are indicative of one or more other actions.