IP Library › Granted Patent US 12,406,500
Granted Patent B2
US 12,406,500 · App. 17/768,815 · Granted Sep 2, 2025

Moment localization in media stream

Inventors: Houwen Peng (Redmond, WA); Jianlong Fu (Beijing, CN)
Assignee: Microsoft Technology Licensing, LLC
G06V20/48G06F16/3344G06F40/10G06V10/62G06V10/7715G06V10/82G06V20/41G06V20/46G06V20/49
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,406,500
App. No.
17/768,815
Granted
Sep 2, 2025
Kind
B2
Abstract

Various implementations of the subject matter relate to moment localization in media stream. In some implementations, a two-dimensional temporal feature map representing a plurality of moments within a media stream is extracted from the media stream, wherein the two-dimensional temporal feature map comprises a first dimension representing a start of a respective one of the plurality of moments and a second dimension representing an end of a respective one of the plurality of moments. A correlation between the plurality of moments and an action in the media stream is determined based on the two-dimensional temporal feature map.

Claims (71)

1. A computer-implemented method, comprising:

extracting, from a media stream, a two-dimensional temporal feature map representing a plurality of moments within the media stream, wherein the two-dimensional temporal feature map comprises a first dimension representing a start of a respective one of the plurality of moments and a second dimension representing an end of a respective one of the plurality of moments;

encoding a sentence feature extracted from an input;

fusing the encoded sentence feature with the two-dimensional temporal feature map into a unified subspace as a fused two-dimensional temporal map;

applying a convolutional layer to the two-dimensional temporal feature map to obtain a further feature map having a same dimension as the two-dimensional temporal feature map, the convolutional layer comprises a dilated convolution and strides of the dilated convolution are configured to increase as lengths of the respective moments increase;

generating a temporal adjacent network using the fused two-dimensional temporal map and the further feature map;

determining, using the temporal adjacent network, a correlation between the plurality of moments and an action in the media stream; and

identifying a matching a set of candidate moments for the input using the temporal adjacent network.

2. The method of claim 1 , wherein extracting the two-dimensional temporal feature map comprises:

segmenting the media stream into a plurality of clips;

extracting features of respective ones of the plurality of clips to obtain a feature map of the media stream; and

extracting, from features of one or more clips corresponding to a moment of the plurality of moments in the feature map of the media stream, features of this moment as a part of the two-dimensional temporal feature map.

3. The method of claim 1 , wherein determining the correlation comprises:

sampling the plurality of moments at respective sample rates to determine a plurality of candidate moments, wherein the sample rates are adaptively adjusted based on lengths of respective ones of the plurality of moments; and

determining a correlation between the plurality of candidate moments and the action in the media stream.

4. The method of claim 3 , wherein the sample rates are configured to decrease as the lengths of the respective moments increase.

5. The method of claim 1 , wherein determining the correlation comprises:

applying a convolutional layer to the two-dimensional temporal feature map to obtain a further feature map having a same dimension as the two-dimensional temporal feature map; and

determining, based on the further feature map, scores of correlation between the plurality of moments and the action in the media stream.

6. The method of claim 1 , wherein determining the correlation comprises:

in response to receiving a query for a particular action in the media stream, extracting a feature vector of the query; and

determining the correlation based on the feature vector of the query and the two-dimensional temporal feature map.

7. The method of claim 6 , wherein determining the correlation comprises:

fusing the feature vector of the query and the two-dimensional temporal feature map to generate a further two-dimensional temporal feature map having a same dimension as the two-dimensional temporal feature map; and

determining, based on the further two-dimensional temporal feature map, the correlation between the plurality of moments and the particular action.

8. The method of claim 7 , wherein fusing the feature vector of the query and the two-dimensional temporal feature map comprises:

generating the further two-dimensional temporal feature map by applying a Hadamard product to the feature vector of the query and the two-dimensional temporal feature map.

9. The method of claim 6 , wherein the query comprises a natural language query.

10. The method of claim 1 , wherein the media stream comprises an untrimmed media stream.

11. A device comprising:

a processing unit; and

a memory coupled to the processing unit and having instructions stored thereon, the instructions, when executed by the processing unit, causing the device to perform acts comprising:

extracting, from a media stream, a two-dimensional temporal feature map representing a plurality of moments within the media stream, wherein the two-dimensional temporal feature map comprises a first dimension representing a start of a respective one of the plurality of moments and a second dimension representing an end of a respective one of the plurality of moments;

encoding a sentence feature extracted from an input;

fusing the encoded sentence feature with the two-dimensional temporal feature map into a unified subspace as a fused two-dimensional temporal map;

applying a convolutional layer to the two-dimensional temporal feature map to obtain a further feature map having a same dimension as the two-dimensional temporal feature map, the convolutional layer comprises a dilated convolution and strides of the dilated convolution are configured to increase as lengths of the respective moments increase;

generating a temporal adjacent network using the fused two-dimensional temporal map and the further feature map;

determining, using the temporal adjacent network, a correlation between the plurality of moments and an action in the media stream; and

identifying a matching a set of candidate moments for the input using the temporal adjacent network.

12. The device of claim 11 , wherein extracting the two-dimensional temporal feature map comprises:

segmenting the media stream into a plurality of clips;

extracting features of respective ones of the plurality of clips to obtain a feature map of the media stream; and

extracting, from features of one or more clips corresponding to a moment of the plurality of moments in the feature map of the media stream, features of this moment as a part of the two-dimensional temporal feature map.

13. The device of claim 11 , wherein determining the correlation comprises:

sampling the plurality of moments at respective sample rates to determine a plurality of candidate moments, wherein the sample rates are adaptively adjusted based on lengths of respective ones of the plurality of moments; and

determining a correlation between the plurality of candidate moments and the action in the media stream.

14. At least one non-transitory machine-readable medium comprising computer-executable instructions which, when executed by a device, cause the device to perform operations to:

extract, from a media stream, a two-dimensional temporal feature map representing a plurality of moments within the media stream, wherein the two-dimensional temporal feature map comprises a first dimension representing a start of a respective one of the plurality of moments and a second dimension representing an end of a respective one of the plurality of moments;

encode a sentence feature extracted from an input;

fuse the encoded sentence feature with the two-dimensional temporal feature map into a unified subspace as a fused two-dimensional temporal map;

apply a convolutional layer to the two-dimensional temporal feature map to obtain a further feature map having a same dimension as the two-dimensional temporal feature map, the convolutional layer comprises a dilated convolution and strides of the dilated convolution are configured to increase as lengths of the respective moments increase;

generate a temporal adjacent network using the fused two-dimensional temporal map and the further feature map;

determine, using the temporal adjacent network, a correlation between the plurality of moments and an action in the media stream; and

identify a matching a set of candidate moments for the input using the temporal adjacent network.

15. The at least one non-transitory machine-readable medium of claim 14 , the instructions to extract the two-dimensional temporal feature map comprising instructions to:

segment the media stream into a plurality of clips;

extract features of respective ones of the plurality of clips to obtain a feature map of the media stream; and

extract, from features of one or more clips corresponding to a moment of the plurality of moments in the feature map of the media stream, features of this moment as a part of the two-dimensional temporal feature map.

16. The at least one non-transitory machine-readable medium of claim 14 , the instructions to determine the correlation comprising instructions to:

sample the plurality of moments at respective sample rates to determine a plurality of candidate moments, wherein the sample rates are adaptively adjusted based on lengths of respective ones of the plurality of moments; and

determine a correlation between the plurality of candidate moments and the action in the media stream.

17. The at least one non-transitory machine-readable medium of claim 16 , wherein the sample rates are configured to decrease as the lengths of the respective moments increase.

18. The at least one non-transitory machine-readable medium of claim 14 , the instructions to determine the correlation comprising instructions to:

applying a convolutional layer to the two-dimensional temporal feature map to obtain a further feature map having a same dimension as the two-dimensional temporal feature map; and

determining, based on the further feature map, scores of correlation between the plurality of moments and the action in the media stream.

19. The device of claim 11 , wherein determining the correlation comprises:

in response to receiving a query for a particular action in the media stream, extracting a feature vector of the query; and

determining the correlation based on the feature vector of the query and the two-dimensional temporal feature map.

20. The at least one non-transitory machine-readable medium of claim 14 , the instructions to determine the correlation further comprising instructions to:

in response to receipt of a query for a particular action in the media stream, extract a feature vector of the query; and

determine the correlation based on the feature vector of the query and the two-dimensional temporal feature map.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Apr 13, 2022
From: PENG, HOUWEN; FU, JIANLONG
To: MICROSOFT TECHNOLOGY LICENSING, LLC
Reel/Frame 059590/0580 →
Priority Claims (1)
CN 201911059082.6 · Nov 1, 2019 · national
Continuity (1)
Related Publication 20230351752A1 · Nov 2, 2023
References Cited (67)
US 9244924B2 · Cheng et al. · 2016 [cited by applicant]
US 20130259390A1 · Dunlop et al. · 2013 [cited by applicant]
US 20170270203A1 · Liu · 2017 [cited by examiner]
US 20190114485A1 · Chan · 2019 [cited by examiner]
US 20200193117A1 · Raff · 2020 [cited by examiner]
US 20200213671A1 · Major · 2020 [cited by examiner]
US 20200327160A1 · Hsieh · 2020 [cited by examiner]
US 20210109966A1 · Ayush · 2021 [cited by examiner]
US 20230108883A1 · Lovell · 2023 [cited by examiner]
US 20240412497A1 · Athar · 2024 [cited by examiner]
CN 102203761A · 2011 [cited by applicant]
CN 102687518A · 2012 [cited by applicant]
CN 109564576A · 2019 [cited by applicant]
CN 110019849A · 2019 [cited by applicant]
Hendricks et al: Localizing Moments in Video with Natural Language (Year: 2017). [cited by examiner]
Zhang, et al., “Learning 2D Temporal Adjacent Networks for Moment Localization with Natural Language”, In Repository of arXiv: 1912.03590v1, Dec. 8, 2019, 9 Pages. (Year: 2019). [cited by examiner]
Chen, et al., “Semantic Proposal for Activity Localization in Videos via Sentence Query”, In Proceedings of Thirty-Third AAAI Conference on Artificial Intelligence, vol. 33, Issue 1, Jul. 17, 2019, pp. 8199-8206. [cited by applicant]
Chen, et al., “Temporally Grounding Natural Sentence in Video”, In Proceedings of the Conference on Empirical Methods in Natural Language Processing, Oct. 31, 2018, pp. 162-171. [cited by applicant]
Chu, et al., “Video Co-Summarization: Video Summarization by Visual Co-Occurrence”, In Proceedings of Conference on Computer Vision and Pattern Recognition, Jun. 7, 2015, pp. 3584-3592. [cited by applicant]
Escorcia, et al., “Temporal Localization of Moments in Video Collections with Natural Language”, In Repository of arXiv:1907.12763v1, Jul. 30, 2019, 14 Pages. [cited by applicant]
Gaidon, et al., “Temporal Localization of Actions with Actoms”, In Journal of IEEE Transactions on Pattern Analysis and Machine Intelligence. vol. 35, Issue 11, Nov. 2013, pp. 2782-2795. [cited by applicant]
Gao, et al., “TALL: Temporal Activity Localization via Language Query”, In Proceedings of International Conference on Computer Vision, Oct. 22, 2017, pp. 5277-5285. [cited by applicant]
Ge, et al., “MAC: Mining Activity Concepts for Language-based Temporal Localization”, In Proceedings of Winter Conference on Applications of Computer Vision, Jan. 7, 2019, pp. 245-253. [cited by applicant]
Gella, et al., “A Dataset for Telling the Stories of Social Media Videos”, In Proceedings of the Conference on Empirical Methods in Natural Language Processing, Oct. 31, 2018, pp. 968-974. [cited by applicant]
Hahn, et al., “Tripping through time: Efficient Localization of Activities in Videos”, In Repository of arXiv:1904.09936v1, Apr. 22, 2019, 10 Pages. [cited by applicant]
Hasan, et al., “Learning Temporal Regularity in Video Sequences”, In Proceedings of Conference on Computer Vision and Pattern Recognition, Jun. 27, 2016, pp. 733-742. [cited by applicant]
Hendricks, et al., “Localizing Moments in Video with Natural Language”, In Proceedings of International Conference on Computer Vision, Oct. 22, 2017, pp. 5804-5813. [cited by applicant]
Hendricks, et al., “Localizing Moments in Video with Temporal Language”, In Proceedings of the Conference on Empirical Methods in Natural Language Processing, Oct. 31, 2018, pp. 1380-1390. [cited by applicant]
Hochreiter, et al., “Long Short-Term Memory”, In Journal of Neural Computation vol. 9, Issue 8, Nov. 15, 1997, pp. 1735-1780. [cited by applicant]
Jiang, et al., “Cross-Modal Video Moment Retrieval with Spatial and Language-Temporal Attention”, In Proceedings of International Conference on Multimedia Retrieval, Jun. 10, 2019, pp. 217-225. [cited by applicant]
Kingma, et al., “Adam: A Method for Stochastic Optimization”, In Repository of arXiv:1412.6980v1, Dec. 22, 2014, 9 Pages. [cited by applicant]
Krishna, et al., “Dense-Captioning Events in Videos”, In Proceedings of International Conference on Computer Vision, Oct. 22, 2017, pp. 706-715. [cited by applicant]
Lei, et al., “TVQA: Localized, Compositional Video Question Answering”, In Proceedings of Conference on Empirical Methods in Natural Language Processing, Oct. 31, 2018, pp. 1369-1379. [cited by applicant]
Li, et al., “Global-Local Temporal Representations For Video Person Re-Identification”, In Repository of arXiv:1908.10049v1, Aug. 27, 2019, 10 Pages. [cited by applicant]
Lin, et al., “BMN: Boundary-Matching Network for Temporal Action Proposal Generation”, In Repository of arXiv:1907.09702v1, Jul. 23, 2019, 10 Pages. [cited by applicant]
Lin, et al., “Single Shot Temporal Action Detection”, In Proceedings of 25th ACM International Conference on Multimedia, Oct. 23, 2017, pp. 988-996. [cited by applicant]
Liu, et al., “Attentive Moment Retrieval in Videos”, In Proceedings of 41st International ACM SIGIR Conference on Research & Development in Information Retrieval, Jul. 8, 2018, pp. 15-24. [cited by applicant]
Liu, et al., “Cross-modal Moment Localization in Videos”, In Proceedings of the 26th ACM International Conference on Multimedia, Oct. 22, 2018, pp. 843-851. [cited by applicant]
Liu, et al., “Temporal Modular Networks for Retrieving Complex Compositional Activities in Videos”, In Proceedings of 15th European Conference on Computer Vision, Sep. 8, 2018, pp. 569-586. [cited by applicant]
“International Search Report and the Written Opinion Issued in PCT Application No. PCT/US20/056390”, Mailed Date: Feb. 9, 2021, 12 Pages. [cited by applicant]
Pennington, et al., “GloVe: Global Vectors for Word Representation”, In Proceedings of Conference on Empirical Methods in Natural Language Processing, Oct. 25, 2014, pp. 1532-1543. [cited by applicant]
Regneri, et al., “Grounding Action Descriptions in Videos”, In Journal of Transactions of the Association for Computational Linguistics, vol. 1, Mar. 1, 2013, pp. 25-36. [cited by applicant]
Rohrbach, et al., “Script Data for Attribute-Based Recognition of Composite Activities”, In Proceedings of 12th European Conference on Computer Vision, Oct. 7, 2012, pp. 144-157. [cited by applicant]
Shao, et al., “Find and Focus: Retrieve and Localize Video Events with Natural Language Queries”, In Proceedings of 15th European Conference on Computer Vision, Sep. 8, 2018, pp. 202-218. [cited by applicant]
Sigurdsson, et al., “Hollywood in Homes: Crowdsourcing Data Collection for Activity Understanding”, In Proceedings of 14th European Conference on Computer Vision, Oct. 11, 2016, pp. 510-526. [cited by applicant]
Simonyan, et al., “Very Deep Convolutional Networks for Large-Scale Image Recognition”, In Repository of arXiv:1409.1556v6, Apr. 10, 2015, 14 Pages. [cited by applicant]
Song, et al., “TVSum: Summarizing Web Videos using Titles”, In Proceedings of Conference on Computer Vision and Pattern Recognition, Jun. 7, 2015, pp. 5179-5187. [cited by applicant]
Song, et al., “VAL: Visual-Attention Action Localizer”, In Proceedings of 19th Pacific-Rim Conference on Multimedia, Sep. 21, 2018, pp. 340-350. [cited by applicant]
Tran, et al., “Learning Spatiotemporal Features with 3D Convolutional Networks”, In Proceedings of International Conference on Computer Vision, Dec. 7, 2015, pp. 4489-4497. [cited by applicant]
Vaswani, et al., “Attention Is All You Need”, In Proceedings of 31st Conference on Neural Information Processing Systems, Dec. 4, 2017, 11 Pages. [cited by applicant]
Vijayanarasimhan, et al., “Capturing Special Video Moments with Google Photos”, Retrieved from: https://ai.googleblog.com/2019/04/capturing-special-video-moments-with.html, Apr. 3, 2019, 5 Pages. [cited by applicant]
Wang, et al., “Language-Driven Temporal Activity Localization: A Semantic Matching Reinforcement Learning Model”, In Proceedings of Conference on Computer Vision and Pattern Recognition, Jun. 15, 2019, pp. 334-343. [cited by applicant]
Wu, et al., “Adaptive Graph Representation Learning for Video Person Re-identification”, In Repository of arXiv:1909.02240v1, Sep. 5, 2019, 10 Pages. [cited by applicant]
Wu, et al., “Multi-modal Circulant Fusion for Video-to-Language and Backward”, In Proceedings of the Twenty-Seventh International Joint Conference on Artificial Intelligence, Jul. 13, 2018, pp. 1029-1035. [cited by applicant]
Xu, et al., “Multilevel Language and Vision Integration for Text-to-Clip Retrieval”, In Proceedings of Thirty-Third AAAI Conference on Artificial Intelligence, vol. 33, Issue 1, Jul. 17, 2019, pp. 9062-9069. [cited by applicant]
Yuan, et al., “To Find Where You Talk: Temporal Sentence Localization in Video with Attention Based Location Regression”, In Proceedings of Thirty-Third AAAI Conference on Artificial Intelligence, vol. 33, Issue 1, Jul.… [cited by applicant]
Zhang, et al., “Cross-Modal Interaction Networks for Query-Based Moment Retrieval in Videos”, In Proceedings of the 42nd International ACM SIGIR Conference on Research and Development in Information Retrieval, Jul. 21, … [cited by applicant]
Zhang, et al., “Exploiting Temporal Relationships in Video Moment Localization with Natural Language”, In Proceedings of 27th ACM International Conference on Multimedia, Oct. 21, 2019, pp. 1230-1238. [cited by applicant]
Zhang, et al., “Learning 2D Temporal Adjacent Networks for Moment Localization with Natural Language”, In Repository of arXiv:1912.03590v1, Dec. 8, 2019, 9 Pages. [cited by applicant]
Zhang, et al., “MAN: Moment Alignment Network for Natural Language Moment Retrieval via Iterative Graph Adjustment”, In Proceedings of Conference on Computer Vision and Pattern Recognition, Jun. 15, 2019, pp. 1247-1257. [cited by applicant]
Zhang, et al., “Temporal Reasoning Graph for Activity Recognition”, In Repository of arXiv:1908.09995v1, Aug. 27, 2019, 14 Pages. [cited by applicant]
Zhao, et al., “Temporal Action Detection with Structured Segment Networks”, In Proceedings of International Conference on Computer Vision, Oct. 22, 2017, pp. 2933-2942. [cited by applicant]
First Office Action Received for Chinese Application No. 201911059082.6, mailed on Jan. 5, 2024, 12 pages. (English Translation Provided). [cited by applicant]
Communication pursuant to article 94(3) received for European Application No. 20804727.4, mailed on Jul. 12, 2024, 06 pages. [cited by applicant]
Third Office Action Received for Chinese Application No. 201911059082.6, mailed on Nov. 6, 2024, 20 pages. (English Translation Provided). [cited by applicant]
Second Office Action Received for Chinese Application No. 201911059082.6, mailed on Aug. 17, 2024, 15 pages. (English Translation Provided). [cited by applicant]
Notice of Allowance Received for Chinese Application No. 201911059082.6, mailed on Mar. 3, 2025, 06 pages. (English Translation Provided). [cited by applicant]