IP Library Granted Patent US 12,222,985
Granted Patent B2
US 12,222,985 · App. 17/883,494 · Granted Feb 11, 2025

Video retrieval method and apparatus using vectorized segmented videos based on key frame detection

Inventors: Seung Joon Lee (Seoul, KR); Sung Jun Kim (Seoul, KR); Raehyuk Jung (Seoul, KR); Haram Jo (Jeollanam-do, KR)
Assignee: Twelve Labs, Inc.
G06F16/7837G06V10/82G06V20/46G06V20/49
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,222,985
App. No.
17/883,494
Granted
Feb 11, 2025
Kind
B2
Abstract

In order to implement the foregoing object, an exemplary embodiment of the present disclosure discloses a video retrieval method performed by a computing device. The video retrieval method may include: segmenting one or more video data into two or more unit video data; encoding, by one or more encoders comprised in a machine learning enabled key frame detection module, two or more unit video data; generating one or more key frame detection vectors for each of the two or more unit video data based on the result of the encoding; and generating feature vector of one or more retrieval video data based on a combination of one or more vectors among the key frame detection vectors.

Claims (29)

1. A video retrieval method performed by a computing device, the video retrieval method comprising:

segmenting one or more video data into two or more unit video data;

encoding, by one or more encoders comprised in a machine learning enabled key frame detecting module, two or more unit video data;

generating one or more key frame detection vectors for each of the two or more unit video data based on the result of the encoding; and

generating feature vector of one or more target retrieval video data based on a combination of one or more vectors among the key frame detection vectors.

2. The video retrieval method of claim 1 , further comprising:

identifying key frame information among the two or more unit video data based on the one or more key frame detection vectors.

3. The video retrieval method of claim 2 ,

wherein the encoding the two or more unit video data includes:

generating, by one or more unit video data encoding modules, one or more unit video data sub tokens for each of the two or more unit video data; and

generating, by the key frame detection module, the one or more key frame detection vectors for the unit video data based on the one or more unit video data sub tokens.

4. The video retrieval method of claim 2 , further comprising:

generating the one or more target retrieval video data by grouping the two or more unit video data based on the identified key frame information.

5. The video retrieval method of claim 1 , wherein each of the one or more target retrieval video data comprises unit video data grouped based on the values of the unit video data sub tokens.

6. The video retrieval method of claim 5 , wherein the each of the one or more target retrieval video data comprises two or more temporally continuous unit video data.

7. The video retrieval method of claim 6 , wherein the each of the one or more target retrieval video data comprises at least one unit video data identified as key frame, and the unit video data identified as key frame is temporally the most preceding or temporally the most trailing among the two or more temporally continuous unit video data.

8. A non-transitory computer readable storage medium storing a computer program, in which when the computer program is executed in one or more processors, the computer program causes the one or more processors to perform operations for performing a video retrieval method, the video retrieval method comprising:

segmenting one or more video data into two or more unit video data;

encoding, by one or more encoders comprised in a machine learning enabled key frame detection module, two or more unit video data;

generating one or more key frame detection vectors for each of the two or more unit video data based on the encoding result; and

generating feature vector of one or more retrieval video data based on a combination of one or more vectors among the key frame detection vectors.

9. A computing device performing a video retrieval method, the computing device comprising:

a processor including at least one core; and

a memory including program codes executable in the processor,

wherein the processor:

segments one or more video data into two or more unit video data;

encodes, by one or more encoders comprised in a machine learning enabled key frame detection module, two or more unit video data;

generates one or more key frame detection vectors for each of the two or more unit video data based on the encoding result; and

generates feature vector of one or more target retrieval video data based on a combination of one or more vectors among the key frame detection vectors.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Aug 9, 2022
From: LEE, SEUNG JOON; KIM, SUNG JUN; JUNG, RAEHYUK; JO, HARAM
To: TWELVE LABS, INC.
Reel/Frame 060763/0214 →
Continuity (2)
Provisional Application 63317359 · Mar 7, 2022
Related Publication 20230306056A1 · Sep 28, 2023
References Cited (29)
US 8645397B1 · Koudas et al. · 2014 [cited by applicant]
US 20190138798A1 · Tang et al. · 2019 [cited by applicant]
US 20210097290A1 · Yang · 2021 [cited by examiner]
US 20210109966A1 · Ayush · 2021 [cited by examiner]
US 20210224550A1 · Zeng et al. · 2021 [cited by applicant]
US 20220121636A1 · Zheng et al. · 2022 [cited by applicant]
US 20220138489A1 · Ye et al. · 2022 [cited by applicant]
US 20220277566A1 · Yang et al. · 2022 [cited by applicant]
CN 104376003A · 2015 [cited by examiner]
CN 105654054A · 2016 [cited by examiner]
CN 107506370A · 2017 [cited by examiner]
CN 108228915A · 2018 [cited by examiner]
CN 111339369A · 2020 [cited by examiner]
CN 111950653A · 2020 [cited by examiner]
CN 112395457A · 2021 [cited by examiner]
CN 113656639A · 2021 [cited by examiner]
Sun et al., “VSRNet: End-to-end video segment retrieval with text query” (Year: 2021). [cited by examiner]
Yan et al., “Self-Supervised Learning to Detect Key Frames in Videos” (Year: 2020). [cited by examiner]
Zhong et al., “Key Frame Extraction Algorithm of Motion Video Based on Priori” (Year: 2020). [cited by examiner]
Yan et al., “Deep Keyframe Detection in Human Action Videos” (Year: 2018). [cited by examiner]
Sze et al., “A New Key Frame Representation for Video Segment Retrieval” (Year: 2005). [cited by examiner]
United States Office Action, U.S. Appl. No. 17/883,489, Nov. 16, 2023, 21 pages. [cited by applicant]
Chen, S., et al., “Shot Contrastive Self-Supervised Learning for Scene Boundary Detection,” Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2021, pp. 9796-9805. [cited by applicant]
Dosovitskiy, A., et al., “An Image is Worth 16×16 Words: Transformers for Image Recognition at Scale,” arXiv preprint arXiv:2010.11929, 2020, 22 pages. [cited by applicant]
Gabeur, V., et al., “Multi-modal transformer for video retrieval,” Computer Vision—ECCV 2020: 16th European Conference, Glasgow, UK, Aug. 23-28, 2020, Proceedings, Part IV 16, 18 pages. [cited by applicant]
Shou, MZ, et al., “Generic event boundary detection: A benchmark for event segmentation,” [cited by applicant]
Souček, T., et al., “Transnet V2: an effective deep network architecture for fast shot transition detection,” arXiv preprint arXiv:2008.04838 (2020), 4 pages. [cited by applicant]
Vaswani, A., et al., “Attention Is All You Need,” 31st Conference on Neural Information Processing Systems (NIPS 2017 NIPS), 15 pages. [cited by applicant]
United States Office Action, U.S. Appl. No. 17/883,489, Apr. 21, 2023, 14 pages. [cited by applicant]
Cited By (1)
US 12,347,462