IP Library › Granted Patent US 12,130,891
Granted Patent B2
US 12,130,891 · App. 17/402,877 · Granted Oct 29, 2024

Method of live video event detection based on natural language queries, and an apparatus for the same

Inventors: Ning Ye (Toronto, CA); Zhiming Hu (Toronto, CA); Caleb Ryan Phillips (Toronto, CA); Iqbal Ismail Mohomed (Toronto, CA)
Assignee: SAMSUNG ELECTRONICS CO., LTD.
G06F18/22G06F16/7343G06F18/214G06N20/00G06V20/46G06V20/44
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,130,891
App. No.
17/402,877
Granted
Oct 29, 2024
Kind
B2
Abstract

A method of real-time video event detection includes: obtaining, based on a natural language query, a query vector; performing multimodal feature extraction on a video stream to obtain a video vector, obtaining a similarity score by comparing the query vector to the video vector; comparing the similarity score to a predetermined threshold; and activating, based on the similarity score being above the predetermined threshold, an action trigger. The multimodal feature extraction is performed using a plurality of overlapping windows that include sequential frames of the video stream.

Claims (69)

1. A method of real-time video event detection comprising:

obtaining, based on a natural language query, a query vector,

performing multimodal feature extraction on a video stream to obtain a video vector,

obtaining a similarity score by comparing the query vector to the video vector;

comparing the similarity score to a predetermined threshold; and

activating, based on the similarity score being above the predetermined threshold, an action trigger,

wherein the performing of the multimodal feature extraction is performed using a plurality of overlapping windows that include sequential frames of the video stream,

wherein the performing of the multimodal feature extraction comprises:

obtaining a latency constraint for the performing of the multimodal feature extraction; and

selecting a plurality of final feature extractors, among a plurality of predetermined feature extractors corresponding to a plurality of modalities, based on the latency constraint, predetermined performances of the plurality of predetermined feature extractors, and predetermined latencies of the plurality of predetermined feature extractors.

2. The method according to claim 1 , wherein each of the plurality of overlapping windows begins at a same starting frame of the video stream.

3. The method according to claim 1 , further comprising:

obtaining a plurality of first sub-video vectors that correspond to the plurality of overlapping windows;

obtaining a similarity score between each of the plurality of first sub-video vectors and the query vector; and

selecting a first sub-video vector having a highest similarity score as the video vector.

4. The method according to claim 3 , further comprising, in response to selection of the video vector, obtaining a plurality of second sub-video vectors that correspond to overlapping windows which begin at a frame following a last frame of a window corresponding to the first sub-video vector.

5. The method according to claim 1 , wherein the selecting the plurality of final feature extractors comprises selecting only a single feature extractor from any of the plurality of modalities.

6. The method according to claim 1 , wherein the performing of the multimodal feature extraction is performed by a deep learning based model, and

the deep learning based model is trained based on training sets in which information provided by one randomly selected feature extractor is masked.

7. The method according to claim 1 , further comprising:

obtaining, based on a plurality of ordered natural language queries, a plurality of ordered query vectors;

obtaining a plurality of similarity scores by comparing only a sequential portion of the plurality of ordered query vectors and the video vector; and

activating, based on one of the plurality of similarity scores being above the predetermined threshold, the action trigger.

8. An apparatus for real-time video event detection, the apparatus comprising:

a memory storing one or more instructions; and

at least one processor configured to execute the one or more instructions to:

obtain, based on a natural language query, a query vector;

perform multimodal feature extraction on a video stream to obtain a video vector,

obtain a similarity score by comparing the query vector to the video vector;

compare the similarity score to a predetermined threshold; and

activate, based on the similarity score being above the predetermined threshold, an action trigger,

wherein the multimodal feature extraction is performed using a plurality of overlapping windows that include sequential frames of the video stream,

wherein the at least one processor is further configured to execute the one or more instructions to:

obtain a latency constraint for performing the multimodal feature extraction; and

select a plurality of final feature extractors, among a plurality of predetermined feature extractors corresponding to a plurality of modalities, based on the latency constraint, predetermined performances of the plurality of predetermined feature extractors, and predetermined latencies of the plurality of predetermined feature extractors.

9. The apparatus according to claim 8 , wherein each of the plurality of overlapping windows begins at a same starting frame of the video stream.

10. The apparatus according to claim 8 , wherein the at least one processor is further configured to execute the one or more instructions to:

obtain a plurality of first sub-video vectors that correspond to the plurality of overlapping windows;

obtain a similarity score between each of the plurality of first sub-video vectors and the query vector; and

select a first sub-video vector having a highest similarity score as the video vector.

11. The apparatus according to claim 10 , wherein the at least one processor is further configured to execute the one or more instructions to obtain, in response to selection of the video vector, a plurality of second sub-video vectors that correspond to overlapping windows which begin at a frame following a last frame of a window corresponding to the first sub-video vector.

12. The apparatus according to claim 8 , wherein the selecting the plurality of final feature extractors comprises selecting only a single feature extractor from any of the plurality of modalities.

13. The apparatus according to claim 8 , wherein the multimodal feature extraction is performed by a deep learning based model, and

the deep learning based model is trained based on training sets in which information provided by one randomly selected feature extractor is masked.

14. The apparatus according to claim 8 , wherein the at least one processor is further configured to execute the one or more instructions to:

obtain, based on a plurality of ordered natural language queries, a plurality of ordered query vectors;

obtain a plurality of similarity scores by comparing only a sequential portion of the plurality of ordered query vectors and the video vector; and

activate, based on one of the plurality of similarity scores being above the predetermined threshold, the action trigger.

15. A non-transitory computer-readable medium storing instructions, the instructions comprising: one or more instructions that, when executed by one or more processors, cause the one or more processors to:

obtain, based on a natural language query, a query vector;

perform multimodal feature extraction on a video stream to obtain a video vector,

obtain a similarity score by comparing the query vector to the video vector;

compare the similarity score to a predetermined threshold; and

activate, based on the similarity score being above the predetermined threshold, an action trigger,

wherein the multimodal feature extraction is performed using a plurality of overlapping windows that include sequential frames of the video stream,

wherein the instructions further cause the one or more processors to:

obtain a latency constraint for performing the multimodal feature extraction; and

select a plurality of final feature extractors, among a plurality of predetermined feature extractors corresponding to a plurality of modalities, based on the latency constraint, predetermined performances of the plurality of predetermined feature extractors, and predetermined latencies of the plurality of predetermined feature extractors,

and

wherein the selecting the plurality of final feature extractors comprises selecting only a single feature extractor from any of the plurality of modalities.

16. The non-transitory computer-readable medium of claim 15 , wherein each of the plurality of overlapping windows begins at a same starting frame of the video stream, and wherein the instructions further cause the one or more processors to:

obtain a plurality of first sub-video vectors that correspond to the plurality of overlapping windows;

obtain a similarity score between each of the plurality of first sub-video vectors and the query vector;

select a first sub-video vector having a highest similarity score as the video vector; and

obtain, in response to selection of the video vector, a plurality of second sub-video vectors that correspond to overlapping windows which begin at a frame following a last frame of a window corresponding to the first sub-video vector.

17. The non-transitory computer-readable medium of claim 15 , wherein the instructions further cause the one or more processors to:

obtain, based on a plurality of ordered natural language queries, a plurality of ordered query vectors;

obtain a plurality of similarity scores by comparing only a sequential portion of the plurality of ordered query vectors and the video vector; and

activate, based on one of the plurality of similarity scores being above the predetermined threshold, the action trigger.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Aug 16, 2021
From: YE, NING; HU, ZHIMING; PHILLIPS, CALEB RYAN; MOHOMED, IQBAL ISMAIL
To: SAMSUNG ELECTRONICS CO., LTD.
Reel/Frame 057188/0982 →
Continuity (2)
Provisional Application 63110019 · Nov 5, 2020
Related Publication 20220138489A1 · May 5, 2022