IP Library Granted Patent US 10,402,655
Granted Patent B2
US 10,402,655 · App. 15/388,666 · Granted Sep 3, 2019

System and method for visual event description and event analysis

Inventors: Mehrsan Javan Roshtkhari (Montréal, CA); Martin Levine (Montréal, CA)
Assignee: Sportlogiq Inc.
G06K9/00718G06F16/7328G06F16/7847G06K9/00758G06K9/00765G06K9/4676G06K9/6215G06K9/6219G06K9/6224G06K9/6284
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 10,402,655
App. No.
15/388,666
Granted
Sep 3, 2019
Kind
B2
Abstract

A system and method are provided for analyzing a video. The method comprises: sampling the video to generate a plurality of spatio-temporal video volumes; clustering similar ones of the plurality of spatio-temporal video volumes to generate a low-level codebook of video volumes; analyzing the low-level codebook of video volumes to generate a plurality of ensembles of volumes surrounding pixels in the video; and clustering the plurality of ensembles of volumes by determining similarities between the ensembles of volumes, to generate at least one high-level codebook. Multiple high-level codebooks can be generated by repeating steps of the method. The method can further include performing visual event retrieval by using the at least one high-level codebook to make an inference from the video, for example comparing the video to a dataset and retrieving at least one similar video, activity and event labeling, and performing abnormal and normal event detection.

Claims (38)

1. A method of analyzing a video, the method comprising:

sampling the video to generate a plurality of spatio-temporal video volumes, each spatio-temporal video volume corresponding to a three-dimensional volume around a pixel in the video comprising a two-dimensional area around the pixel and a depth in time, to capture local information in space and time around the pixel;

clustering similar ones of the plurality of spatio-temporal video volumes to generate a low-level codebook of video volumes;

analyzing the low-level codebook of video volumes to generate a plurality of ensembles of volumes surrounding pixels in the video; and

clustering the plurality of ensembles of volumes by determining similarities between the ensembles of volumes, to generate at least one high-level codebook.

2. The method of claim 1 , further comprising generating multiple high-level codebooks by repeating the analyzing and clustering using spatial and temporal contextual structures.

3. The method of claim 1 , wherein the similarities between the ensembles of volumes are determined using a probabilistic model.

4. The method of claim 3 , wherein the probabilistic model utilizes a star graph model.

5. The method of claim 1 , further comprising removing non-informative regions from the at least one high-level codebook.

6. The method of claim 5 , wherein the non-informative regions comprise at least one background region in the video.

7. The method of claim 1 , wherein each ensemble of volumes is characterized by a set of video volumes, a central video volume, and a relative distance of each of the volumes in the ensemble to the central video volume.

8. The method of claim 1 , wherein the clustering is performed using a spectral clustering method.

9. The method of claim 1 , wherein the codebooks comprise bags of visual words.

10. The method of claim 9 , wherein the high-level codebook provides a multi-level hierarchical bag of visual words.

11. The method of claim 1 , further comprising performing visual event retrieval by using the at least one high-level codebook to make an inference from the video.

12. The method of claim 11 , wherein the visual event retrieval comprises comparing the video to a dataset and retrieving at least one similar video.

13. The method of claim 12 , wherein the comparison determines videos in the dataset comprising similar events to the video.

14. The method of claim 12 , further comprising generating a similarity map between the video and at least one video stored in the dataset.

15. The method of claim 14 , wherein the similarity map is generated using a pre-trained hierarchical bag of video words.

16. The method of claim 11 , wherein the visual event retrieval comprises activity and event labeling.

17. The method of claim 16 , further comprising generating a similarity map between the video and at least one video stored in the dataset.

18. The method of claim 17 , wherein the similarity map is generated using a pre-trained hierarchical bag of video words.

19. The method of claim 11 , wherein the visual event retrieval comprises performing abnormal and normal event detection.

20. The method of claim 19 , further comprising performing a decomposition of contextual information in the at least one high-level codebook.

21. The method of claim 19 , further comprising generating a similarity map between the video and the previously observed frames in the same video.

22. The method of claim 19 , further comprising generating a similarity map between the video and at least one video stored in the dataset.

23. The method of claim 21 , wherein the similarity map is generated using training data comprising a video database.

24. The method of claim 19 , further comprising performing online model updating.

25. A non-transitory computer readable medium comprising computer executable instructions for analyzing a video, comprising instructions for:

sampling the video to generate a plurality of spatio-temporal video volumes, each spatio-temporal video volume corresponding to a three-dimensional volume around a pixel in the video comprising a two-dimensional area around the pixel and a depth in time, to capture local information in space and time around the pixel;

clustering similar ones of the plurality of spatio-temporal video volumes to generate a low-level codebook of video volumes;

analyzing the low-level codebook of video volumes to generate a plurality of ensembles of volumes surrounding pixels in the video; and

clustering the plurality of ensembles of volumes by determining similarities between the ensembles of volumes, to generate at least one high-level codebook.

26. A video processing system comprising a processor, an interface for receiving videos, and memory, the memory comprising computer executable instructions for analyzing a video, comprising instructions for:

sampling the video to generate a plurality of spatio-temporal video volumes, each spatio-temporal video volume corresponding to a three-dimensional volume around a pixel in the video comprising a two-dimensional area around the pixel and a depth in time, to capture local information in space and time around the pixel;

clustering similar ones of the plurality of spatio-temporal video volumes to generate a low-level codebook of video volumes;

analyzing the low-level codebook of video volumes to generate a plurality of ensembles of volumes surrounding pixels in the video; and

clustering the plurality of ensembles of volumes by determining similarities between the ensembles of volumes, to generate at least one high-level codebook.

Assignments (3)
SECURITY INTEREST Recorded Feb 25, 2026
From: SPORTLOGIQ INC.
To: CANADIAN IMPERIAL BANK OF COMMERCE IN ITS CAPACITY AS ADMINISTRATIVE AGENT
Reel/Frame 073891/0430 →
NUNC PRO TUNC ASSIGNMENT Recorded Dec 22, 2016
From: LEVINE, MARTIN; JAVAN ROSHTKHARI, MEHRSAN
To: THE ROYAL INSTITUTION FOR THE ADVANCEMENT OF LEARNING/MCGILL UNIVERSITY
Reel/Frame 040754/0283 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Dec 22, 2016
From: THE ROYAL INSTITUTION FOR THE ADVANCEMENT OF LEARNING/MCGILL UNIVERSITY
To: SPORTLOGIQ INC.
Reel/Frame 040754/0375 →
Continuity (3)
Continuation PCTCA2015050569 · Jun 19, 2015
Provisional Application 62016133 · Jun 24, 2014
Related Publication 20170103264A1 · Apr 13, 2017
Cited By (6)
US 12,450,902 US 12,475,694 US 12,561,365 US 12,586,425 US 12,670,567 US 12,688,587