IP Library Granted Patent US 12,469,287
Granted Patent B2
US 12,469,287 · App. 18/048,380 · Granted Nov 11, 2025

Computer-implemented method and non- transitory computer-readable medium for generating a thumbnail from a video stream or file, and video surveillance system

Inventor: Morten Engel Kristiansen (Brøndby, DK)
Assignee: MILESTONE SYSTEMS A/S
G06V20/46G01S15/42G06F3/16G06V20/44G06V20/52
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,469,287
App. No.
18/048,380
Granted
Nov 11, 2025
Kind
B2
Abstract

A computer-implemented method of generating a thumbnail of a video stream or file of a surveillance area in a video management system, including setting, in the video management system, at least one sound event to be detected in at least one audio stream or file corresponding to the video stream or file; detecting, in the at least one audio stream or file, at least one point in time at which the at least one sound event occurs; generating the thumbnail based on at least part of at least one frame of the video stream or file, wherein the frame is selected based on the point in time; and displaying the thumbnail in the video management system.

Claims (79)

1 . A computer-implemented method of generating a thumbnail of a video stream or a thumbnail of a file of a surveillance area in a video management system, comprising:

setting, in the video management system, at least one sound event to be detected in at least one audio stream or file corresponding to the video stream or file;

detecting, in the at least one audio stream or file, at least one point in time at which the at least one sound event occurs;

generating the thumbnail based on at least part of at least one frame of the said video stream or the thumbnail of the file, wherein the frame is selected based on the point in time;

determining a distance between a sound source that causes the sound event in the surveillance area and a video capturing device from which the video stream or file originates;

selecting the frame based on the point in time and the distance; and

displaying the thumbnail in the video management system,

wherein determining the distance comprises using passive acoustic location by triangulating a location of the sound source using different audio capturing devices,

wherein when a sound arrives at one of the different audio capturing devices, the method further comprises waiting for a predetermined time before checking whether that sound has arrived at another one of the different audio capturing devices, and if so calculating the distance.

2 . The computer-implemented method according to claim 1 , comprising applying a time correction value to the point in time to define a revised point in time, wherein the frame corresponds to the revised point in time.

3 . The computer-implemented method according to claim 2 , wherein the time correction value is calculated based on the following formula:

time

correction

value

=

base

value

for

sound

event

-

distance

of

sound

source

from

video

capturing

device

speed

of

sound

wherein ‘base value for sound event’ corresponds to a predetermined value associated with a length, pitch or loudness of the sound event or a type of the sound event;

‘distance of sound source from video capturing device’ corresponds to the determined distance; and

‘speed of sound’ corresponds to a speed at which sound travels from the sound source towards the video capturing device.

4 . The computer-implemented method according to claim 3 , wherein the revised point in time is calculated based on the following formula:

revised point in time=point in time+time correction value.

5 . The computer-implemented method according to claim 3 , wherein when the sound event corresponds to a predetermined sound level in the audio stream or file, the base value is set to 0.

6 . The computer-implemented method according to claim 3 , further comprising setting the base value such that when the sound event to be detected in the video management system corresponds to an accident in the surveillance area, the time correction value is a positive number.

7 . The computer-implemented method according to claim 1 , wherein the sound event corresponds to a predetermined sound level or a change in a sound level in the audio stream or file.

8 . The computer-implemented method according to claim 1 , wherein the sound event corresponds to a type of sound.

9 . The computer-implemented method according to claim 1 , wherein the audio stream or file is captured by at least one audio capturing device and wherein the video stream or file is captured by a video camera, and wherein the audio capturing device is disposed for capturing sounds outside of a field-of-view of the video camera.

10 . The computer-implemented method according to claim 1 , further comprising using a motion detection algorithm for detecting at least one event of interest in the surveillance area and at least one audio capturing device for detecting the sound event.

11 . The computer-implemented method according to claim 1 , wherein the audio stream or file is captured by at least one audio capturing device and wherein the video stream or file is captured by a video camera, and wherein the at least one audio capturing device is attached to the video camera in the video management system and not attached to other video cameras in the video management system.

12 . The computer-implemented method according to claim 1 , wherein selecting the frame is based on the point in time, on the distance, and on a length, pitch or loudness of the sound event or a type of the sound event.

13 . The computer-implemented method according to claim 1 , wherein determining the distance comprises determining a focal distance between the video capturing device and a point in the surveillance area where the video capturing device focuses on and setting the distance to that focal distance.

14 . A non-transitory computer-readable medium storing a program that, when implemented by a video management system, causes the video management system to perform a method of generating a thumbnail of a video stream or a thumbnail of a file of a surveillance area in the video management system, the method comprising:

setting, in the video management system, at least one sound event to be detected in at least one audio stream or file corresponding to the video stream or file;

detecting, in the at least one audio stream or file, at least one point in time at which the at least one sound event occurs;

generating the thumbnail based on at least part of at least one frame of the video stream or the thumbnail of the file, wherein the frame is selected based on the point in time;

determining a distance between a sound source that causes the sound event in the surveillance area and a video capturing device from which the video stream or file originates;

selecting the frame based on the point in time and the distance; and

displaying the thumbnail in the video management system,

wherein determining the distance comprises using passive acoustic location by triangulating a location of the sound source using different audio capturing devices,

wherein when a sound arrives at one of the different audio capturing devices, the method further comprises waiting for a predetermined time before checking whether that sound has arrived at another one of the different audio capturing devices, and if so calculating the distance.

15 . A video surveillance system comprising a video management system, an apparatus configured to generate a thumbnail of a video stream or a thumbnail of a file of a surveillance area in the video management system, a plurality of video cameras and at least one audio capturing device, the apparatus comprising one or more processors configured to:

set, in the video management system, at least one sound event to be detected in at least one audio stream or file corresponding to the video stream or file;

detect, in the at least one audio stream or file, at least one point in time at which the at least one sound event occurs;

generate a thumbnail based on at least part of at least one frame of the video stream or the thumbnail of the file, wherein the frame is selected based on the point in time;

determine a distance between a sound source that causes the sound event in the surveillance area and a video capturing device from which the video stream or file originates;

select the frame based on the point in time and the distance; and

display the thumbnail in the video management system,

wherein determining the distance comprises using passive acoustic location by triangulating a location of the sound source using different audio capturing devices,

wherein when a sound arrives at one of the different audio capturing devices, the method further comprises waiting for a predetermined time before checking whether that sound has arrived at another one of the different audio capturing devices, and if so calculating the distance.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jun 27, 2024
From: KRISTIANSEN, MORTEN ENGEL
To: MILESTONE SYSTEMS A/S
Reel/Frame 068457/0144 →
Priority Claims (1)
GB 2115244 · Oct 22, 2021 · national
Continuity (1)
Related Publication 20230125724A1 · Apr 27, 2023
References Cited (36)
US 9643722B1 · Myslinski · 2017 [cited by examiner]
US 10120536B2 · Cha · 2018 [cited by examiner]
US 11064102B1 · Helpingstine · 2021 [cited by examiner]
US 11282353B1 · Fowler · 2022 [cited by examiner]
US 11417183B1 · Onofrio · 2022 [cited by examiner]
US 20120127308A1 · Eldershaw · 2012 [cited by examiner]
US 20130091432A1 · Shet · 2013 [cited by examiner]
US 20130176442A1 · Shuster · 2013 [cited by examiner]
US 20140362231A1 · Bietsch · 2014 [cited by examiner]
US 20150085114A1 · Ptitsyn · 2015 [cited by examiner]
US 20150106721A1 · Cha et al. · 2015 [cited by applicant]
US 20160292935A1 · Patron · 2016 [cited by examiner]
US 20170223302A1 · Conlan · 2017 [cited by examiner]
US 20170223314A1 · Collings, III · 2017 [cited by examiner]
US 20170228603A1 · Johnson · 2017 [cited by examiner]
US 20170323540A1 · Boykin · 2017 [cited by examiner]
US 20180050800A1 · Boykin · 2018 [cited by examiner]
US 20180115788A1 · Burns et al. · 2018 [cited by applicant]
US 20180310698A1 · Rao · 2018 [cited by examiner]
US 20190158788A1 · Mcbride · 2019 [cited by examiner]
US 20190208168A1 · Collings, III · 2019 [cited by examiner]
US 20210034882A1 · Johnson · 2021 [cited by examiner]
US 20220011533A1 · Matikainen · 2022 [cited by examiner]
US 20220033077A1 · Myslinski · 2022 [cited by examiner]
US 20220172700A1 · Xiong · 2022 [cited by examiner]
US 20220263994A1 · Kawamoto · 2022 [cited by examiner]
US 20220360705A1 · Ogino · 2022 [cited by examiner]
US 20230056155A1 · Fukuda · 2023 [cited by examiner]
US 20230072905A1 · Chen · 2023 [cited by examiner]
US 20230093631A1 · Kim · 2023 [cited by examiner]
US 20230125724A1 · Kristiansen · 2023 [cited by examiner]
US 20230126960A1 · Ogino · 2023 [cited by examiner]
US 20230396741A1 · Bendtson · 2023 [cited by examiner]
US 20240054787A1 · Son · 2024 [cited by examiner]
WO 2016007988A1 · 2016 [cited by applicant]
WO 2021167374A1 · 2021 [cited by applicant]