IP Library Granted Patent US 12,456,301
Granted Patent B2
US 12,456,301 · App. 17/929,995 · Granted Oct 28, 2025

Method of training a machine learning algorithm to identify objects or activities in video surveillance data

Inventor: Kamal Nasrollahi (Brøndby, DK)
Assignee: MILESTONE SYSTEMS A/S
G06V20/52G06N20/00G06T7/194G06T7/70G06T17/00G06V10/7747G06T2207/10028G06T2207/20081G06V2201/07
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,456,301
App. No.
17/929,995
Granted
Oct 28, 2025
Kind
B2
Abstract

A method of training a machine learning algorithm to identify objects or activities in video surveillance data comprises generating a 3D simulation of a real environment from video surveillance data captured by at least one video surveillance camera installed in the real environment. Objects or activities are synthesized within the simulated 3D environment and the synthesized objects or activities within the simulated 3D environment are used as training data to train the machine learning algorithm to identify objects or activities, wherein the synthesized objects or activities within the simulated 3D environment used as training data are all viewed from the same viewpoint in the simulated 3D environment.

Claims (27)

1. A video surveillance method of training a machine learning algorithm to identify activities in video surveillance data by a method comprising the steps of:

generating a 3D simulation of a real environment from video surveillance data captured by at least one video surveillance camera installed in the real environment;

synthesizing activities within the simulated 3D environment; and

using the synthesized activities within the simulated 3D environment as training data to train the machine learning algorithm to identify objects or activities, wherein the synthesized activities within the simulated 3D environment used as training data are all viewed from the same viewpoint in the simulated 3D environment;

installing a video surveillance camera at the same viewpoint in the real environment as the viewpoint from which the synthesized activities within the simulated 3D environment used as training data are viewed; and

applying the trained machine learning algorithm to video surveillance data captured by the video surveillance camera.

2. The video surveillance method according to claim 1 , wherein the machine learning algorithm is pre-trained using training data comprising image data of activities in different environments, before being trained using the synthesized activities within the simulated 3D environment.

3. The video surveillance method according to claim 1 , wherein the simulated 3D environment has a fixed configuration and the method comprises varying imaging conditions and/or weather conditions within the simulated 3D environment and synthesizing the activities under the different imaging and/or weather conditions.

4. The video surveillance method according to claim 1 , wherein the video surveillance data used to generate the simulated 3D environment is captured from multiple viewpoints in the real environment.

5. The video surveillance method according to claim 1 , wherein the video surveillance data used to generate the simulated 3D environment is captured only from the same viewpoint in the real environment as the viewpoint from which the synthesized activities are viewed in the simulated 3D environment to train the machine learning algorithm.

6. The video surveillance method according to claim 1 , wherein the machine learning algorithm runs on a processor in a video surveillance camera and the training of the machine learning algorithm is carried out in the camera.

7. The video surveillance method according to claim 6 , wherein the steps of generating the 3D simulation and synthesising the activities to generate the training data is carried out in a server of a video management system and the training data is sent to the video surveillance camera.

8. The video surveillance method according to claim 1 , wherein the step of generating a 3D simulation of the real environment comprises:

acquiring image data of the real environment from the video surveillance camera;

generating a depth map from the image data; and

using a semantic segmentation algorithm to label background information and foreground information.

9. A non-transitory computer-readable storage medium storing a computer program comprising code which, when run on a processor causes it to carry out the method according to claim 1 .

10. A video surveillance system, comprising:

a processor configured to:

generate a 3D simulation of a real environment from video surveillance data captured by at least one video surveillance camera installed in the real environment;

synthesize activities within the simulated 3D environment;

generate training data comprising image data comprising the synthesized activities within the simulated 3D environment viewed from a single viewpoint in the simulated 3D environment; and

use the training data to train a machine learning algorithm to identify activities in video surveillance data; and

a video surveillance camera installed at the same viewpoint in the real environment as the viewpoint from which the synthesized activities within the simulated 3D environment used as training data are viewed,

wherein the processing means is configured to apply the trained machine learning algorithm to video surveillance data captured by the video surveillance camera installed at the same viewpoint in the real environment.

11. The video surveillance system according to claim 10 , wherein the processing means includes a first processor in the video surveillance camera that captures the video surveillance data used to generate the 3D simulation, wherein the machine learning algorithm runs on the first processor and the first processor is configured to use the training data to train the machine learning algorithm to identify objects or activities.

12. The video surveillance system according to claim 11 , wherein the processing means includes a second processor in a video management system, wherein the second processor is configured to receive the video surveillance data from the video surveillance camera, generate the 3D simulation, synthesize the activities within the simulated 3D environment and generate the training data and send the training data to the first processor.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Sep 6, 2022
From: NASROLLAHI, KAMAL
To: MILESTONE SYSTEMS A/S
Reel/Frame 061001/0313 →
Priority Claims (1)
EP 21196002 · Sep 10, 2021 · regional
Continuity (1)
Related Publication 20230081908A1 · Mar 16, 2023
References Cited (31)
US 10497257B2 · Sohn · 2019 [cited by examiner]
US 11366981B1 · Paz-Perez · 2022 [cited by examiner]
US 12014446B2 · Koh · 2024 [cited by examiner]
US 20190065853A1 · Sohn · 2019 [cited by examiner]
US 20200175759A1 · Russell et al. · 2020 [cited by applicant]
US 20200294194A1 · Sun · 2020 [cited by applicant]
US 20220157049A1 · Inoshita · 2022 [cited by examiner]
US 20220157136A1 · Metzler · 2022 [cited by examiner]
US 20220237908A1 · Yang · 2022 [cited by examiner]
US 20220406066A1 · Rangarajan · 2022 [cited by examiner]
US 20230072293A1 · Koh · 2023 [cited by examiner]
US 20230177811A1 · Nadler · 2023 [cited by examiner]
WO 2019113510A1 · 2019 [cited by applicant]
WO 2020228766A1 · 2020 [cited by applicant]
Olaf Ronneberger, et al., U-Net: Convolutional Networks for Biomedical Image Segmentation, arXiv: 1505.04597v1 [cs.CV] May 18, 2015. [cited by applicant]
Andrey Kuryenkov, et al., DeformNet: Free-Form Deformation Network for 3D Shape Reconstruction from a Single Image, arXiv:1708.04672v1 [cs.CV] Aug. 11, 2017. [cited by applicant]
Jhony K. Pontes, et al., Image2Mesh: A Learning Framework for Single Image 3D Reconstruction, arXiv: 1711.10669v1 [cs.CV] Nov. 29, 2017. [cited by applicant]
Mennatullah Siam, et al., RTSEG: Real-Time Semantic Segmentation Comparative Study, arXiv: 1803.02758v5 [cs.CV] May 16, 2020. [cited by applicant]
Chaoqiang Zhao, et al., Monocular Depth Estimation Based On Deep Learning: An Overview, arXiv:2003.06620v2 [cs.CV] Jul. 3, 2020. [cited by applicant]
Jae-Han Lee, et al., Single-Image Depth Estimation Based on Fourier Domain Analysis, 2018 IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2018. [cited by applicant]
Johannes Kopf, et al., One Shot 3D Photography, ACM Trans. Graph., vol. 39, No. 4, Article 76. Jul. 2020. [cited by applicant]
Changqian Yu, et al., BiSeNet: Bilateral Segmentation Network for Real-time Semantic Segmentation, ECCV 2018: Computer Vision—ECCV 2018, Computer Vision Foundation. [cited by applicant]
Di Lin, et al., Multi-Scale Context Intertwining for Semantic Segmentation, ECCV 2018: Computer Vision—ECCV 2018, Computer Vision Foundation. [cited by applicant]
Haoqiang Fan, et al., A Point Set Generation Network for 3D Object Reconstruction from a Single Image, 2017 IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 2017. [cited by applicant]
Zilong Huang, et al., CCNet: Criss-Cross Attention for Semantic Segmentation, 2019 IEEE/CVF International Conference on Computer Vision (ICCV), 2019. [cited by applicant]
Andrey Ignatov, et al., Fast and Accurate Single-Image Depth Estimation on Mobile Devices, Mobile AI 2021 Challenge: Report, 2021 IEEE/CVF Conference on Computer Vision and Pattern Recognition Workshops (CVPRW), 2021. [cited by applicant]
Xin Wu, et al., Improvement of Mask-RCNN Object Segmentation Algorithm, Intelligent Robotics and Applications. ICIRA 2019. Lecture Notes in Computer Science, vol. 11740, 2019. [cited by applicant]
Fabio Tosi, et al., Learning monocular depth estimation infusing traditional stereo knowledge, 2019 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2019. [cited by applicant]
Chaoyang Wang, et al., Learning Depth from Monocular Videos using Direct Methods, The IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 2018. [cited by applicant]
Haozhe Xie, et al., Pix2Vox: Context-aware 3D Reconstruction from Single and Multi-view Images, 2019 IEEE/CVF International Conference on Computer Vision (ICCV), 2019. [cited by applicant]
Shuran Song, et al., Sun RGB-D: A RGB-D Scene Understanding Benchmark Suite, CVPR, 2015, https://openaccess.thecvf.com/content_cvpr_2015/papers/Song_SUN_RGB-D_A_2015_CVPR_paper.pdf. [cited by applicant]