IP Library Granted Patent US 12,008,794
Granted Patent B2
US 12,008,794 · App. 17/239,536 · Granted Jun 11, 2024

Systems and methods for intelligent video surveillance

Inventor: Zhong Zhang (Great Falls, VA)
Assignee: SHANGHAI TRUTHVISION INFORMATION TECHNOLOGY CO., LTD.
G06V10/255G06F18/214G06N3/08G06T7/246G06T7/292G06T7/80G06V20/10G06V20/52
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,008,794
App. No.
17/239,536
Granted
Jun 11, 2024
Kind
B2
Abstract

A method may include obtaining a video collected by a visual sensor, the video including a plurality of frames and detecting one or more objects from the video in at least a portion of the plurality of frames. The method may also include determining a first detection result associated with the one or more objects with a trained self-learning model. The method may further include selecting a target moving object of interest from the one or more objects at least in part based on the first detection result. The trained self-learning model may be provided based on a plurality of training samples collected by the visual sensor.

Claims (60)

1. A system, comprising:

a storage device storing a set of instructions; and

at least one processor configured to communicate with the storage device, wherein when executing the set of instructions, the at least one processor is directed to cause the system to perform operations including:

obtaining a video collected by a visual sensor, the video including a plurality of frames;

detecting, in at least a portion of the plurality of frames, one or more objects from the video;

determining, with a trained self-learning model, a first detection result associated with the one or more objects;

determining, based on the at least a portion of the plurality of frames, one or more behavior features associated with each of the one or more objects;

determining, based on the one or more behavior features associated with each of the one or more objects, a second detection result associated with each of the one or more objects; and

determining, based on the first detection result and the second detection result, a tartlet moving object of interest from the one or more objects, wherein the trained self-learning model is provided based on a plurality of training samples collected by the visual sensor.

2. The system of claim 1 , wherein to detect, in at least a portion of the plurality of frames, one or more objects from the video, the at least one processor is directed to cause the system to perform the operations including:

detecting the one or more objects from the video using an object detection model.

3. The system of claim 2 , wherein the object detection model is constructed based on a deep learning model.

4. The system of claim 1 , wherein to determine, based on the at least a portion of the plurality of frames, one or more behavior features associated with each of the one or more objects, the at least one processor is directed to cause the system to perform the operations including:

determining, based on the at least a portion of the plurality of frames and a prior calibration model of the visual sensor determined in a last calibration, the one or more behavior features.

5. The system of claim 1 , wherein the one or more behavior features of the one or more objects include at least one of a speed, an acceleration, a trajectory, a movement amplitude, a direction, a movement frequency, or voice information.

6. The system of claim 1 , wherein the trained self-learning model is generated by a process including:

obtaining the plurality of training samples each of which includes a historical video collected by the visual sensor;

detecting one or more motion subjects from the historical video for each of the plurality of training samples; and

training a self-learning model using information associated with the detected one or more motion subjects to obtain the trained self-learning model.

7. The system of claim 6 , wherein the information associated with the detected one or more motion subjects includes at least one of:

time information when the detected one or more motion subjects recorded by the historical video;

spatial information associated with the detected one or more motion subjects;

weather information when the detected one or more motion subjects recorded by the historical video; or

motion information of the detected one or more motion subjects.

8. The system of claim 1 , wherein the trained self-learning model includes a first part relating to reference knowledge of different scenes and a second part relating to learned knowledge generated from a training process of the trained self-learning model.

9. The system of claim 8 , wherein the reference knowledge of different scenes includes characteristics of one or more subjects appearing in each of the different scenes.

10. The system of claim 1 , wherein the first detection result includes one or more first candidate moving objects of interest and the second detection result includes one or more second candidate moving objects of interest, and to determine, based on the first detection result and the second detection result, the target moving object of interest from the one or more objects, the at least one processor is directed to cause the system to perform the operations including:

designating a same candidate moving object of interest from the one or more first candidate moving objects of interest and the one or more second candidate moving objects of interest as the target moving object of interest.

11. The system of claim 1 , wherein the first detection result includes a first probability that each of the one or more objects is a moving object of interest, the second detection result includes a second probability that each of the one or more objects is a moving object of interest, and to determine, based on the first detection result and the second detection result, the target moving object of interest from the one or more objects, the at least one processor is directed to cause the system to perform the operations including:

designating a moving object having a first probability exceeding a first threshold and a second probability exceeding a second threshold as the target moving object of interest.

12. The system of claim 1 , wherein the at least one processor is directed to cause the system to perform additional operations including:

in response to a detection of the target moving object of interest from the video, generating feedback relating to the detection of the target moving object of interest; and

transmitting the feedback relating to the detection of the target moving object of interest to a terminal.

13. The system of claim 12 , wherein the feedback includes a notification indicating that a moving object exists.

14. The system of claim 1 , wherein the at least one processor is directed to cause the system to perform additional operations including:

in response to a detection of each of at least a portion of the one or more objects from the video, generating candidate feedbacks each of which relates to the detection of one of the one or more objects;

determining, based on at least one of the first detection result or the second detection result, target feedback from the candidate feedbacks; and

transmitting the target feedback to a terminal.

15. The system of claim 1 , wherein the at least one processor is further configured to cause the system to perform additional operations including:

determining, based on the target moving object of interest, a calibration model of the visual sensor, the calibration model describing a transform relationship between a two-dimensional (2D) coordinate system and a three-dimensional (3D) coordinate system of the visual sensor.

16. The system of claim 15 , wherein to determine, based on the target moving object of interest, a calibration model of the visual sensor, the at least one processor is further configured to cause the system to perform the additional operations including:

determining, based on the at least a portion of the plurality of frames, an estimated value of a characteristic of the target moving object of interest denoted by the 2D coordinate system; and

determining, based on the estimated value and a reference value of the characteristic of the target moving object of interest denoted by the 3D coordinate system, the calibration model.

17. The system of claim 16 , wherein the characteristic of the target moving object of interest includes a physical size of at least a portion of the target moving object of interest.

18. The system of claim 1 , wherein the target moving object of interest includes at least one of a person, a vehicle, or an animal whose motion includes an anomaly.

19. A method implemented on a computing device having at least one processor and at least one computer-readable storage medium for abnormal scene detection, the method comprising:

obtaining a video collected by a visual sensor, the video including a plurality of frames;

detecting, in at least a portion of the plurality of frames, one or more objects from the video;

determining, with a trained self-learning model, a first detection result associated with the one or more objects;

determining, based on the at least a portion of the plurality of frames, one or more behavior features associated with each of the one or more objects;

determining, based on the one or more behavior features associated with each of the one or more objects, a second detection result associated with each of the one or more objects; and

determining, based on the first detection result and the second detection result, a target moving object of interest from the one or more objects, wherein the trained self-learning model is provided based on a plurality of training samples collected by the visual sensor.

20. A non-transitory computer readable medium, comprising:

instructions being executed by at least one processor, causing the at least one processor to implement a method, comprising:

obtaining a video collected by a visual sensor, the video including a plurality of frames;

detecting, in at least a portion of the plurality of frames, one or more objects from the video;

determining, with a trained self-learning model, a first detection result associated with the one or more objects;

determining, based on the at least a portion of the plurality of frames, one or more behavior features associated with each of the one or more objects;

determining, based on the one or more behavior features associated with each of the one or more objects, a second detection result associated with each of the one or more objects; and

determining, based on the first detection result and the second detection result, a target moving object of interest from the one or more objects, wherein the trained self-learning model is provided based on a plurality of training samples collected by the visual sensor.

Assignments (2)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Apr 25, 2021
From: ZHANG, ZHONG
To: TRUTHVISION, INC.
Reel/Frame 056029/0951 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Apr 25, 2021
From: TRUTHVISION, INC.
To: SHANGHAI TRUTHVISION INFORMATION TECHNOLOGY CO., LTD.
Reel/Frame 056029/0958 →
Continuity (4)
Continuation PCTCN2019113176 · Oct 25, 2019
Provisional Application 62750795 · Oct 25, 2018
Provisional Application 62750797 · Oct 25, 2018
Related Publication 20210241468A1 · Aug 5, 2021