IP Library Granted Patent US 9,959,630
Granted Patent B2
US 9,959,630 · App. 15/019,759 · Granted May 1, 2018

Background model for complex and dynamic scenes

View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 9,959,630
App. No.
15/019,759
Granted
May 1, 2018
Kind
B2
Abstract

Systems and methods for viewing a scene depicted in a sequence of video frames and identifying and tracking objects between separate frames of the sequence. Each tracked object is classified based on known categories and a stream of context events associated with the object is generated. A sequence of primitive events based on the stream of context events is generated and stored together, along with detailed data and generalized data related to an event. All of the data is then evaluated to learn patterns of behavior that occur within the scene.

Claims (66)

1. A computer-implemented method, comprising:

receiving a sequence of video frames from a video camera;

receiving a request to view a scene depicted in the sequence of video frames;

identifying and tracking at least one object between separate frames of the sequence of video frames;

classifying each tracked object based on a known category of objects;

generating a stream of context events associated with each tracked object;

generating a sequence of primitive events based on the stream of context events;

storing the stream of context events and the sequence of primitive events in one or more adaptive resonance theory (ART) networks;

storing detailed data in the one or more ART networks related to an event based on the stream of context events and the sequence of primitive events;

storing generalized data in the one or more ART networks related to an event based on the stream of context events and the sequence of primitive events; and

evaluating the stream of context events, the sequence of primitive events, the detailed data, and the generalized data with the one or more ART networks to learn patterns of behavior that occur within the scene.

2. The computer-implemented method of claim 1 , wherein the stream of context events includes a collection of kinematic information related to the at least one object.

3. The computer-implemented method of claim 2 , wherein the kinematic information includes one or more of: position, current trajectory, projected trajectory, direction, orientation, velocity, acceleration, size and color.

4. The computer-implemented method of claim 1 , wherein classifying further includes identifying features of the object.

5. The computer-implemented method of claim 4 , wherein the identifying features include one or more: of height/width in pixels, average color values, shape and area.

6. The computer-implemented method of claim 1 , wherein the object is a person and wherein the identifying features include one or more of: prediction of a gender, an estimation of a pose, and an indication of whether the person is carrying an object.

7. A system, comprising:

a processor; and

one or more adaptive resonance theory (ART) networks in communication with the processor when the system is in operation, the one or more ART networks having stored thereon instructions that upon execution by the processor at least cause the system to:

receive a sequence of video frames from a video camera;

receive a request to view a scene depicted in the sequence of video frames;

identify and track at least one object between separate frames of the sequence of video frames;

classify each tracked object based on a known category of objects;

generate a stream of context events associated with each tracked object;

generate a sequence of primitive events based on the stream of context events;

store the stream of context events and the sequence of primitive events in the one or more ART networks;

store detailed data in the one or more ART networks related to an event based on the stream of context events and the sequence of primitive events;

store generalized data in the one or more ART networks related to an event based on the stream of context events and the sequence of primitive events; and

evaluate the stream of context events, the sequence of primitive events, the detailed data, and the generalized data with the one or more ART networks to learn patterns of behavior that occur within the scene.

8. The system of claim 7 , wherein the stream of context events includes a collection of kinematic information related to the at least one object.

9. The system of claim 8 , wherein the kinematic information includes one or more of: position, current trajectory, projected trajectory, direction, orientation, velocity, acceleration, size and color.

10. The system of claim 7 , wherein classifying further includes identifying features of the object.

11. The system of claim 10 , wherein the identifying features include one or more of: height/width in pixels, average color values, shape and area.

12. The system of claim 7 , wherein the object is a person and wherein the identifying features include one or more of: prediction of a gender, an estimation of a pose, and an indication of whether the person is carrying an object.

13. A computer-implemented method, comprising:

receiving a sequence of video frames from a video camera;

receiving a request to view a scene depicted in the sequence of video frames;

retrieving a background image and one or more foreground images associated with the scene;

identifying and tracking at least some of the one or more foreground images between separate frames of the sequence of video frames;

classifying each tracked foreground image based on a known category of objects;

generating a stream of context events associated with each tracked foreground image;

generating a sequence of primitive events based on the stream of context events;

storing the stream of context events and the sequence of primitive events in an adaptive resonance theory (ART) network;

storing detailed data in the adaptive resonance theory (ART) network related to an event based on the stream of context events and the sequence of primitive events;

storing generalized data in the adaptive resonance theory (ART) network related to an event based on the stream of context events and the sequence of primitive events; and

evaluating the stream of context events, the sequence of primitive events, the detailed data, and the generalized data with the one or more ART networks to learn patterns of behavior that occur within the scene.

14. The computer-implemented method of claim 13 , wherein the stream of context events includes a collection of kinematic information related to the at least one object.

15. The computer-implemented method of claim 14 , wherein the kinematic information includes one or more of: position, current trajectory, projected trajectory, direction, orientation, velocity, acceleration, size and color.

16. The computer-implemented method of claim 13 , wherein classifying further includes identifying features of the object.

17. The computer-implemented method of claim 16 , wherein the identifying features include one or more of: height/width in pixels, average color values, shape and area.

18. The computer-implemented method of claim 13 , wherein the object is a person and wherein the identifying features include one or more of: prediction of a gender, an estimation of a pose, and an indication of whether the person is carrying an object.

19. A system, comprising:

a processor; and

one or more adaptive resonance theory (ART) networks in communication with the processor when the system is in operation, the one or more ART networks having stored thereon instructions that upon execution by the processor at least cause the system to:

receive a sequence of video frames from a video camera;

receive a request to view a scene depicted in the sequence of video frames;

retrieve background image and one or more foreground images associated with the scene;

identify and tracking at least some of the one or more foreground images between separate frames of the sequence of video frames;

classify each tracked foreground image based on a known category of objects;

generate a stream of context events associated with each tracked foreground image;

generate a sequence of primitive events based on the stream of context events;

store the stream of context events and the sequence of primitive events in the one or more ART networks;

store detailed data in the one or more ART networks related to an event based on the stream of context events and the sequence of primitive events;

store generalized data in the one or more ART networks related to an event based on the stream of context events and the sequence of primitive events; and

evaluate the stream of context events, the sequence of primitive events, the detailed data, and the generalized data with the one or more ART networks to learn patterns of behavior that occur within the scene.

20. The system of claim 19 , wherein the stream of context events includes a collection of kinematic information related to the at least one object.

Assignments (4)
NUNC PRO TUNC ASSIGNMENT Recorded Oct 13, 2022
From: AVIGILON PATENT HOLDING 1 CORPORATION
To: MOTOROLA SOLUTIONS, INC.
Reel/Frame 062034/0176 →
CHANGE OF NAME Recorded Dec 12, 2016
From: 9051147 CANADA INC.
To: AVIGILON PATENT HOLDING 1 CORPORATION
Reel/Frame 040886/0579 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Mar 8, 2016
From: BEHAVIORAL RECOGNITION SYSTEMS, INC.
To: 9051147 CANADA INC.
Reel/Frame 037925/0645 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Feb 15, 2016
From: COBB, WESLEY KENNETH; SEOW, MING-JUNG; YANG, TAO
To: BEHAVIORAL RECOGNITION SYSTEMS, INC.
Reel/Frame 037736/0703 →