IP Library › Granted Patent US 10,997,421
Granted Patent B2
US 10,997,421 · App. 15/947,032 · Granted May 4, 2021

Neuromorphic system for real-time visual activity recognition

Inventors: Deepak Khosla (Camarillo, CA); Ryan M. Uhlenbrock (Camarillo, CA); Yang Chen (Westlake Village, CA)
Assignee: HRL Laboratories, LLC
G06K9/00718G06K9/00744G06K9/00771G06K9/4628G06K9/6271G06N3/04G06N3/0445G06N3/0454G06N3/08G05D1/0202G05D1/0246G06K2009/00738
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 10,997,421
App. No.
15/947,032
Granted
May 4, 2021
Kind
B2
Abstract

Described is a system for visual activity recognition that includes one or more processors and a memory, the memory being a non-transitory computer-readable medium having executable instructions encoded thereon, such that upon execution of the instructions, the one or more processors perform operations including detecting a set of objects of interest in video data and determining an object classification for each object in the set of objects of interest, the set including at least one object of interest. The one or more processors further perform operations including forming a corresponding activity track for each object in the set of objects of interest by tracking each object across frames. The one or more processors further perform operations including, for each object of interest and using a feature extractor, determining a corresponding feature in the video data. The system may provide a report to a user's cell phone or central monitoring facility.

Claims (47)

1. A system for visual activity recognition, the system comprising:

one or more processors and a memory, the memory being a non-transitory computer-readable medium having executable instructions encoded thereon, such that upon execution of the instructions, the one or more processors perform operations of:

detecting a set of objects of interest in video data and determining an object classification for each object in the set of objects of interest, the set comprising at least one object of interest;

forming a corresponding activity track for each object in the set of objects of interest by tracking each object across a plurality of frames, wherein a filter is used to predict a centroid of each activity track in a current frame and update a moving bounding box such that each activity track is a sequence of moving bounding boxes across the plurality of frames, with the centroid representing movement of the object of interest across the plurality of frames in the video data and a size of each moving bounding box representing a size of the object of interest;

for each object of interest and using a feature extractor comprising a convolutional neural network, determining a corresponding feature in the video data by performing feature extraction based on the corresponding activity track by independently processing each activity track with the convolutional neural network to determine the corresponding feature; and

for each object of interest, based on the output of the feature extractor, determining a corresponding activity classification for each object of interest.

2. The system of claim 1 , wherein the one or more processors further perform operations of:

controlling a device based on at least one of the corresponding activity classifications.

3. The system of claim 2 , wherein controlling the device comprises using a machine to send at least one of a visual, audio, or electronic alert regarding the activity classification.

4. The system of claim 2 , wherein controlling the device comprises causing a ground-based or aerial vehicle to initiate a physical action.

5. The system of claim 1 , wherein the feature extractor further comprises a recurrent neural network, and the one or more processors further perform operations of:

for each object of interest and using the recurrent neural network, extracting a corresponding temporal sequence feature based on at least one of the corresponding activity track and the corresponding feature.

6. The system of claim 5 , wherein the recurrent neural network uses Long Short-Term Memory as a temporal component.

7. The system of claim 1 , wherein the convolutional neural network comprises at least five layers of convolution-rectification-pooling.

8. The system of claim 1 , wherein the convolutional neural network further comprises at least two fully-connected layers.

9. The system of claim 1 , wherein the activity classification comprises at least one of a probability and a confidence score.

10. The system of claim 5 , wherein the set of objects of interest includes multiple objects of interest, and the convolutional neural network, the recurrent neural network, and the activity classifier operate in parallel on multiple corresponding activity tracks.

11. The system of claim 1 , wherein the activity classification comprises at least one of a probability and a confidence score.

12. The system of claim 1 , wherein the one or more processors further perform operations of:

reporting the corresponding activity classification for each object of interest to a user's cell phone or to a central monitoring facility.

13. A computer program product for biofeedback, the computer program product comprising:

a non-transitory computer-readable medium having executable instructions encoded thereon, such that upon execution of the instructions by one or more processors, the one or more processors perform operations of:

detecting a set of objects of interest in video data and determining an object classification for each object in the set of objects of interest, the set comprising at least one object of interest;

forming a corresponding activity track for each object in the set of objects of interest by tracking each object across a plurality of frames, wherein a filter is used to predict a centroid of each activity track in a current frame and update a moving bounding box such that each activity track is a sequence of moving bounding boxes across the plurality of frames, with the centroid representing movement of the object of interest across the plurality of frames in the video data and a size of each moving bounding box representing a size of the object of interest;

for each object of interest and using a feature extractor comprising a convolutional neural network, determining a corresponding feature in the video data by performing feature extraction based on the corresponding activity track by independently processing each activity track with the convolutional neural network to determine the corresponding feature; and

for each object of interest, based on the output of the feature extractor, determining a corresponding activity classification for each object of interest.

14. The computer program product of claim 13 , wherein the one or more processors further perform operations of:

controlling a device based on at least one of the corresponding activity classifications.

15. The computer program product of claim 14 , wherein controlling the device comprises using a machine to send at least one of a visual, audio, or electronic alert regarding the activity classification.

16. The computer program product of claim 14 , wherein controlling the device comprises causing a ground-based or aerial vehicle to initiate a physical action.

17. The computer program product of claim 13 , wherein the feature extractor further comprises a recurrent neural network, and the one or more processors further perform operations of:

for each object of interest and using the recurrent neural network, extracting a corresponding temporal sequence feature based on at least one of the corresponding activity track and the corresponding feature.

18. The computer program product of claim 16 , wherein the recurrent neural network uses Long Short-Term Memory as a temporal component.

19. A computer implemented method for biofeedback, the method comprising an act of:

causing one or more processers to execute instructions encoded on a non-transitory computer-readable medium, such that upon execution, the one or more processors perform operations of:

detecting a set of objects of interest in video data and determining an object classification for each object in the set of objects of interest, the set comprising at least one object of interest;

forming a corresponding activity track for each object in the set of objects of interest by tracking each object across a plurality of frames, wherein a filter is used to predict a centroid of each activity track in a current frame and update a moving bounding box such that each activity track is a sequence of moving bounding boxes across the plurality of frames, with the centroid representing movement of the object of interest across the plurality of frames in the video data and a size of each moving bounding box representing a size of the object of interest;

for each object of interest and using a feature extractor comprising a convolutional neural network, determining a corresponding feature in the video data by performing feature extraction based on the corresponding activity track by independently processing each activity track with the convolutional neural network to determine the corresponding feature; and

for each object of interest, based on the output of the feature extractor, determining a corresponding activity classification for each object of interest.

20. The method of claim 19 , wherein the one or more processors further perform operations of:

controlling a device based on at least one of the corresponding activity classifications.

21. The method of claim 20 , wherein controlling the device comprises using a machine to send at least one of a visual, audio, or electronic alert regarding the activity classification.

22. The method of claim 20 , wherein controlling the device comprises causing a ground-based or aerial vehicle to initiate a physical action.

23. The method of claim 19 , wherein the feature extractor further comprises a recurrent neural network, and the one or more processors further perform operations of:

for each object of interest and using the recurrent neural network, extracting a corresponding temporal sequence feature based on at least one of the corresponding activity track and the corresponding feature.

24. The method of claim 23 , wherein the recurrent neural network uses Long Short-Term Memory as a temporal component.

25. The system as set forth in claim 5 , wherein the recurrent neural network concatenates features from multiple frames.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jul 9, 2018
From: KHOSLA, DEEPAK; UHLENBROCK, RYAN M.; CHEN, YANG
To: HRL LABORATORIES, LLC
Reel/Frame 046298/0870 →
Continuity (4)
Continuation In Part 15883822 · Jan 30, 2018
Provisional Application 62479204 · Mar 30, 2017
Provisional Application 62516217 · Jun 7, 2017
Related Publication 20180300553A1 · Oct 18, 2018
Cited By (3)
US 12,249,128 US 12,333,806 US 12,555,378