IP Library Granted Patent US 10,691,949
Granted Patent B2
US 10,691,949 · App. 15/812,685 · Granted Jun 23, 2020

Action recognition in a video sequence

Inventors: Niclas Danielsson (Lund, SE); Simon Molin (Lund, SE)
Assignee: Axis AB
G06K9/00718G06K9/00335G06K9/00744G06K9/00979G06K9/3233G06K9/685G06K2009/00738
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 10,691,949
App. No.
15/812,685
Granted
Jun 23, 2020
Kind
B2
Abstract

A method and system for action recognition in a video sequence is disclosed. The system comprises a camera configured to capture the video sequence and a server configured to perform action recognition. The camera comprises an object identifier that identifies an object of interest in an object image frame of the video sequence; an action candidate recognizer configured to apply a first action recognition algorithm to the object image frame to detect presence of an action candidate; an video extractor configured to produce action image frames of an action video sequence by extracting video data pertaining to a plurality of image frames from the video sequence; and a network interface configured to transfer the action video sequence to the server. The server comprises an action verifier configured to apply a second action recognition algorithm to the action video sequence to verify or reject that the action candidate is an action.

Claims (27)

1. A method for action recognition in a video sequence captured by a camera, the method comprising:

by circuitry of the camera:

identifying an object of interest in an image frame of the video sequence;

applying a first action recognition algorithm to the image frame to detect an action candidate, wherein the image frame is a single image comprising the object of interest, wherein the first action recognition algorithm uses contextual and/or spatial recognition information of the single image frame to detect the action candidate within the image frame;

producing image frames of an action video sequence by extracting video data pertaining to a plurality of image frames from the video sequence, wherein one or more of the plurality of image frames from which the video data is extracted comprises the object of interest; and

transferring the action video sequence to a server configured to perform action recognition; and

by circuitry of the server:

applying a second action recognition algorithm to the action video sequence to verify or reject that the action candidate is an action of a predefined type, wherein the second action recognition algorithm uses temporal information of a plurality of image frames of the action video sequence.

2. The method according to claim 1 , wherein the act of producing the image frames of the action video sequence comprises cropping the plurality of image frames of the video sequence such that the image frames comprising the object of interest comprises at least a portion of the object of interest.

3. The method according to claim 2 , wherein the image frames of the action video sequence comprising the object of interest comprises a portion of background at least partly surrounding the object of interest.

4. The method according to claim 1 , wherein the act of transferring the action video sequence comprises transferring coordinates within the action video sequence to the object of interest.

5. The method according to claim 1 , wherein the method further comprises, by the circuitry of the camera:

detecting an object of interest in the video sequence, wherein the act of producing the image frames of the action video sequence comprises extracting video data pertaining to a first predetermined number of image frames of the video sequence related to a point of time before detection of the object of interest.

6. The method according to claim 1 , wherein the method further comprises, by the circuitry of the camera:

detecting an object of interest in the video sequence, wherein the act of producing the image frames of the action video sequence comprises extracting video data pertaining to a second predetermined number of image frames of the video sequence related to a point of time after detection of the object of interest.

7. The method according to claim 1 , wherein the camera and the server are separate physical entities positioned at a distance from each other and are configured to communicate with each other via a digital network.

8. A system for action recognition in a video sequence, the system comprising:

a camera configured to capture the video sequence and a server configured to perform action recognition, the camera comprising:

an object identifier configured to identify an object of interest in an image frame of the video sequence;

an action candidate recognizer configured to apply a first action recognition algorithm to the image frame to detect an action candidate, wherein the image frame is a single image comprising the object of interest, wherein the first action recognition algorithm uses contextual and/or spatial recognition information of the single image frame to detect the action candidate within the image frame;

a video extractor configured to produce image frames of an action video sequence by extracting video data pertaining to a plurality of image frames from the video sequence, wherein one or more of the plurality of image frames from which the video data is extracted comprises the object of interest; and

a network interface configured to transfer the action video sequence to the server, the server comprising:

an action verifier configured to apply a second action recognition algorithm to the action video sequence to verify or reject that the action candidate is an action of a predefined type, wherein the second action recognition algorithm uses temporal information of a plurality of image frames of the action video sequence.

9. The system according to claim 8 , wherein the video extractor is further configured to crop the plurality of images frames of the video sequence such that the image frames of the video sequence comprising the object of interest comprises at least a portion of the object of interest.

10. The system according to claim 8 , wherein the video extractor is further configured to crop the plurality of images frames of the video sequence such that the image frames of the video sequence comprising the object of interest comprises a portion of background at least partly surrounding the object of interest.

11. The system according to claim 8 , wherein the object identifier is further configured to detect an object of interest in the video sequence, wherein the video extractor is further configured to extract video data pertaining to a first predetermined number of image frames of the video sequence related to a point of time before detection of the object of interest.

12. The system according to claim 8 , wherein object identifier is further configured to detect an object of interest in the video sequence, wherein the video extractor is further configured to extract video data pertaining to a second predetermined number of image frames of the video sequence related to a point of time after detection of the object of interest.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Nov 14, 2017
From: DANIELSSON, NICLAS; MOLIN, SIMON
To: AXIS AB
Reel/Frame 044124/0386 →
Priority Claims (1)
EP 16198678 · Nov 14, 2016 · regional
Continuity (1)
Related Publication 20180137362A1 · May 17, 2018
Cited By (2)
US 12,412,426 US 12,423,975