IP Library Granted Patent US 8,774,499
Granted Patent B2
US 8,774,499 · App. 13/405,986 · Granted Jul 8, 2014

Embedded optical flow features

Inventors: Jinjun Wang (San Jose, CA); Jing Xiao (Cupertino, CA)
Assignee: Seiko Epson Corporation
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 8,774,499
App. No.
13/405,986
Granted
Jul 8, 2014
Kind
B2
Abstract

Aspects of the present invention include systems and methods for generating an optical flow-based feature. In embodiments, to extract an optical flow feature, the optical flow at sparse interest points is obtained, and Locality-constrained Linear Coding (LLC) is applied to the sparse interest points to embed each flow into a higher-dimensional code. In embodiments, for an image frame, the multiple codes are combined together using a weighted pooling that is related to the distribution of the optical flows in the image frame. In embodiments, the feature may be used in training models to detect actions, in trained models for action detection, or both.

Claims (55)

1. A computer-implemented method comprising:

obtaining a set of local motion features for an image frame stored in a memory;

for each local motion feature from the set of local motion features, generating a sparse coding vector; and

forming an image feature that represents the image frame by pooling the sparse coding vectors based upon a weighting of a distribution of local motion features in the image frame, including:

generating the image feature for a frame by pooling sparse coding vectors for the image frame by weighting each sparse coding vector in relation to a posterior of a local motion feature corresponding to the sparse coding vector, including weighting each sparse coding vector by an inverse proportional to a square root of a posterior of a local motion feature corresponding to the sparse coding vector.

2. The computer-implemented method of claim 1 wherein weighting each sparse coding vector in relation to a posterior of a local motion feature corresponding to the sparse coding vector comprises:

using an equation:

y

=

C

P

o

-

1

2

(

X

)

,

where y represents the image feature for the image frame, C represents a matrix of spare coding vectors for the image frame, P o (X) represents a matrix of posterior values for local motion features.

3. The computer-implemented method of claim 1 wherein the step of forming an image feature that represents the image frame by pooling the sparse coding vectors based upon a weighting of the distribution of local motion features in the image frame further comprises:

normalizing the pooled sparse coding vectors to form the image feature.

4. The computer-implemented method of claim 1 wherein the step of generating a sparse coding vector comprises:

using Locality-constrained Linear Coding (LLC) and a codebook to generate the spares coding vector.

5. The computer-implemented method of claim 1 wherein the method further comprises:

using the image feature to recognize an action.

6. A non-transitory computer-readable medium or media comprising one or more sequences of instructions which, when executed by one or more processors, causes steps for forming an optical flow feature to represent an image in a video comprising:

for each frame from a set of frames from the video:

extracting optical flows from the frame;

using Locality-constrained Linear Coding (LLC) and an optical flow codebook to convert each optical flows into a higher-dimensional code; and

using a weighted pooling to combine the higher-dimensional codes into the optical flow feature to represent the frame, wherein the weighted pooling is related to a probability distribution of the optical flows, including weighting each higher-dimensional codes by an inverse proportional to a square root of a posterior of an optical flow corresponding to the higher-dimensional code.

7. The computer-readable medium or media of claim 6 wherein the method further comprises:

normalizing the weighted pooled higher-dimensional codes to obtain the optical flow feature for the frame.

8. The computer-readable medium or media of claim 6 further comprising:

extracting the set of frames from the video.

9. The computer-readable medium or media of claim 6 wherein the method further comprises:

building the optical flow codebook using at least some of the optical flows.

10. The computer-readable medium or media of claim 6 wherein the method further comprises:

using at least some of the optical flow features to train a model.

11. The computer-readable medium or media of claim 6 wherein the method further comprises:

using at least some of the optical flow features from the frames to recognize one or more actions.

12. A non-transitory computer-readable medium or media comprising one or more sequences of instructions which, when executed by one or more processors, causes steps to be performed comprising:

obtaining a set of local motion features for an image frame;

for each local motion feature from the set of local motion features, generating a sparse coding vector; and

forming an image feature that represents the image frame by pooling the sparse coding vectors based upon a weighting of a distribution of local motion features in the image frame, including:

generating the image feature for a frame by pooling sparse coding vectors for the image frame by weighting each sparse coding vector in relation to a posterior of a local motion feature corresponding to the sparse coding vector, including weighting each sparse coding vector by an inverse proportional to a square root of a posterior of a local motion feature corresponding to the sparse coding vector.

13. The computer-readable medium or media of claim 12 wherein the step of forming an image feature that represents the image frame by pooling the sparse coding vectors based upon a weighting of the distribution of local motion features in the image frame further comprises:

normalizing the pooled sparse coding vectors to form the image feature.

14. The computer-readable medium or media of claim 12 wherein the step of generating a sparse coding vector comprises:

using Locality-constrained Linear Coding (LLC) and a codebook to generate the spares coding vector.

15. The computer-readable medium or media of claim 12 wherein the steps further comprise:

using the image feature to recognize an action.

Assignments (2)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded May 1, 2012
From: EPSON RESEARCH AND DEVELOPMENT, INC.
To: SEIKO EPSON CORPORATION
Reel/Frame 028136/0868 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Feb 27, 2012
From: WANG, JINJUN; XIAO, JING
To: EPSON RESEARCH AND DEVELOPMENT, INC.
Reel/Frame 027769/0210 →
Continuity (2)
Provisional Application 61447502 · Feb 28, 2011
Related Publication 20120219213A1 · Aug 30, 2012