IP Library Granted Patent US 8,098,730
Granted Patent B2
US 8,098,730 · App. 11/278,487 · Granted Jan 17, 2012

Generating a motion attention model

View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 8,098,730
App. No.
11/278,487
Granted
Jan 17, 2012
Kind
B2
Abstract

Systems and methods to generate a motion attention model of a video data sequence are described. In one aspect, a motion saliency map B is generated to precisely indicate motion attention areas for each frame in the video data sequence. The motion saliency maps are each based on intensity I, spatial coherence Cs, and temporal coherence Ct values. These values are extracted from each block or pixel in motion fields that are extracted from the video data sequence. Brightness values of detected motion attention areas in each frame are accumulated to generate, with respect to time, the motion attention model.

Claims (199)

1. A method implemented at least in part by a computing device for generating a motion attention model of a video data sequence, the method comprising:

generating, by the computing device, a motion saliency map B to precisely indicate motion attention areas for each frame in the video data sequence, the motion saliency map being based on intensity I, spatial coherence C s , and temporal coherence C t values from each location of a block MB ij , in motion fields extracted from the video data sequence; and

accumulating brightness associated with each of the intensity I, the spatial coherence C s , and the temporal coherence C t values of detected motion attention areas to generate, with respect to time, a motion attention model for the video data sequence, wherein:

the intensity I values are motion intensity values each represented on a motion intensity map by a brightness associated with each location of the block MB ij ;

the spatial coherence C s values are each represented on a spatial coherence map by a brightness associated with each location of the block MB ij ;

the temporal coherence C t values are each represented on a temporal coherence map by a brightness associated with each location of the block MB ij ; and

a combination is represented on the motion saliency map B by a brightness of the detected motion attention areas associated with each location of the block MB ij .

2. The method of claim 1 , wherein the video data sequence is in an MPEG data format.

3. The method of claim 1 , wherein the motion field is a motion vector field, or an optical flow field.

4. A computing device for generating a motion attention model of a video data sequence, the computing device comprising:

a motion attention modeling module for:

generating a motion saliency map B to precisely indicate motion attention areas for each frame in the video data sequence, the motion saliency map being based on intensity I, spatial coherence C s , and temporal coherence C t values from each location of a block MB ij in motion fields extracted from the video data sequence, wherein the saliency map B is calculated according to B=I×C t ×(1−I×C s ); and

accumulating brightness represented by the intensity I, the spatial coherence C s and the temporal coherence C t values of detected motion attention areas to generate, with respect to time, a motion attention model for the video data sequence.

5. The method of claim 1 , wherein a linear combination is represented on the motion saliency map B.

6. The method of claim 1 , wherein the intensity I is a normalized motion intensity.

7. The method of claim 1 , wherein the spatial coherence C s is a normalized spatial coherence.

8. The method of claim 1 , wherein the temporal coherence C t is a normalized temporal coherence.

9. The method of claim 1 , wherein a combination represented on the motion saliency map B is a combination of the brightness associated with each location of the block MB ij from the intensity map, the spatial coherence map and the temporal coherence map.

10. The method of claim 1 , wherein I is a normalized magnitude of a motion vector that is calculated according to:

I ( i,j )=√{square root over ( dx i,j 2 +dy i,j 2 )}/MaxMag,

wherein (dx i,j , dy i,j ) denotes two components of the motion vector in motion field, and MaxMag is the maximum magnitude of motion vectors.

11. The method of claim 1 , wherein C s is calculated with respect to spatial widow w as follows:

Cs

(

i

,

j

)

=

-

t

=

1

n

p

s

(

t

)

Log

(

p

s

(

t

)

)

,

p

s

(

t

)

=

SH

i

,

j

w

(

t

)

/

k

=

1

n

SH

i

,

j

w

(

k

)

;

and

wherein SH w i,j (t) is a spatial phase histogram whose probability distribution function is p s (t), and n is a number of histogram bins.

12. The method of claim 1 , wherein C t is calculated with respect to a sliding window of size L frames along time t axis as:

Ct

(

i

,

j

)

=

-

t

=

1

n

p

t

(

t

)

Log

(

p

t

(

t

)

)

,

p

t

(

t

)

=

TH

i

,

j

L

(

t

)

/

k

=

1

n

TH

i

,

j

L

(

k

)

;

and

wherein TH L i,j (t) is a temporal phase histogram whose probability distribution function is p t (t), and n is a number of histogram bins.

13. The method of claim 1 , wherein B is calculated according to

B=I×C t ×(1 −I×C s ).

14. The method of claim 1 , wherein the motion attention model is calculated according to:

M

motion

=

(

r

Λ

q

Ω

r

B

q

)

/

N

MB

,

B q being brightness of a block in the motion saliency map, Λ being the set of detected motion attention areas, Ω r denoting a set of blocks in each detected motion attention area, N MB being a number of blocks in a motion field; and

wherein an M motion value for each frame in the video data sequence represents a continuous motion attention curve with respect to time.

15. The method of claim 1 , further comprising calculating motion intensity I values, spatial coherence C s values and temporal coherence C t values from each location of the blocks in the motion fields extracted from the video data sequence.

16. A computer-readable media storing executable instructions that, when executed by one or more processors, perform a method comprising:

generating a motion saliency map B to precisely indicate motion attention areas for each frame in the video data sequence, the motion saliency map being based on intensity I values being represented by a brightness associated with each location on a motion intensity map, spatial coherence C s values being represented by a brightness associated with each location on a spatial coherence map and temporal coherence C t values being represented by a brightness associated with each location on a temporal coherence map in motion fields extracted from the video data sequence;

calculating the motion intensity I values, the spatial coherence C s values and the temporal coherence C t values from each location of blocks in the motion fields extracted from the video data sequence; and

accumulating brightness from the motion intensity map, the spatial coherence map and the temporal coherence map of detected motion attention areas to generate, with respect to time, a motion attention model for the video data sequence.

17. The method of claim 16 , wherein the saliency map B is calculated according to

B=I×C t× (1 −I×C s ).

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Dec 9, 2014
From: MICROSOFT CORPORATION
To: MICROSOFT TECHNOLOGY LICENSING, LLC
Reel/Frame 034543/0001 →