IP Library Granted Patent US 9,445,047
Granted Patent B1
US 9,445,047 · App. 14/220,721 · Granted Sep 13, 2016

Method and apparatus to determine focus of attention from video

View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 9,445,047
App. No.
14/220,721
Granted
Sep 13, 2016
Kind
B1
Abstract

A method and system include identifying, by a processing device, at least one media clip captured by at least one camera for an event, detecting at least one human object in the at least one media clip, and calculating, by the processing device, a region in the at least one media clip containing a focus of attention of the detected human object.

Claims (51)

1. A method, comprising:

identifying, by a processing device, at least one media clip representing a scene including a plurality of human objects;

determining a respective location and a respective gaze direction associated with each of the plurality of human objects;

partitioning the scene into a plurality of cells;

calculating a respective viewing cone associated with each of the plurality of human objects based on the respective location and the respective gaze direction, wherein the respective viewing cone overlaps one or more of the plurality of cells;

for each cell, determining an accumulated attention value based on viewing cones that overlap a respective cell; and

calculating, by the processing device, a region in the at least one media clip containing a focus of attention of the plurality of human objects based on accumulated attention values associated with the plurality of cells.

2. The method of claim 1 , further comprising:

detecting the plurality of human objects in the at least one media clip, wherein detecting the plurality of human objects comprises:

detecting a plurality of faces in the at least one media clip;

determining a location of each of the plurality of human objects based on a respective detected face; and

determining a gaze direction of each of the plurality of human objects based on the respective detected face.

3. The method of claim 1 , wherein an attention value for a respective human object of the plurality of human objects is calculated as a function of a distance from a corresponding cell to the respective human object.

4. The method of claim 1 , wherein an attention value for a respective human object of the plurality of human objects is calculated as a function of a distance from a central axis of a viewing cone of the respective human object to a corresponding cell.

5. The method of claim 1 , wherein an attention value from for a respective human object of the plurality of human objects is determined based on whether a viewing cone of the respective human object overlaps with a corresponding cell.

6. The method of claim 1 , wherein the scene is partitioned into one of 2D cells and 3D cells.

7. The method of claim 1 , further comprising:

generating an edited media clip of the at least one media clip using the region containing the focus of attention, wherein the at least one media clip comprises multiple video clips, wherein generating the edited media clip comprises selecting, from the multiple video clips, video frames which best display the region containing the focus of attention.

8. The method of claim 1 , further comprising:

generating an edited media clip of the at least one media clip using the region containing the focus of attention, wherein the at least one media clip comprises an image, and wherein generating the edited media clip comprises selecting a region within the image which best displays the region containing the focus of attention.

9. A non-transitory machine-readable storage medium storing instructions which, when executed, cause a processing device to perform operations comprising:

identifying, by the processing device, at least one media clip representing a scene including a plurality of human objects;

determining a respective location and a respective gaze direction associated with each of the plurality of human objects;

partitioning the scene into a plurality of cells;

calculating a respective viewing cone associated with each of the plurality of human objects based on the respective location and the respective gaze direction, wherein the respective viewing cone overlaps one or more of the plurality of cells;

for each cell, determining an accumulated attention value based on viewing cones that overlap a respective cell; and

calculating a region in the at least one media clip containing a focus of attention of the plurality of human objects based on accumulated attention values associated with the plurality of cells.

10. The machine-readable storage medium of claim 9 , wherein the operations further comprise:

detecting the plurality of human objects in the at least one media clip, wherein detecting the plurality of human objects comprises:

detecting a plurality of faces in the at least one media clip;

determining a location of each of the plurality of human objects based on a respective detected face;

determining a gaze direction of each of plurality of human objects based on the respective detected face.

11. The machined-readable storage medium of claim 9 , wherein an attention value for a respective human object of the plurality of human objects is calculated as a function of a distance from a corresponding cell to the respective human object.

12. The machine-readable storage medium of claim 9 , wherein an attention value for a respective human object of the plurality of human objects is calculated as a function of a distance from a central axis of a viewing cone of the respective human object to a corresponding cell.

13. The machine-readable storage medium of claim 9 , wherein the operations further comprises:

generating an edited media clip of the at least one media clip using the region containing the focus of attention, wherein the at least one media clip comprises multiple video clips, wherein generating the edited media clip comprises selecting, from the multiple video clips, video frames which best display the region containing the focus of attention.

14. A system comprising:

a memory; and

a processing device, communicably coupled to the memory, to:

identify at least one media clip representing a scene including a plurality of human objects;

determine a respective location and a respective gaze direction associated with each of the plurality of human objects;

partition the scene into a plurality of cells;

calculate a respective viewing cone associated with each of the plurality of human objects based on the respective location and the respective gaze direction, wherein the respective viewing cone overlaps one or more of the plurality of cells;

for each cell, determine an accumulated attention value based on viewing cones that overlap a respective cell; and

calculate a region in the at least one media clip containing a focus of attention of the plurality of human objects based on accumulated attention values associated with the plurality of cells.

15. The system of claim 14 , wherein the processing device is further to detect the plurality of human objects in the at least one media clip, and wherein to detect the plurality of human objects, the processing device is further to:

detect a plurality of faces in the at least one media clip;

determine a location of each of the plurality of human objects based on a respective detected face; and

determine a gaze direction of each of the plurality of human objects based on the respective detected face.

16. The system of claim 14 , wherein an attention value for a respective human object of the plurality of human objects is calculated as a function of a distance from a corresponding cell to the respective human object.

17. The system of claim 14 , wherein an attention value for a respective human object of the plurality of human objects is calculated as a function of a distance from a central axis of a viewing cone of the respective human object to a corresponding cell.

Assignments (2)
CHANGE OF NAME Recorded Oct 2, 2017
From: GOOGLE INC.
To: GOOGLE LLC
Reel/Frame 044566/0657 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Mar 24, 2014
From: FRUEH, CHRISTIAN; BHARAT, KRISHNA; YAGNIK, JAY
To: GOOGLE, INC.
Reel/Frame 032506/0419 →