IP Library Granted Patent US 9,703,373
Granted Patent B2
US 9,703,373 · App. 14/694,131 · Granted Jul 11, 2017

User interface control using gaze tracking

Inventors: Kai Jochen Kohlhoff (Foster City, CA); Xiaoyi Zhang (Los Angeles, CA)
Assignee: Google Inc.
G06F3/013G06F3/012G06F3/0482G06F3/04842G06K9/0061G06K9/00248G06K9/00261G06K9/00604G06T7/74G06T2207/30201
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 9,703,373
App. No.
14/694,131
Granted
Jul 11, 2017
Kind
B2
Abstract

Methods, systems, and apparatus, including computer programs encoded on a computer storage medium, for identifying a direction in which a user is looking. In one aspect, a method includes receiving an image of a sequence of images. The image can depict a face of a user. A template image for each particular facial feature point can be compared to one or more image portions of the image. The template image for the particular facial feature point can include a portion of a previous image of the sequence of images that depicted the facial feature point. Based on the comparison, a matching image portion of the image that matches the template image for the particular facial feature point is identified. A location of the matching image portion is identified in the image. A direction in which the user is looking is determined based on the identified location for each template image.

Claims (70)

1. A method performed by data processing apparatus, the method comprising:

receiving a first image of a sequence of images, the first image determined to depict at least a face of a user;

in response to determining that the first image depicts the face, fitting a shape model to the face, the fitting including:

identifying, for the face, a shape model that includes one or more facial feature points; and

fitting the shape model to the face in the first image to generate a fitted shape model by adjusting a location of each of the facial feature points of the shape model to overlap with a corresponding facial feature point of the face in the first image; and

generating, from the first image and based on the fitted shape model, a template image for each facial feature point of the face in the first image, the template image for each facial feature point of the face depicting a portion of the face at a location of the facial feature point of the face in the first image, the portion of the face for each template image being less than all of the face;

for each subsequent image in the sequence of images:

for each facial feature point of the face:

comparing the template image for the facial feature point of the face to a respective image portion of the subsequent image located at a same location in the subsequent image as a location at which the facial feature point of the face was identified in a previous image; and

for at least one facial feature point of the face for which the facial feature point's template image does not match the respective image portion, comparing the template image for the at least one facial feature point to one or more additional image portions of the subsequent image until a match is found between the template image for the at least one facial feature point and one of the one or more additional image portions; and

determining, for the subsequent image, a direction in which the user is looking based on a location of a matching image portion in the subsequent image for each facial feature point of the face.

2. The method of claim 1 , further comprising identifying the one additional image portion of the subsequent image for the at least one facial feature point by:

determining a similarity score for each of the one or more additional image portions of the subsequent image, the similarity score for a particular additional image portion of the subsequent image specifying a level of similarity between the template image for the at least one facial feature point and the particular additional image portion; and

selecting the additional image portion having the greatest similarity score as a matching image portion.

3. The method of claim 1 , wherein determining the direction in which the user is looking based on the location of a matching image portion in the subsequent image for each facial feature point of the face comprises determining an orientation of the user's face in the subsequent image based on a relative location of the matching image portion for each facial feature point of the face.

4. The method of claim 1 , further comprising:

determining that the user is looking at a particular object presented on a display based on the determined direction in which the user is looking; and

causing an operation to be performed on the particular object in response to determining that the user is looking at the particular object.

5. The method of claim 4 , wherein determining, for the subsequent image, that the user is looking at a particular object presented on the display comprises determining that the user is looking at the particular object based on a location of a pupil of the user in the subsequent image and the determined direction.

6. The method of claim 1 , further comprising:

determining a portion of a display at which the user is looking based on the determined direction in which the user is looking;

detecting a location of a pupil of the user in the subsequent image; and

determining a sub-portion of the portion of the display at which the user is looking based on the location of the pupil.

7. The method of claim 6 , further comprising:

receiving an additional subsequent image of the sequence of images, the additional subsequent image being received subsequent to receipt of the subsequent image;

determining that the user is looking at the portion of the display based on the additional subsequent image;

determining that a location of the pupil in the additional subsequent image is different from the location of the pupil in the subsequent image; and

causing an operation to be performed on a user interface element displayed in the portion of the display in response to determining that the location of the pupil in the additional subsequent image is different from the location of the pupil in the subsequent image.

8. The method of claim 7 , wherein the operation comprises moving a cursor across the display based on the location of the pupil in the subsequent image and the location of the pupil in the additional subsequent image.

9. A system, comprising:

a data processing apparatus; and

a memory storage apparatus in data communication with the data processing apparatus, the memory storage apparatus storing instructions executable by the data processing apparatus and that upon such execution cause the data processing apparatus to perform operations comprising:

receiving a first image of a sequence of images, the first image determined to depict at least a face of a user;

in response to determining that the first image depicts the face, fitting a shape model to the face, the fitting including:

identifying, for the face, a shape model that includes one or more facial feature points; and

fitting the shape model to the face in the first image to generate a fitted shape model by adjusting a location of each of the facial feature points of the shape model to overlap with a corresponding facial feature point of the face in the first image; and

generating, from the first image and based on the fitted shape model, a template image for each facial feature point of the face in the first image, the template image for each facial feature point of the face depicting a portion of the face at a location of the facial feature point of the face in the first image, the portion of the face for each template image being less than all of the face;

for each subsequent image in the sequence of images:

for each facial feature point of the face:

comparing the template image for the facial feature point of the face to a respective image portion of the subsequent image located at a same location in the subsequent image as a location at which the facial feature point of the face was identified in a previous image; and

for at least one facial feature point of the face for which the facial feature point's template image does not match the respective image portion, comparing the template image for the at least one facial feature point to one or more additional image portions of the subsequent image until a match is found between the template image for the at least one facial feature point and one of the one or more additional image portions; and

determining, for the subsequent image, a direction in which the user is looking based on a location of a matching image portion in the subsequent image for each facial feature point of the face.

10. The system of claim 9 , wherein the operations further comprise identifying the one additional image portion of the subsequent image for the at least one facial feature point by:

determining a similarity score for each of the one or more additional image portions of the subsequent image, the similarity score for a particular additional image portion of the subsequent image specifying a level of similarity between the template image for the at least one facial feature point and the particular additional image portion; and

selecting the additional image portion having the greatest similarity score as a matching image portion.

11. The system of claim 9 , wherein determining the direction in which the user is looking based on the location of a matching image portion in the subsequent image for each facial feature point of the face comprises determining an orientation of the user's face in the subsequent image based on a relative location of the matching image portion for each facial feature point of the face.

12. The system of claim 9 , wherein the instructions upon execution cause the data processing apparatus to perform further operations comprising:

determining that the user is looking at a particular object presented on a display based on the determined direction in which the user is looking; and

causing an operation to be performed on the particular object in response to determining that the user is looking at the particular object.

13. The system of claim 9 , wherein the instructions upon execution cause the data processing apparatus to perform further operations comprising:

determining a portion of a display at which the user is looking based on the determined direction in which the user is looking;

detecting a location of a pupil of the user in the subsequent image; and

determining a sub-portion of the portion of the display at which the user is looking based on the location of the pupil.

14. The system of claim 13 , wherein the instructions upon execution cause the data processing apparatus to perform further operations comprising:

receiving an additional subsequent image of the sequence of images, the additional subsequent image being received subsequent to receipt of the subsequent image;

determining that the user is looking at the portion of the display based on the additional subsequent image;

determining that a location of the pupil in the additional subsequent image is different from the location of the pupil in the subsequent image; and

causing an operation to be performed on a user interface element displayed in the portion of the display in response to determining that the location of the pupil in the additional subsequent image is different from the location of the pupil in the subsequent image.

15. The system of claim 14 , wherein the operation comprises moving a cursor across the display based on the location of the pupil in the subsequent image and the location of the pupil in the additional subsequent image.

16. A non-transitory computer storage medium encoded with a computer program, the program comprising instructions that when executed by a data processing apparatus cause the data processing apparatus to perform operations comprising:

receiving a first image of a sequence of images, the first image determined to depict at least a face of a user;

in response to determining that the first image depicts the face, fitting a shape model to the face, the fitting including:

identifying, for the face, a shape model that includes one or more facial feature points; and

fitting the shape model to the face in the first image to generate a fitted shape model by adjusting a location of each of the facial feature points of the shape model to overlap with a corresponding facial feature point of the face in the first image; and

generating, from the first image and based on the fitted shape model, a template image for each facial feature point of the face in the first image, the template image for each facial feature point of the face depicting a portion of the face at a location of the facial feature point of the face in the first image, the portion of the face for each template image being less than all of the face;

for each subsequent image in the sequence of images:

for each facial feature point of the face:

comparing the template image for the facial feature point of the face to a respective image portion of the subsequent image located at a same location in the subsequent image as a location at which the facial feature point of the face was identified in a previous image; and

for at least one facial feature point of the face for which the facial feature point's template image does not match the respective image portion, comparing the template image for the at least one facial feature point to one or more additional image portions of the subsequent image until a match is found between the template image for the at least one facial feature point and one of the one or more additional image portions; and

determining, for the subsequent image, a direction in which the user is looking based on a location of a matching image portion in the subsequent image for each facial feature point of the face.

Assignments (2)
CHANGE OF NAME Recorded Oct 2, 2017
From: GOOGLE INC.
To: GOOGLE LLC
Reel/Frame 044097/0658 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Aug 21, 2015
From: KOHLHOFF, KAI JOCHEN; ZHANG, XIAOYI
To: GOOGLE INC.
Reel/Frame 036387/0923 →
Continuity (2)
Provisional Application 61983232 · Apr 23, 2014
Related Publication 20150309569A1 · Oct 29, 2015