IP Library › Granted Patent US 12,597,290
Granted Patent B2
US 12,597,290 · App. 18/246,625 · Granted Apr 7, 2026

Three-dimensional (3D) facial feature tracking for autostereoscopic telepresence systems

Inventors: Sascha Haeberling (San Francisco, CA); Jason Lawrence (Seattle, WA)
Assignee: GOOGLE LLC
G06V40/174G06T15/00G06V10/30G06V40/171G06V2201/07
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,597,290
App. No.
18/246,625
Granted
Apr 7, 2026
Kind
B2
Abstract

A method including capturing at least one image using a plurality of cameras, identifying a plurality of facial features associated with the at least one image, stabilizing a location of at least one facial feature of the plurality of facial features, converting the location of the at least one facial feature to a three-dimensional (3D) location, predicting a future location of the at least one facial feature, and rendering a 3D image, on a flat panel display, using the future location of the at least one facial feature.

Claims (92)

1 . A method comprising:

capturing at least one image using a plurality of cameras;

identifying a plurality of facial features associated with the at least one image;

stabilizing a location of at least one facial feature of the plurality of facial features;

converting the location of the at least one facial feature to a three-dimensional (3D) location;

reducing a noise associated with the 3D location of the at least one facial feature;

predicting a future location of the at least one facial feature; and

rendering a 3D image, on a flat panel display, using the future location of the at least one facial feature.

2 . The method of claim 1 , wherein

the at least one image includes two or more images, and

the two or more images are captured at least one of sequentially by the same camera and at the same time by two or more cameras.

3 . The method of claim 1 , wherein the identifying a plurality of facial features includes:

identifying a face associated with the at least one image,

generating a plurality of landmarks corresponding to the plurality of facial features on the face, and

associating a location with each of the plurality of landmarks.

4 . The method of claim 1 , wherein the identifying a plurality of facial features includes:

identifying a first face associated with the at least one image,

identifying a second face associated with the at least one image,

generating a first plurality of landmarks corresponding to the plurality of facial features on the first face,

generating a second plurality of landmarks corresponding to the plurality of facial features on the second face,

associating a first location with each of the first plurality of landmarks, and

associating a second location with each of the second plurality of landmarks.

5 . The method of claim 1 , wherein the stabilizing of the location of the at least one facial feature includes:

selecting a facial feature from at least two of the at least one image,

selecting at least two landmarks associated with the facial feature, and

averaging the location of the at least two landmarks.

6 . The method of claim 1 , wherein the stabilizing of the location of the at least one facial feature includes:

generating a plurality of landmarks corresponding to the plurality of facial features,

selecting a subset of the plurality of landmarks,

determining a motion of a face based on a velocity associated with the location of each landmark of the subset of the plurality of landmarks, and

the stabilizing of the location of the at least one facial feature is based on the motion of the face.

7 . The method of claim 1 , wherein the converting of the location of the at least one facial feature to a 3D location includes:

triangulating the location of the at least one facial feature based on location data generated using images captured by three or more of the plurality of cameras.

8 . The method of claim 1 , wherein the noise is reduced by applying a double exponential filter to the 3D location of the at least one facial feature.

9 . The method of claim 1 , wherein the predicting of the future location of the at least one facial feature includes:

applying a double exponential filter to a sequence of the 3D location of the at least one facial feature, and

selecting a value at a time greater than zero (0) that is based on the double exponential filtered sequence of the 3D location before the time zero (0).

10 . The method of claim 1 , further comprising generating audio using the future location of the at least one facial feature.

11 . A three-dimensional (3D) content system comprising:

a memory including code segments representing a plurality of computer instructions; and

a processor configured to execute the code segments, the computer instructions including:

capturing at least one image using a plurality of cameras;

identifying a plurality of facial features associated with the at least one image;

stabilizing a location of at least one facial feature of the plurality of facial features;

converting the location of the at least one facial feature to a three-dimensional (3D) location;

reducing a noise associated with the 3D location of the at least one facial feature;

predicting a future location of the at least one facial feature; and

rendering a 3D image, on a flat panel display, using the future location of the at least one facial feature.

12 . The 3D content system of claim 11 , wherein

the at least one image includes two or more images, and

the two or more images are captured at least one of sequentially by the same camera and at the same time by two or more cameras.

13 . The 3D content system of claim 11 , wherein the identifying a plurality of facial features includes:

identifying a face associated with the at least one image,

generating a plurality of landmarks corresponding to the plurality of facial features on the face, and

associating a location with each of the plurality of landmarks.

14 . The 3D content system of claim 11 , wherein the identifying a plurality of facial features includes:

identifying a first face associated with the at least one image,

identifying a second face associated with the at least one image,

generating a first plurality of landmarks corresponding to the plurality of facial features on the first face,

generating a second plurality of landmarks corresponding to the plurality of facial features on the second face,

associating a first location with each of the first plurality of landmarks, and

associating a second location with each of the second plurality of landmarks.

15 . The 3D content system of claim 11 , wherein the stabilizing of the location of the at least one facial feature includes:

selecting a facial feature from at least two of the at least one image,

selecting at least two landmarks associated with the facial feature, and

averaging the location of the at least two landmarks.

16 . The 3D content system of claim 11 , wherein the stabilizing of the location of the at least one facial feature includes:

generating a plurality of landmarks corresponding to the plurality of facial features,

selecting a subset of the plurality of landmarks,

determining a motion of a face based on a velocity associated with the location of each landmark of the subset of the plurality of landmarks, and

the stabilizing of the location of the at least one facial feature is based on the motion of the face.

17 . The 3D content system of claim 11 , wherein the converting of the location of the at least one facial feature to a 3D location includes:

triangulating the location of the at least one facial feature based on location data generated using images captured by three or more of the plurality of cameras.

18 . The 3D content system of claim 11 , wherein the noise is reduced by applying a double exponential filter to the 3D location of the at least one facial feature.

19 . The 3D content system of claim 11 , wherein the predicting of the future location of the at least one facial feature includes:

applying a double exponential filter to a sequence of the 3D location of the at least one facial feature, and

selecting a value at a time greater than zero (0) that is based on the double exponential filtered sequence of the 3D location before the time zero (0).

20 . The 3D content system of claim 11 , wherein the computer instructions further include generating audio using the future location of the at least one facial feature.

21 . A non-transitory computer-readable storage medium having stored thereon computer executable program code which, when executed on a computer system, causes the computer system to perform a method comprising:

capturing at least one image using a plurality of cameras;

identifying a plurality of facial features associated with the at least one image;

stabilizing a location of at least one facial feature of the plurality of facial features;

converting the location of the at least one facial feature to a three-dimensional (3D) location;

applying a double exponential filter to the 3D location of the at least one facial feature to reduce a noise associated with the 3D location of the at least one facial feature;

predicting a future location of the at least one facial feature;

determining a display position for rendering a 3D image, on a flat panel display, using the future location of the at least one facial feature; and

determining an audio position for generating audio using the future location of the at least one facial feature.

22 . The non-transitory computer-readable storage medium of claim 21 , wherein the stabilizing of the location of the at least one facial feature includes:

generating a plurality of landmarks corresponding to the plurality of facial features,

selecting a subset of the plurality of landmarks,

determining a motion of a face based on a velocity associated with the location of each landmark of the subset of the plurality of landmarks, and

the stabilizing of the location of the at least one facial feature is based on the motion of the face.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded May 30, 2023
From: HAEBERLING, SASCHA; LAWRENCE, JASON
To: GOOGLE LLC
Reel/Frame 063788/0038 →
Continuity (1)
Related Publication 20230316810A1 · Oct 5, 2023
References Cited (26)
US 7653213B2 · Longhurst et al. · 2010 [cited by applicant]
US 8884949B1 · Lambert et al. · 2014 [cited by applicant]
US 9538167B2 · Welch · 2017 [cited by examiner]
US 10171738B1 · Liang · 2019 [cited by examiner]
US 20140347262A1 · Paek · 2014 [cited by examiner]
US 20190228556A1 · Wang et al. · 2019 [cited by applicant]
US 20190325633A1 · Miller, IV · 2019 [cited by examiner]
US 20200272806A1 · Walker · 2020 [cited by examiner]
CA 3038584A1 · 2020 [cited by examiner]
CN 109683335A · 2019 [cited by applicant]
EP 3429195A1 · 2019 [cited by applicant]
JP 2012029209A · 2012 [cited by applicant]
JP 2015513833A · 2015 [cited by applicant]
KR 20150053730A · 2015 [cited by applicant]
WO 2014190221A1 · 2014 [cited by applicant]
WO 2018049201A1 · 2018 [cited by applicant]
First Office Action (with English translation) for Chinese Application No. 202080106726.7, mailed Sep. 10, 2024, 19 pages. [cited by applicant]
Park, et al., “Gaze Point Detection by Computing the 3D Positions and 3D Motion of Face”, IEICE Transactions on Information and Systems, Information & Systems Society, vol. E83-D, No. 4, pp. 884-894, Apr. 2000, XP000975… [cited by applicant]
Notice of Allowance for Japanese Application No. 2023-532766 (with English translation), mailed Dec. 17, 2024, 6 pages. [cited by applicant]
Office Action for Japanese Application No. 2023-532766 (with English translation), mailed Jul. 16, 2024, 10 pages. [cited by applicant]
International Search Report and Written Opinion for PCT Application No. PCT/US2020/070831, mailed on Aug. 3, 2021, 16 pages. [cited by applicant]
Cech, et al., “A 3D Approach to Facial Landmarks: Detection, Refinement, and Tracking”, 2014 22nd International Conference on Pattern Recognition, 2014, pp. 2173-2178. [cited by applicant]
Laviola, et al., “Double Exponential Smoothing: An Alternative to Kalman Filter-Based Predicitive Tracking”, The Eurographics Association, 2003, 8 pages. [cited by applicant]
Nescher, et al., “Analysis of Short Term Path Prediction of Human Locomotion for Augmented and Virtual Reality Applications”, International Conference on Cyberworlds, 2012, pp. 15-22. [cited by applicant]
Maimone, et al., “Holographic Near-Eye Displays for Virtual and Augmented Reality”, ACM Transactions on Graphics, vol. 36, No. 4, Article 85, Jul. 2017, 16 pages. [cited by applicant]
Perlin, et al., “An Autostereoscopic Display”, SIGGRAPH 2000, pp. 319-326. [cited by applicant]