IP Library Granted Patent US 12,597,207
Granted Patent B2
US 12,597,207 · App. 18/498,919 · Granted Apr 7, 2026

Camera reprojection for faces

Inventors: James Allan Booth (Zurich, CH); Elif Albuz (Los Gatos, CA); Peihong Guo (San Mateo, CA); Tong Xiao (San Jose, CA)
Assignee: Meta Platforms Technologies, LLC
G06T17/20G06T7/73G06T15/04G06V20/64G06V40/171G06T2207/30244
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,597,207
App. No.
18/498,919
Granted
Apr 7, 2026
Kind
B2
Abstract

In one embodiment, a computing system may access a first image of a first portion of a face of a user captured by a first camera from a first viewpoint and a second image of a second portion of the face captured by a second camera from a second viewpoint. The system may generate, using a machine-learning model and the first and second images, a synthesized image corresponding to a third portion of the face of the user as viewed from a third viewpoint. The system may access a three-dimensional (3D) facial model representative of the face and generate a texture image for the face by projecting at least the synthesized image onto the 3D facial model from a predetermined camera pose corresponding to the third viewpoint. The system may cause an output image to be rendered using at least the 3D facial model and the texture image.

Claims (41)

1 . A method comprising, by one or more computing systems:

generating, using a machine-learning model, a first image of a first portion of a face of a user from a first viewpoint, and a second image of a second portion of the face of the user from a second viewpoint, a synthesized image corresponding to a third portion of the face of the user as viewed from a third viewpoint, wherein the third viewpoint is different from the first viewpoint and the second viewpoint;

generating a texture image for the face of the user by projecting at least the synthesized image onto a 3D facial model from a predetermined camera pose that captured a ground-truth image for training the machine-learning model;

reprojecting the texture image onto the 3D facial model at the predetermined camera pose, the predetermined camera pose representative of a position of a camera that captured the ground-truth image; and

causing an output image of a facial representation of the user.

2 . The method of claim 1 , wherein the first, second, and third viewpoints are different viewpoints.

3 . The method of claim 1 , wherein the generating of the synthesized image comprises: compiling the first image and the second image into the synthesized image from the third viewpoint.

4 . The method of claim 1 , further comprising training the machine-learning model, wherein the training of the machine-learning model comprises:

accessing multiple sets of images corresponding to different portions of a plurality of faces of a plurality of users; and

processing, using the machine-learning model, each set of the multiple sets of images corresponding to a respective particular portion of the face of a particular user, wherein the processing comprises:

generating, using the machine-learning model, a particular synthesized image from a desired rendering viewpoint by compiling at least one set of the multiple sets of images, the at least one set corresponding to the particular portion of the face of the particular user;

accessing a particular ground-truth image corresponding to the particular portion of the face from the desired rendering viewpoint; and

comparing the particular synthesized image generated using the machine-learning model to the particular ground-truth image.

5 . The method of claim 1 , further comprising:

blending the texture image with a predetermined texture to generate a second texture image, wherein the output image of the facial representation of the user is generated based on the second texture image and the 3D facial model.

6 . The method of claim 5 , wherein the second texture image represents a state of the user at a current time.

7 . The method of claim 1 , wherein causing the output image of the facial representation of the user to be rendered comprises:

sending a rendering package to a second artificial-reality system worn by a second user, the rendering package including (1) the 3D facial model of the user, (2) the texture image for the face of the user, and (3) instructions to render the facial representation of the user based on the 3D facial model and the texture image; and

rendering the output image from a viewpoint of the second user with respect to the user using the rendering package.

8 . The method of claim 1 , further comprising:

accessing a three-dimensional (3D) facial model representative of the face of the user;

identifying one or more facial features in the synthesized image; and

selecting the predetermined camera pose based on comparing locations of the one or more facial features in the synthesized image to predetermined feature locations on the 3D facial model.

9 . The method of claim 8 , further comprising: morphing the 3D facial model based on comparison of the locations of the one or more facial features in the synthesized image to the predetermined feature locations on the 3D facial model.

10 . The method of claim 1 , wherein the first image is captured from a first camera and the second image is captured from a second camera, each of the first camera and the second camera is located inside of an artificial-reality system worn by the user.

11 . The method of claim 1 , wherein the 3D facial model is generated by: morphing, based on at least one or more facial features in the synthesized image, a predetermined 3D facial model representative of a plurality of faces of a plurality of users.

12 . The method of claim 1 , wherein the 3D facial model is a predetermined 3D facial model representative of a plurality of faces of a plurality of users.

13 . The method of claim 1 , wherein the output image of the facial representation of the user is photorealistic.

14 . One or more computer-readable non-transitory storage media embodying software that is operable when executed to:

generate, using a machine-learning model, a first image of a first portion of a face of a user from a first viewpoint, and a second image of a second portion of the face of the user from a second viewpoint, a synthesized image corresponding to a third portion of the face of the user as viewed from a third viewpoint, wherein the third viewpoint is different from the first viewpoint and second viewpoint;

generate a texture image for the face of the user by projecting at least the synthesized image onto a 3D facial model from a predetermined camera pose that captured a ground-truth image for training the machine-learning model;

reproject the texture image onto the 3D facial model at the predetermined camera pose, the predetermined camera pose representative of a position of a camera that captured the ground-truth image; and

cause an output image of a facial representation of the user.

15 . The media of claim 14 , wherein the first, second, and third viewpoints are different viewpoints.

16 . A system comprising: one or more processors; and one or more computer-readable non-transitory storage media coupled to one or more of the processors and comprising instructions operable when executed by one or more of the processors to cause the system to:

generate, using a machine-learning model, a first image of a first portion of a face of a user from a first view point, and a second image of a second portion of the face of the user from a second viewpoint, a synthesized image corresponding to a third portion of the face of the user as viewed from a third viewpoint, wherein the third viewpoint is different from the first viewpoint and the second viewpoint;

generate a texture image for the face of the user by projecting at least the synthesized image onto a 3D facial model from a predetermined camera pose that captured a ground-truth image for training the machine-learning model;

reproject the texture image onto the 3D facial model at the predetermined camera pose, the predetermined camera pose representative of a position of a camera that captured the ground-truth image; and

cause an output image of a facial representation of the user.

17 . The system of claim 16 , wherein the first, second, and third viewpoints are different viewpoints.

18 . The system of claim 16 , wherein the predetermined camera pose is a specific camera pose that captured a ground-truth image used to train the machine-learning model to generate the synthesized image from the third viewpoint.

Assignments (2)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jun 11, 2024
From: BOOTH, JAMES ALLAN; ALBUZ, ELIF; GUO, PEIHONG; XIAO, TONG
To: FACEBOOK TECHNOLOGIES, LLC
Reel/Frame 067690/0476 →
CHANGE OF NAME Recorded Jun 11, 2024
From: FACEBOOK TECHNOLOGIES , LLC
To: META PLATFORMS TECHNOLOGIES, LLC
Reel/Frame 067698/0202 →
Continuity (3)
Continuation 18145592 · Dec 22, 2022
Continuation 17028927 · Sep 22, 2020
Related Publication 20240078754A1 · Mar 7, 2024
References Cited (12)
US 8228315B1 · Starner et al. · 2012 [cited by applicant]
US 8928589B2 · Bi · 2015 [cited by applicant]
US 11270515B2 · Beith et al. · 2022 [cited by applicant]
US 20160070952A1 · Kim et al. · 2016 [cited by applicant]
US 20180158246A1 · Grau et al. · 2018 [cited by applicant]
US 20180336737A1 · Varady et al. · 2018 [cited by applicant]
US 20200245873A1 · Frank et al. · 2020 [cited by applicant]
US 20210007607A1 · Frank · 2021 [cited by examiner]
US 20210142045A1 · Noest · 2021 [cited by examiner]
Green S., “Augmented Reality Keyboard for an Interactive Gesture Recognition System,” Master in Computer Science Thesis, School of Computer Science and Statistics, Trinity College Dublin, May 2016, 76 pages. [cited by applicant]
EPO—International Search Report and Written Opinion for International Application No. PCT/US2021/046060, mailed Nov. 30, 2021, 9 pages. [cited by applicant]
Wall Street Journal, “NEC Turns a Person's Arm into a Keyboard,” Nov. 6, 2015, Retrieved from the Internet: URL: https://www.youtube.com/watch?v=MVDWG33YTsg, 2 pages. [cited by applicant]