IP Library › Granted Patent US 12,573,108
Granted Patent B2
US 12,573,108 · App. 18/295,741 · Granted Mar 10, 2026

Head-pose and gaze redirection

Inventors: Zhen Wang (San Diego, CA); Shiwei Jin (San Diego, CA); Lei Wang (San Diego, CA); Ning Bi (San Diego, CA)
Assignee: QUALCOMM INCORPORATED
G06T11/60G06F3/013G06T3/60G06T7/70G06T9/00G06V10/44G06V10/82G06T2207/30201
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,573,108
App. No.
18/295,741
Granted
Mar 10, 2026
Kind
B2
Abstract

Systems and techniques are described herein for generating an image. For instance, a method for generating an image is provided. The method may include obtaining a source image of a face having source attributes and exhibiting a source pose and source gaze; obtaining at least one of a target pose and a target gaze; and generating a modified image of the face having the source attributes and exhibiting at least one of the target pose and the target gaze.

Claims (104)

1 . An apparatus for generating an image, the apparatus comprising:

at least one memory; and

at least one processor coupled to the at least one memory and configured to:

obtain a source image of a face having source attributes and exhibiting a source pose and a source gaze;

determine a pose-normalization matrix based on the source pose;

determine a gaze-normalization matrix based on the source gaze;

determine a plurality of source features based on the source image, wherein the plurality of source features include a pose feature and a gaze feature;

apply the pose-normalization matrix to the pose feature to generate a normalized pose feature;

apply the gaze-normalization matrix to the gaze feature to generate a normalized gaze feature;

obtain a target pose;

obtain a target gaze;

determine a pose-rotation matrix based on the target pose;

determine a gaze-rotation matrix based on the target gaze;

apply the pose-rotation matrix to the normalized pose feature to generate a rotated pose feature;

apply the gaze-rotation matrix to the normalized gaze feature to generate a rotated gaze feature; and

generate a modified image of the face based on the plurality of source features the rotated pose feature, and the rotated gaze feature, wherein the modified image exhibits the source attributes, the target pose, and the target gaze.

2 . The apparatus of claim 1 , wherein, to determine the plurality of source features, the at least one processor is configured to encode the source image using a machine-learning model.

3 . The apparatus of claim 2 , wherein:

the source image is encoded into the plurality of source features using a convolutional neural network as an encoder;

the plurality of source features are based on a plurality of layers of the convolutional neural network; and

the modified image is generated using a deconvolutional network as a decoder.

4 . The apparatus of claim 1 , wherein the plurality of source features include a plurality of attribute features, wherein, to generate the modified image of the face, the at least one processor is configured to decode the plurality of attribute features the rotated pose feature, and the rotated gaze feature to generate the modified image.

5 . The apparatus of claim 4 , wherein:

the source image is encoded into the plurality of source features using a multi-level attribute encoder;

the plurality of source features are based on a plurality of layers of the multi-level attribute encoder; and

the at least one processor is configured to generate the modified image using multi-channel adaptive attentional denormalization residual blocks to process the plurality of attribute features, the rotated pose feature, and the rotated gaze feature to generate the modified image.

6 . The apparatus of claim 1 , wherein, to obtain the target pose and the target gaze, the at least one processor is further configured to extract the target pose and the target gaze from a target image using a machine-learning model.

7 . The apparatus of claim 1 , wherein at least one of the target pose or the target gaze is directed toward a viewing angle from which the source image of the face is captured.

8 . A method for generating an image, the method comprising:

obtaining a source image of a face having source attributes and exhibiting a source pose and a source gaze;

determining a pose-normalization matrix based on the source pose;

determining a gaze-normalization matrix based on the source gaze;

determining a plurality of source features based on the source image, wherein the plurality of source features include a pose feature and a gaze feature;

applying the pose-normalization matrix to the pose feature to generate a normalized pose feature;

applying the gaze-normalization matrix to the gaze feature to generate a normalized gaze feature;

obtaining a target pose;

obtaining a target gaze;

determining a pose-rotation matrix based on the target pose;

determining a gaze-rotation matrix based on the target gaze;

applying the pose-rotation matrix to the normalized pose feature to generate a rotated pose feature;

applying the gaze-rotation matrix to the normalized gaze feature to generate a rotated gaze feature; and

generating a modified image of the face based on the plurality of source features, the rotated pose feature, and the rotated gaze feature, wherein the modified image exhibits the source attributes, the target pose, and the target gaze.

9 . The method of claim 8 , wherein determining the plurality of source features comprises encoding the source image using a machine-learning model.

10 . The method of claim 9 , wherein:

the source image is encoded into the plurality of source features using a convolutional neural network as an encoder;

the plurality of source features are based on a plurality of layers of the convolutional neural network; and

the modified image is generated using a deconvolutional network as a decoder.

11 . The method of claim 8 , wherein the plurality of source features include a plurality of attribute features, wherein generating the modified image of the face comprises decoding the plurality of attribute features, the rotated pose feature, and the rotated gaze feature to the modified image.

12 . The method of claim 11 , wherein:

the source image is encoded into the plurality of source features using a multi-level attribute encoder;

the plurality of source features are based on a plurality of layers of the multi-level attribute encoder; and

the modified image is generated using multi-channel adaptive attentional denormalization residual blocks to process the plurality of attribute features, the rotated pose feature, and the rotated gaze feature to generate the modified image.

13 . The method of claim 8 , wherein obtaining the target pose and the target gaze comprises extracting the target pose and the target gaze from a target image using a machine-learning model.

14 . The method of claim 8 , wherein at least one of the target pose or the target gaze is directed toward a viewing angle from which the source image of the face is captured.

15 . A non-transitory computer-readable storage medium having stored thereon instructions that, when executed by at least one processor, cause the at least one processor to:

obtain a source image of a face having source attributes and exhibiting a source pose and a source gaze;

determine a pose-normalization matrix based on the source pose;

determine a gaze-normalization matrix based on the source gaze;

determine a plurality of source features based on the source image, wherein the plurality of source features include a pose feature and a gaze feature;

apply the pose-normalization matrix to the pose feature to generate a normalized pose feature;

apply the gaze-normalization matrix to the gaze feature to generate a normalized gaze feature;

obtain a target pose;

obtain a target gaze;

determine a pose-rotation matrix based on the target pose;

determine a gaze-rotation matrix based on the target gaze;

apply the pose-rotation matrix to the normalized pose feature to generate a rotated pose feature;

apply the gaze-rotation matrix to the normalized gaze feature to generate a rotated gaze feature; and

generate a modified image of the face based on the plurality of source features, the rotated pose feature, and the rotated gaze feature, wherein the modified image exhibits the source attributes and at least one of the target pose and the target gaze.

16 . The non-transitory computer-readable storage medium of claim 15 , wherein, to determine the plurality of source features, the instructions, when executed by the at least one processor, cause the at least one processor to encode the source image using a machine-learning model.

17 . The non-transitory computer-readable storage medium of claim 16 , wherein:

the source image is encoded into the plurality of source features using a convolutional neural network as an encoder;

the plurality of source features are based on a plurality of layers of the convolutional neural network; and

the modified image is generated using a deconvolutional network as a decoder.

18 . The non-transitory computer-readable storage medium of claim 15 , wherein the plurality of source features include a plurality of attribute features, wherein, to generate the modified image of the face, the instructions, when executed by the at least one processor, cause the at least one processor to: decode the plurality of attribute features and at least one of the rotated pose feature or the rotated gaze feature to generate the modified image.

19 . The non-transitory computer-readable storage medium of claim 18 , wherein:

the source image is encoded into the plurality of source features using a multi-level attribute encoder;

the plurality of source features are based on a plurality of layers of the multi-level attribute encoder; and

the modified image is generated using multi-channel adaptive attentional denormalization residual blocks to process the plurality of attribute features, the rotated pose feature, and the rotated gaze feature to generate the modified image.

20 . The non-transitory computer-readable storage medium of claim 15 , wherein to obtain the target pose and the target gaze the instructions, when executed by the at least one processor, cause the at least one processor to extract the target pose and the target gaze from a target image using a machine-learning model.

21 . The non-transitory computer-readable storage medium of claim 15 , wherein at least one of the target pose or the target gaze is directed toward a viewing angle from which the source image of the face is captured.

22 . An apparatus for generating an image, the apparatus comprising:

means for obtaining a source image of a face having source attributes and exhibiting a source pose and a source gaze;

means for determining a pose-normalization matrix based on the source pose;

means for determining a gaze-normalization matrix based on the source gaze;

means for determining a plurality of source features based on the source image, wherein the plurality of source features include a pose feature and a gaze feature;

means for applying the pose-normalization matrix to the pose feature to generate a normalized pose feature;

means for applying the gaze-normalization matrix to the gaze feature to generate a normalized gaze feature;

means for obtaining a target pose;

means for obtaining a target gaze;

means for determining a pose-rotation matrix based on the target pose;

means for determining a gaze-rotation matrix based on the target gaze;

means for applying the pose-rotation matrix to the normalized pose feature to generate a rotated pose feature;

means for applying the gaze-rotation matrix to the normalized gaze feature to generate a rotated gaze feature; and

means for generating a modified image of the face based on the plurality of source features, the rotated pose feature, and the rotated gaze feature, wherein the modified image exhibits the source attributes and at least one of the target pose and the target gaze.

23 . The apparatus of claim 22 , further comprising means for encoding the source image into the plurality of source features using a machine-learning model.

24 . The apparatus of claim 23 , wherein:

the source image is encoded into the plurality of source features using a convolutional neural network as an encoder;

the plurality of source features are based on a plurality of layers of the convolutional neural network; and

the modified image is generated using a deconvolutional network as a decoder.

25 . The apparatus of claim 22 , wherein the plurality of source features include a plurality of attribute features, further comprising means for decoding the plurality of attribute features, the rotated pose feature, and the rotated gaze feature to generate modified features to generate the modified image.

26 . The apparatus of claim 25 , wherein:

the source image is encoded into the plurality of source features using a multi-level attribute encoder;

the plurality of source features are based on a plurality of layers of the multi-level attribute encoder; and

the modified image is generated using multi-channel adaptive attentional denormalization residual blocks to process the plurality of attribute features, the rotated pose feature, and the rotated gaze feature to generate the modified image.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Apr 26, 2023
From: WANG, ZHEN; JIN, SHIWEI; WANG, LEI; BI, NING
To: QUALCOMM INCORPORATED
Reel/Frame 063454/0344 →
Continuity (1)
Related Publication 20240338868A1 · Oct 10, 2024
References Cited (9)
US 20210065418A1 · Han · 2021 [cited by examiner]
US 20220050521A1 · Drozdov · 2022 [cited by examiner]
US 20220237829A1 · Ren · 2022 [cited by examiner]
US 20230154088A1 · Duarte · 2023 [cited by examiner]
US 20230254448A1 · Binder · 2023 [cited by examiner]
Ganin Y., et al., “DeepWarp: Photorealistic Image Resynthesis for Gaze Manipulation”, Sep. 17, 2016, Topics in Cryptology—CT-RSA 2020, The Cryptographers Track at the RSA Conference 2020, San Francisco, CA, USA, Feb. 24… [cited by applicant]
International Search Report and Written Opinion—PCT/US2024/012649—ISA/EPO—May 14, 2024. [cited by applicant]
Jindal S., et al., “CUDA-GHR: Controllable Unsupervised Domain Adaptation for Gaze and Head Redirection”, 2023 IEEE/CVF Winter Conference on Applications of Computer Vision (WACV), IEEE, Jan. 2, 2023, XP034291204, Abstr… [cited by applicant]
Zheng Y., et al., “Self-Learning Transformations for Improving Gaze and Head Redirection”, arXiv:2010.12307v1 [cs.Cv], arxiv.org, Cornell University Library, 201 Olin Library Cornell University Ithaca, NY, 14853, Oct. 2… [cited by applicant]