IP Library Granted Patent US 12,106,413
Granted Patent B1
US 12,106,413 · App. 17/806,543 · Granted Oct 1, 2024

Joint autoencoder for identity and expression

Inventor: Andrew P. Mason (Cottesloe, AU)
Assignee: Apple Inc.
G06T13/40G06T17/20G06V40/172G06V40/174
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,106,413
App. No.
17/806,543
Granted
Oct 1, 2024
Kind
B1
Abstract

Rendering an avatar for a user in a communication session includes obtaining latent variables for expression and identity of a face, applying a concatenation of the latent variables to an expression decoder of a trained asymmetric joint autoencoder for expression and identification to obtain an expression mesh of a face, and rendering an avatar using the expression mesh of the face. The asymmetric joint autoencoder includes two encoders and two decoders, where an expression decoder utilizes the concatenated latents and the identity decoder utilizes the identity latents and not the expression latents.

Claims (52)

1. A method for generating an expressive avatar comprising:

obtaining expression latent variables for a face, wherein the expression latent variables are derived from an input mesh of a facial expression of the face;

obtaining identity latent variables for the face;

applying a concatenation of the expression latent variables to an expression decoder of a trained asymmetric joint autoencoder for expression and identification to obtain an expression mesh of the face; and

generating an avatar based on the expression mesh of the face.

2. The method of claim 1 , further comprising:

applying the identity latent variables to an identity decoder of the asymmetric joint autoencoder to obtain an identity mesh for the face.

3. The method of claim 2 , wherein the expression latent variables are usable as input for the expression decoder, and wherein the expression latent variables are not usable as input for the identity decoder.

4. The method of claim 1 , further comprising:

receiving additional expression latent variables for a face;

applying a concatenation of the additional expression latent variables and identity latent variables to the expression decoder to obtain an additional expression mesh of the face; and

generating an avatar based on the additional expression mesh of the face.

5. The method of claim 4 , wherein the expression latent variables are received from a remote device and wherein the identity latent variables are received less frequently than the expression latent variables.

6. The method of claim 1 , wherein the expression latent variables are obtained from an expression encoder of the trained asymmetric joint autoencoder, and wherein the identity latent variables are obtained from an identity encoder of the trained asymmetric joint autoencoder.

7. The method of claim 1 , wherein the asymmetric joint autoencoder comprises:

an identity encoder portion configured to produce the identity latent variables,

an expression encoder portion configured to produce the expression latent variables, and

the expression decoder configured to produce an expression mesh based on the identity latent variables and the expression latent variables.

8. The method of claim 7 , wherein the asymmetric joint autoencoder further comprises an identity decoder portion configured to generate an output mesh from the identity latent variables and not from the expression latent variables.

9. A non-transitory computer-readable medium comprising computer-readable code executable by one or more processors to:

obtain expression latent variables for a face, wherein the expression latent variables are derived from an input mesh of a facial expression of the face;

obtain identity latent variables for the face;

apply a concatenation of the expression latent variables to an expression decoder of a trained asymmetric joint autoencoder for expression and identification to obtain an expression mesh of the face; and

generate an avatar based on the expression mesh of the face.

10. The non-transitory computer-readable medium of claim 9 , further comprising computer-readable code to:

apply the identity latent variables to an identity decoder of the asymmetric joint autoencoder to obtain an identity mesh for the face.

11. The non-transitory computer-readable medium of claim 10 , wherein the expression latent variables are usable as input for the expression decoder and wherein the expression latent variables are not usable as input for the identity decoder.

12. The non-transitory computer-readable medium of claim 9 , further comprising:

receive additional expression latent variables for a face;

apply a concatenation of the additional expression latent variables and identity latent variables to the expression decoder to obtain an additional expression mesh of the face; and

generate an avatar based on the additional expression mesh of the face.

13. The non-transitory computer-readable medium of claim 12 , wherein the expression latent variables are received from a remote device and wherein the identity are received less frequently than the expression latent variables.

14. The non-transitory computer-readable medium of claim 9 , wherein the expression latent variables are obtained from an expression encoder of the trained asymmetric joint autoencoder and wherein the identity latent variables are obtained from an identity encoder of the trained asymmetric joint autoencoder.

15. The non-transitory computer-readable medium of claim 9 , wherein the asymmetric joint autoencoder comprises:

an identity encoder portion configured to produce the identity latent variables,

an expression encoder portion configured to produce the expression latent variables, and

the expression decoder configured to produce an expression mesh based on the identity latent variables and the expression latent variables.

16. The non-transitory computer readable medium of claim 15 , wherein the asymmetric joint autoencoder further comprises an identity decoder portion configured to generate an output mesh from the identity latent variables and not from the expression latent variables.

17. A system comprising:

one or more processors;

one or computer-readable media comprising computer-readable code executable by the one or more processors to:

obtain expression latent variables for a face, wherein the expression latent variables are derived from an input mesh of a facial expression of the face;

obtain identity latent variables for the face;

apply a concatenation of the expression latent variables to an expression decoder of a trained asymmetric joint autoencoder for expression and identification to obtain an expression mesh of the face; and

generate an avatar based on the expression mesh of the face.

18. The system of claim 17 , further comprising computer readable code to:

apply the identity latent variables to an identity decoder of the asymmetric joint autoencoder to obtain an identity mesh for the face.

19. The system of claim 18 , wherein the expression latent variables are usable as input for the expression decoder and wherein the expression latent variables are not usable as input for the identity decoder.

20. The system of claim 17 , further comprising:

receive additional expression latent variables for a face;

apply a concatenation of the additional expression latent variables and identity latent variables to the expression decoder to obtain an additional expression mesh of the face; and

generate an avatar based on the additional expression mesh of the face.

Continuity (1)
Provisional Application 63210687 · Jun 15, 2021
Cited By (2)
US 12,531,575 US 12,682,530