IP Library Granted Patent US 6,940,454
Granted Patent B2
US 6,940,454 · App. 09/929,516 · Granted Sep 6, 2005

Method and system for generating facial animation values based on a combination of visual and audio information

Assignee: Nevengineering, Inc.
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 6,940,454
App. No.
09/929,516
Granted
Sep 6, 2005
Kind
B2
Abstract

Facial animation values are generated using a sequence of facial image frames and synchronously captured audio data of a speaking actor. In the technique, a plurality of visual-facial-animation values are provided based on tracking of facial features in the sequence of facial image frames of the speaking actor, and a plurality of audio-facial-animation values are provided based on visemes detected using the synchronously captured audio voice data of the speaking actor. The plurality of visual facial animation values and the plurality of audio facial animation values are combined to generate output facial animation values for use in facial animation.

Claims (123)

1. Method for generating facial animation values using a sequence of facial image frames and synchronously captured audio data of a speaking actor, comprising the steps for:

providing a plurality of visual-facial-animation values based on tracking of facial features in the sequence of facial image frames of the speaking actor;

providing a plurality of audio-facial-animation values based on visemes detected using the synchronously captured audio voice data of the speaking actor; and

combining the plurality of visual facial animation values and the plurality of audio facial animation values to generate output facial animation values for use in facial animation.

2. Method for generating facial animation values as defined in claim 1 , wherein the output facial animation values associated with a mouth for a facial animation are based only on the respective mouth-associated values of the plurality of audio facial animation values.

3. Method for generating facial animation values as defined in claim 1 , wherein the output facial animation values associated with a mouth for a facial animation are based on a weighted average of the respective mouth-associated values of the plurality of visual facial animation values and the respective mouth-associated values of the plurality of audio facial animation values.

4. Method for generating facial animation values as defined in claim 3 , wherein the output facial animation values are calculated using the following equation:

(

f

_

n

=

σ

_

n

a

σ

_

n

a

+

σ

_

n

v

·

a

_

n

+

σ

_

n

v

σ

_

n

v

+

σ

_

n

a

·

v

_

n

)

i

where:

f n are the output facial animation values;

v n are the visual facial animation values;

a n are the respective mouth-associated values of the audio facial animation values;

σ n a are the weights for the audio facial animation values; and

σ n v are the weights for the visual facial animation values.

5. Method for generating facial animation values as defined in claim 1 , wherein the output facial animation values associated with a mouth for a facial animation are based on Kalman filtering of the respective mouth-associated values of the plurality of visual facial animation values and the respective mouth-associated values of the plurality of audio facial animation values.

6. Method for generating facial animation values as defined in claim 1 , wherein the step of combining the plurality of visual facial animation values and the plurality of audio facial animation values to generate output facial animation values includes detecting whether speech is occurring in the synchronously captured audio voice data of the speaking actor and, while speech is detected as occurring, generating the output facial animation values associated with a mouth based only on the respective mouth-associated values of the plurality of audio facial animation values and, while speech is not detected as occurring, generating the output facial animation values associated with a mouth based only on the respective mouth-associated values of the plurality of visual facial animation values.

7. Method for generating facial animation values as defined in claim 1 , wherein the tracking of facial features in the sequence of facial image frames of the speaking actor is performed using bunch graph matching.

8. Method for generating facial animation values as defined in claim 1 , wherein the tracking of facial features in the sequence of facial image frames of the speaking actor is performed using transformed facial image frames generated based on wavelet transformations.

9. Method for generating facial animation values as defined in claim 1 , wherein the tracking of facial features in the sequence of facial image frames of the speaking actor is performed using transformed facial image frames generated based on Gabor wavelet transformations.

10. Method for generating facial animation values as defined in claim 1 , wherein the tracking of facial features in the sequence of facial image frames of the speaking actor is performed without using markers attached to the speaking actor's face.

11. Apparatus for generating facial animation values using a sequence of facial image frames and synchronously captured audio data of a speaking actor, comprising:

means for providing a plurality of visual-facial-animation values based on tracking of facial features in the sequence of facial image frames of the speaking actor;

means for providing a plurality of audio-facial-animation values based on visemes detected using the synchronously captured audio voice data of the speaking actor; and

means for providing a plurality of visual-facial-animation values based on tracking of facial features in the sequence of facial image frames of the speaking actor;

means for combining the plurality of visual facial animation values and the plurality of audio facial animation values to generate output facial animation values for use in facial animation.

12. Apparatus for generating facial animation values as defined in claim 11 , wherein the output facial animation values associated with a mouth for a facial animation are based only on the respective mouth-associated values of the plurality of audio facial animation values.

13. Apparatus for generating facial animation values as defined in claim 11 , wherein the output facial animation values associated with a mouth for a facial animation are based on a weighted average of the respective mouth-associated values of the plurality of visual facial animation values and the respective mouth-associated values of the plurality of audio facial animation values.

14. Apparatus for generating facial animation values as defined in claim 13 , wherein the output facial animation values are calculated using the following equation:

(

f

_

n

=

σ

_

n

a

σ

_

n

a

+

σ

_

n

v

·

a

_

n

+

σ

_

n

v

σ

_

n

v

+

σ

_

n

a

·

v

_

n

)

i

where:

f n are the output facial animation values;

v n are the visual facial animation values;

a n are the respective mouth-associated values of the audio facial animation values;

σ n a are the weights for the audio facial animation values; and

σ n v are the weights for the visual facial animation values.

15. Apparatus for generating facial animation values as defined in claim 11 , wherein the output facial animation values associated with a mouth for a facial animation are based on Kalman filtering of the respective mouth-associated values of the plurality of visual facial animation values and the respective mouth-associated values of the plurality of audio facial animation values.

16. Apparatus for generating facial animation values as defined in claim 11 , wherein the means for combining the plurality of visual facial animation values and the plurality of audio facial animation values to generate output facial animation values includes means for detecting whether speech is occurring in the synchronously captured audio voice data of the speaking actor and, while speech is detected as occurring, generating the output facial animation values associated with a mouth based only on the respective mouth-associated values of the plurality of audio facial animation values and, while speech is not detected as occurring, generating the output facial animation values associated with a mouth based only on the respective mouth-associated values of the plurality of visual facial animation values.

17. Apparatus for generating facial animation values as defined in claim 11 , wherein the tracking of facial features in the sequence of facial image frames of the speaking actor is performed using bunch graph matching.

18. Apparatus for generating facial animation values as defined in claim 11 , wherein the tracking of facial features in the sequence of facial image frames of the speaking actor is performed using transformed facial image frames generated based on wavelet transformations.

19. Apparatus for generating facial animation values as defined in claim 11 , wherein the tracking of facial features in the sequence of facial image frames of the speaking actor is performed using transformed facial image frames generated based on Gabor wavelet transformations.

20. Apparatus for generating facial animation values as defined in claim 11 , wherein the tracking of facial features in the sequence of facial image frames of the speaking actor is performed without using markers attached to the speaking actor's face.

Assignments (5)
CHANGE OF NAME Recorded Oct 5, 2017
From: GOOGLE INC.
To: GOOGLE LLC
Reel/Frame 044127/0735 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Dec 13, 2006
From: NEVENGINEERING, INC.
To: GOOGLE INC.
Reel/Frame 018616/0814 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Mar 2, 2004
From: EYEMATIC INTERFACES, INC.
To: NEVENGINEERING, INC.
Reel/Frame 015023/0038 →
SECURITY INTEREST Recorded Oct 9, 2003
From: NEVENGINEERING, INC.
To: EYEMATIC INTERFACES, INC.
Reel/Frame 014572/0440 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jan 29, 2002
From: PAETZOLD, FRANK; BUDDENMEIER, ULRICH F.; DZHURINKSIY, YEVGENIY V.; DERLICH, KARIN M.; NEVEN, HARTMUT
To: EYEMATIC INTERFACES, INC.
Reel/Frame 012565/0612 →
Continuity (4)
Continuation In Part 0987137000 · May 31, 2001
Continuation 0918807900 · Nov 6, 1998
Provisional Application 6008161500 · Apr 13, 1998
Related Publication 20020118195A1 · Aug 29, 2002