IP Library Granted Patent US 8,078,466
Granted Patent B2
US 8,078,466 · App. 12/627,373 · Granted Dec 13, 2011

Coarticulation method for audio-visual text-to-speech synthesis

View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 8,078,466
App. No.
12/627,373
Granted
Dec 13, 2011
Kind
B2
Abstract

A method for generating animated sequences of talking heads in text-to-speech applications wherein a processor samples a plurality of frames comprising image samples. The processor reads first data comprising one or more parameters associated with noise-producing orifice images of sequences of at least three concatenated phonemes which correspond to an input stimulus. The processor reads, based on the first data, second data comprising images of a noise-producing entity. The processor generates an animated sequence of the noise-producing entity.

Claims (27)

1. A method of synchronizing synthesized speech and animation, the method comprising:

associating, by a computing device, a received stimulus with a phoneme having corresponding mouth parameters in a coarticulation library;

selecting, by the computing device, a parameter set corresponding to the mouth parameters from an animation library, the parameter set representing frame segments; and

generating, via a noise producing entity, speech associated with the stimulus that is synchronized with the frame segments and overlaying the frame segments on a larger entity to synthesize a whole animated image.

2. The method of claim 1 , wherein the stimulus is text.

3. The method of claim 2 , wherein the stimulus is derived from speech recognition.

4. The method of claim 2 , wherein the stimulus is derived from speech recognition.

5. The method of claim 1 , wherein the speech is output using a phoneme transcript stored in the coarticulation library.

6. The method of claim 1 , further comprising iteratively applying the method to phoneme sequences in the stimulus to form a complete animation.

7. The method of claim 1 , wherein the parameter set is associated with images of at least three concatenated phonemes with correspond to the stimulus.

8. The method of claim 1 , wherein the stimulus is text.

9. The method of claim 1 , wherein the speech is output using a phoneme transcript stored in the coarticulation library.

10. The method of claim 1 , further comprising iteratively applying the method to phoneme sequences in the stimulus to form a complete animation.

11. A system for synchronizing synthesized speech and animation, the system comprising:

a processor;

a first module controlling the processor to associate a received stimulus with a phoneme having corresponding mouth parameters in a coarticulation library;

a second module controlling the processor to select a parameter set corresponding to the mouth parameters from an animation library, the parameter set representing frame segments; and

a third module controlling the processor to generate, via a noise producing entity, speech associated with the stimulus that is synchronized with the frame segments and to overlay the frame segments on a larger entity to synthesize a whole animated image.

12. The system of claim 11 , wherein the stimulus is text.

13. The system of claim 12 , wherein the stimulus is derived from speech recognition.

14. The system of claim 11 , wherein the speech is output using a phoneme transcript stored in the coarticulation library.

15. The system of claim 11 , further comprising a fourth module controlling the processor to iteratively apply the method to phoneme sequences in the stimulus to form a complete animation.

16. The system of claim 11 , wherein the parameter set is associated with images of at least three concatenated phonemes with correspond to the stimulus.

17. A method of synchronizing synthesized speech and animation, the method comprising:

associating, by a computing device, a received stimulus with a phoneme having corresponding mouth parameters in a coarticulation library;

selecting, by the computing device, a parameter set corresponding to the mouth parameters from an animation library, the parameter set representing frame segments; and

generating, via a noise producing entity, speech associated with the stimulus that is synchronized with the frame segments and overlaying the frame segments on a larger entity to synthesize a whole animated image.

Assignments (4)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jan 26, 2017
From: AT&T INTELLECTUAL PROPERTY II, L.P.
To: NUANCE COMMUNICATIONS, INC.
Reel/Frame 041498/0316 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Apr 26, 2016
From: AT&T CORP.
To: AT&T PROPERTIES, LLC
Reel/Frame 038529/0164 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Apr 26, 2016
From: AT&T PROPERTIES, LLC
To: AT&T INTELLECTUAL PROPERTY II, L.P.
Reel/Frame 038529/0240 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Apr 14, 2016
From: COSATTO, ERIC; GRAF, HANS PETER; SCHROETER, JUERGEN
To: AT&T CORP.
Reel/Frame 038279/0587 →