IP Library Granted Patent US 7,366,670
Granted Patent B1
US 7,366,670 · App. 11/464,018 · Granted Apr 29, 2008

Method and system for aligning natural and synthetic video to speech synthesis

View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 7,366,670
App. No.
11/464,018
Granted
Apr 29, 2008
Kind
B1
Abstract

Facial animation in MPEG-4 can be driven by a text stream and a Facial Animation Parameters (FAP) stream. Text input is sent to a TTS converter that drives the mouth shapes of the face. FAPs are sent from an encoder to the face over the communication channel. Disclosed are codes bookmarks in the text string transmitted to the TTS converter. Bookmarks are placed between and inside words and carry an encoder time stamp. The encoder time stamp does not relate to real-world time. The FAP stream carries the same encoder time stamp found in the bookmark of the text. The system reads the bookmark and provides the encoder time stamp as well as a real-time time stamp to the facial animation system. The facial animation system associates the correct facial animation parameter with the real-time time stamp using the encoder time stamp of the bookmark as a reference.

Claims (23)

1. A computer-readable medium storing instructions for controlling a computing device to encode an animation comprising at least one animation mimic within a first stream and speech associated with a second stream, the instructions comprising:

assigning a predetermined code that points to an animation mimic within a first stream; and

synchronizing a second stream with the animation mimics stream by placing the predetermined code within the second stream.

2. The computer-readable medium of claim 1 , wherein the first stream is an animation mimics stream and the second stream is a text stream.

3. The computer-readable medium of claim 1 , wherein the animation mimic is a facial mimic.

4. The computer-readable medium of claim 1 , wherein the predetermined code comprises an escape sequence followed by a plurality of bits, which define one of a set of possible animation mimics.

5. The computer-readable medium of claim 2 , wherein the instructions further comprise encoding the animation mimic stream and the text stream containing the predetermined code.

6. The computer-readable medium of claim 1 , wherein the instructions further comprise placing the predetermined code in between words in the second stream.

7. The computer-readable medium of claim 1 , wherein the instructions further comprise placing the predetermined code in between letters in the second stream.

8. A computer-readable medium storing instructions for controlling a computing device to decode an animation including speech and at least one animation mimic, the instructions comprising:

monitoring a first stream for a predetermined code that points to an animation mimic within a second stream thereby indicating a synchronization relationship between the first stream and the second stream; and

sending a signal to a visual decoder to start the animation mimic that is pointed to by the predetermined code.

9. The computer-readable medium of claim 8 , wherein the first stream is a text stream and the second stream is an animation mimics stream.

10. The computer-readable medium of claim 8 , wherein the correspondence between the predetermined code and the animation mimic is established during an encoding process of the first stream.

11. The computer-readable medium of claim 8 , wherein the animation mimic is a facial mimic.

12. The computer-readable medium of claim 8 , wherein the first stream and the second stream are distinctly maintained streams.

13. A system for decoding at least one stream of data, the system comprising:

a module that monitors a first stream for a predetermined code that points to an animation mimic within a second stream thereby indicating a synchronization relationship between the first stream and the second stream; and

a module that sends a signal to a visual decoder to start the animation mimic that is pointed to by the predetermined code.

14. The system of claim 13 , wherein the first stream is a text stream and the second stream is an animation mimic stream.

15. The system of claim 13 , wherein the animation mimic is a facial mimic.

16. The system of claim 13 , wherein the first stream and the second stream are distinctly maintained streams.

17. The system of claim 13 , wherein the predetermined code is placed according to one of: between words in the first stream, inside words in the first stream, or between phonemes within the first stream.

Assignments (4)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jan 26, 2017
From: AT&T INTELLECTUAL PROPERTY II, L.P.
To: NUANCE COMMUNICATIONS, INC.
Reel/Frame 041498/0316 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Apr 26, 2016
From: AT&T CORP.
To: AT&T PROPERTIES, LLC
Reel/Frame 038529/0164 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Apr 26, 2016
From: AT&T PROPERTIES, LLC
To: AT&T INTELLECTUAL PROPERTY II, L.P.
Reel/Frame 038529/0240 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Apr 14, 2016
From: BASSO, ANDREA; BEUTNAGEL, MARK CHARLES; OSTERMANN, JOERN
To: AT&T CORP.
Reel/Frame 038279/0845 →