IP Library Granted Patent US 7,848,920
Granted Patent B2
US 7,848,920 · App. 12/166,845 · Granted Dec 7, 2010

Method and system of dynamically adjusting a speech output rate to match a speech input rate

View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 7,848,920
App. No.
12/166,845
Granted
Dec 7, 2010
Kind
B2
Abstract

A method ( 10 ) and system of adjusting a speech output rate to match a speech input rate can include the steps of receiving ( 12 ) speech input, computing ( 14 ) a speech input rate, and dynamically adjusting ( 18 or 26 ) a speech output rate to match the speech input rate. If the type of speech output is TTS, then a rate of TTS can be adjusted ( 18 ). If the type of speech output is recorded and alternate text is available, then steps ( 22 and 24 ) of counting alternate text available from a recorded output and determining an audio file length is used to compute a default output rate to adjust a recorded output rate. If the type is recorded and alternate text is unavailable, then steps ( 21 and 24 ) of obtaining an output word count from a transcription of a recorded speech output and determining an audio file length is used.

Claims (10)

1. A system for dynamically and automatically adjusting a speech output rate to match a speech input rate, the system comprising:

a memory; and

a processor programmed: to receive a speech input, compute a speech input rate from the speech input, determine whether a type of speech output to be provided at the speech output rate is text-to-speech or recorded speech output, and dynamically adjust the speech output rate to match the speech input rate, wherein the processor is programmed to adjust the speech output rate based upon the type of speech output, and wherein the process is further programmed to determine, if the type of speech is recorded, whether alternate text is available, and if alternate text is available, to count the alternate text available from a recorded output and determine an audio file length to compute a default output rate which is used to adjust a recorded output rate to match the input speech rate.

2. The system of claim 1 , wherein the processor is further programmed to adjust a rate of text-to-speech synthesis to match the speech input rate if the type of speech output is text-to-speech.

3. The system of claim 1 , wherein the processor is further programmed to obtain an output word count from a transcription of a recorded speech output and determine an audio file length to compute a default output rate which is used to adjust a recorded output rate to match the input speech rate when the type of speech is recorded and alternate text is unavailable.

4. The system of claim 1 , wherein the processor is further programmed to compute a running average of the rates computed for the last n utterances of the speech input when computing the speech input rate.

5. The system of claim 1 , wherein the processor is further programmed to feed back an estimate of the speech input rate to a speech production mechanism to adjust the speech output rate.

6. A machine-readable storage, having stored thereon a computer program having a plurality of code sections executable by a machine for causing the machine to perform the steps of receiving a speech input, computing a speech input rate from the speech input, determining whether a type of speech output to be provided at the speech output rate is text-to-speech or recorded speech output, and dynamically adjusting the speech output rate to match the speech input rate, wherein the machine-readable storage is programmed to cause the machine to adjust the speech output rate based upon the type of speech output, and wherein the machine-readable storage is further programmed to determine, when the type of speech is recorded, whether alternate text is available, and if alternate text is available, to count the alternate text available from a recorded output and determine an audio file length to compute a default output rate which is used to adjust a recorded output rate to match the input.

7. The machine-readable storage of claim 6 , wherein the machine-readable storage is further programmed to adjust a rate of text-to-speech synthesis to match the speech input rate if the type of speech output is text-to-speech.

8. The machine-readable storage of claim 6 , wherein the machine-readable storage is further programmed to obtain an output word count from a transcription of a recorded speech output and determine an audio file length to compute a default output rate which is used to adjust a recorded output rate to match the input speech rate when the type of speech is recorded and alternate text is unavailable.

Assignments (2)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Nov 9, 2023
From: NUANCE COMMUNICATIONS, INC.
To: MICROSOFT TECHNOLOGY LICENSING, LLC
Reel/Frame 065530/0871 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Mar 2, 2009
From: INTERNATIONAL BUSINESS MACHINES CORPORATION
To: NUANCE COMMUNICATIONS, INC.
Reel/Frame 022330/0088 →