IP Library Granted Patent US 7,280,968
Granted Patent B2
US 7,280,968 · App. 10/396,077 · Granted Oct 9, 2007

Synthetically generated speech responses including prosodic characteristics of speech inputs

View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 7,280,968
App. No.
10/396,077
Granted
Oct 9, 2007
Kind
B2
Abstract

A method for digitally generating speech with improved prosodic characteristics can include receiving a speech input, determining at least one prosodic characteristic contained within the speech input, and generating a speech output including the prosodic characteristic within the speech output.

Claims (63)

1. A method for synthetically generating speech with improved prosodic characteristics comprising the steps of:

storing at least one pre-defined prosodic characteristic;

receiving a speech input;

extracting at least one prosodic characteristic contained within said speech input;

selecting at least one prosodic characteristic for generating a speech output, wherein said at least one prosodic characteristic is selected from said at least one extracted prosodic characteristic and said at least one pre-defined prosodic characteristic, wherein said at least one pre-defined prosodic characteristic is selected if said at least one extracted prosodic characteristic matches at least one condition; and,

generating a speech output including said at least one selected prosodic characteristic within said speech output.

2. The method of claim 1 , wherein said speech output includes a portion of said speech input, wherein said portion of said speech output utilizes said at least one prosodic characteristic.

3. The method of claim 1 , further comprising the steps of:

upon completing said determining step, storing said prosodic characteristic into a data store; and,

before said generating step, retrieving said prosodic characteristic from said data store.

4. The method of claim 1 , wherein said receiving step occurs during a first session and wherein said generating step occurs during a second session, and wherein said first session and said second session represent two different interactive periods for a common user.

5. The method of claim 1 , wherein at least one prosodic characteristic is selected from the group consisting of the speed before and after a word, the pause before and after a word, the rhyme of words, the relative tones of a word, the relative stresses applied to a word, the relative stresses applied to a syllable, and the relative stresses applied to a syllable combination.

6. The method of claim 1 , wherein said receiving step and said generating step are performed by an interactive voice response system.

7. The method of claim 1 , further comprising the steps of:

converting said speech input into an input text string; and,

performing a function responsive to said converting step.

8. The method of claim 7 , further comprising the steps of:

generating an output text string responsive to said performing step; and,

converting said output text string into said speech output.

9. The method of claim 1 , further comprising the steps of:

identifying a part of speech associated with at least one word within said speech input; and,

detecting said at least one prosodic characteristic for said selected part of speech.

10. The method of claim 9 , wherein said part of speech is a proper noun.

11. A system for generating synthetic speech comprising:

a speech recognition component capable of extracting at least one prosodic characteristic from speech input;

a prosodic characteristic store configured to store and permit retrieval of said at least one extracted prosodic characteristic and at least one pre-defined prosodic characteristic; and,

a text-to-speech component capable of modifying at least a portion of synthetically generated speech based upon at least one prosodic characteristic, wherein said at least one prosodic characteristic is selected from said at least one extracted prosodic characteristic and said at least pre-defined prosodic characteristics, wherein said at least one pre-defined prosodic characteristic is selected if said at least one extracted prosodic characteristic matches at least one condition.

12. The system of claim 11 , wherein said system is an interactive voice response system.

13. A machine readable storage having stored thereon, a computer program having a plurality of code sections, said code sections executable by a machine for causing the machine to perform the steps of:

storing at least one pre-defined prosodic characteristic;

receiving a speech input;

extracting at least one prosodic characteristic contained within said speech input;

selecting at least one prosodic characteristic for generating a speech output, wherein said at least one prosodic characteristic is selected from said at least one extracted prosodic characteristic and said at least one pre-defined prosodic characteristic, wherein said at least one pre-defined prosodic characteristic is selected if said at least one extracted prosodic characteristic matches at least one condition; and,

generating a speech output including said at least one selected prosodic characteristic within said speech output.

14. The machine readable storage of claim 13 , wherein said speech output includes a portion of said speech input, wherein said portion of said speech output utilizes said at least one prosodic characteristic.

15. The machine readable storage of claim 13 , further comprising the steps of:

upon completing said determining step, storing said prosodic characteristic into a data store; and,

before said generating step, retrieving said prosodic characteristic from said data store.

16. The machine readable storage of claim 13 , wherein said receiving step occurs during a first session and wherein said generating step occurs during a second session, and wherein said first session and said second session represent two different interactive periods for a common user.

17. The machine readable storage of claim 13 , wherein at least one prosodic characteristic is selected from the group consisting of the speed before and after a word, the pause before and after a word, the rhyme of words, the relative tones of a word, the relative stresses applied to a word, the relative stresses applied to a syllable, and the relative stresses applied to a syllable combination.

18. The machine readable storage of claim 13 , wherein said receiving step and said generating step are performed by an interactive voice response system.

19. The machine readable storage of claim 13 , further comprising the steps of:

converting said speech input into an input text string; and,

performing a function responsive to said converting step.

20. The machine readable storage of claim 19 , further comprising the steps of:

generating an output text string responsive to said performing step; and,

converting said output text string into said speech output.

21. The machine readable storage of claim 13 , further comprising the steps of:

identifying a part of speech associated with at least one word within said speech input; and,

detecting said at least one prosodic characteristic for said selected part of speech.

22. The machine readable storage of claim 21 , wherein said part of speech is a proper noun.

23. A method for synthetically generating speech comprising the steps of:

receiving a speech input;

analyzing said speech input to generate special handling instructions, said instructions comprising generating speech using at least one prosodic characteristic selected from said at least one extracted prosodic characteristic and said at least one pre-defined prosodic characteristic, wherein said at least one pre-defined prosodic characteristic is selected if said at least one extracted prosodic characteristic matches at least one condition; and,

altering at least one speech generation characteristic of a text-to-speech application based upon said special handling instructions.

24. The method of claim 23 , wherein said at least one condition is based upon at least one of a language proficiency level and an emotional state of the listener.

25. The method of claim 23 , wherein said speech generation characteristic alters at least one of clarity and pace of speech output.

26. A machine readable storage having stored thereon, a computer program having a plurality of code sections, said code sections executable by a machine for causing the machine to perform the steps of:

receiving a speech input;

analyzing said speech input to generate special handling instructions, said instructions comprising generating speech using at least one prosodic characteristic selected from said at least one extracted prosodic characteristic and said at least one pre-defined prosodic characteristic, wherein said at least one pre-defined prosodic characteristic is selected if said at least one extracted prosodic characteristic matches at least one condition; and,

altering at least one speech generation characteristic of a text-to-speech application based upon said special handling instructions.

27. The machine readable storage of claim 26 , wherein wherein said at least one condition is based upon at least one of a language proficiency level and an emotional state of the listener.

28. The machine readable storage of claim 26 , wherein said speech generation characteristic alters at least one of clarity and pace of speech output.

Assignments (9)
RELEASE (REEL 052935 / FRAME 0584) Recorded Jan 2, 2025
From: WELLS FARGO BANK, NATIONAL ASSOCIATION
To: CERENCE OPERATING COMPANY
Reel/Frame 069797/0818 →
CORRECTIVE ASSIGNMENT TO CORRECT THE REPLACE THE CONVEYANCE DOCUMENT WITH THE NEW ASSIGNMENT PREVIOUSLY RECORDED AT REEL: 050836 FRAME: 0191. ASSIGNOR(S) HEREBY CONFIRMS THE ASSIGNMENT. Recorded Apr 19, 2022
From: NUANCE COMMUNICATIONS, INC.
To: CERENCE OPERATING COMPANY
Reel/Frame 059804/0186 →
SECURITY AGREEMENT Recorded Jun 15, 2020
From: CERENCE OPERATING COMPANY
To: WELLS FARGO BANK, N.A.
Reel/Frame 052935/0584 →
RELEASE OF SECURITY INTEREST Recorded Jun 12, 2020
From: BARCLAYS BANK PLC
To: CERENCE OPERATING COMPANY
Reel/Frame 052927/0335 →
SECURITY AGREEMENT Recorded Nov 7, 2019
From: CERENCE OPERATING COMPANY
To: BARCLAYS BANK PLC
Reel/Frame 050953/0133 →
CORRECTIVE ASSIGNMENT TO CORRECT THE ASSIGNEE NAME PREVIOUSLY RECORDED AT REEL: 050836 FRAME: 0191. ASSIGNOR(S) HEREBY CONFIRMS THE INTELLECTUAL PROPERTY AGREEMENT. Recorded Oct 29, 2019
From: NUANCE COMMUNICATIONS, INC.
To: CERENCE OPERATING COMPANY
Reel/Frame 050871/0001 →
INTELLECTUAL PROPERTY AGREEMENT Recorded Oct 23, 2019
From: NUANCE COMMUNICATIONS, INC.
To: CERENCE INC.
Reel/Frame 050836/0191 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Mar 6, 2009
From: INTERNATIONAL BUSINESS MACHINES CORPORATION
To: NUANCE COMMUNICATIONS, INC.
Reel/Frame 022354/0566 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Mar 25, 2003
From: BLASS, OSCAR J.
To: INTERNATIONAL BUSINESS MACHINES CORPORATION
Reel/Frame 013912/0591 →