IP Library Granted Patent US 9,495,954
Granted Patent B2
US 9,495,954 · App. 15/049,592 · Granted Nov 15, 2016

System and method of synthetic voice generation and modification

View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 9,495,954
App. No.
15/049,592
Granted
Nov 15, 2016
Kind
B2
Abstract

Disclosed herein are systems, methods, and non-transitory computer-readable storage media for generating a synthetic voice. A system configured to practice the method combines a first database of a first text-to-speech voice and a second database of a second text-to-speech voice to generate a combined database, selects from the combined database, based on a policy, voice units of a phonetic category for the synthetic voice to yield selected voice units, and synthesizes speech based on the selected voice units. The system can synthesize speech without parameterizing the first text-to-speech voice and the second text-to-speech voice. A policy can define, for a particular phonetic category, from which text-to-speech voice to select voice units. The combined database can include multiple text-to-speech voices from different speakers. The combined database can include voices of a single speaker speaking in different styles. The combined database can include voices of different languages.

Claims (37)

1. A method comprising:

storing, in a database, voice data according to user emotions;

identifying, from user speech, a user emotion;

identifying, via at least one processor and according to the user emotion, a first portion of the voice data, wherein the first portion of the voice data comprises a first emotional content for a first speaker;

identifying, via the at least one processor and according to the user emotion, a second portion of the voice data, wherein the second portion of the voice data comprises a second emotional content for a second speaker; and

synthesizing synthesized speech using the first portion of the voice data and the second portion of the voice data.

2. The method of claim 1 , wherein the first portion of the voice data comprises first text-to-speech voice data and wherein the second portion of the voice data comprises second text-to-speech voice data.

3. The method of claim 1 , wherein the synthesized speech corresponds to the user emotion.

4. The method of claim 1 , further comprising generating a plurality of synthetic voices from the database, wherein each synthetic voice in the plurality of synthetic voices is generated according to a respective selection policy.

5. The method of claim 1 , wherein a selection policy defines, for a particular phonetic category, which voice to use for synthesizing the synthesized speech.

6. The method of claim 1 , wherein the first portion of the voice data and the second portion of the voice data have a similar pitch range and fundamental frequency with respect to one another.

7. The method of claim 3 , wherein the synthesized speech is synthesized using selected voice units from the database, wherein the selected voice units comprise a first voice unit from the first portion of the voice data, and comprise a second voice unit from the second portion of the voice data.

8. The method of claim 7 , wherein the synthesized speech is synthesized without parameterizing the first portion of the voice data and without parameterizing the second portion of the voice data.

9. A system comprising:

a processor; and

a computer-readable storage medium having instructions stored which, when executed by the processor, result in the processor performing operations comprising:

storing, in a database, voice data according to user emotions;

identifying, from user speech, a user emotion;

identifying, via at least one processor and according to the user emotion, a first portion of the voice data, wherein the first portion of the voice data comprises a first emotional content for a first speaker;

identifying, via the at least one processor and according to the user emotion, a second portion of the voice data, wherein the second portion of the voice data comprises a second emotional content for a second speaker; and

synthesizing synthesized speech using the first portion of the voice data and the second portion of the voice data.

10. The system of claim 9 , wherein the first portion of the voice data comprises first text-to-speech voice data and wherein the second portion of the voice data comprises second text-to-speech voice data.

11. The system of claim 9 , wherein the synthesized speech corresponds to the user emotion.

12. The system of claim 9 , the computer-readable storage medium having additional instructions stored which, when executed by the processor, result in operations comprising generating a plurality of synthetic voices from the database, wherein each synthetic voice in the plurality of synthetic voices is generated according to a respective selection policy.

13. The system of claim 9 , wherein a selection policy defines, for a particular phonetic category, which voice to use for synthesizing the synthesized speech.

14. The system of claim 11 , wherein the synthesized speech is synthesized using selected voice units from the database, wherein the selected voice units comprise a first voice unit from the first portion of the voice data, and comprise a second voice unit from the second portion of the voice data.

15. The system of claim 14 , wherein the synthesized speech is synthesized without parameterizing the first portion of the voice data and without parameterizing the second portion of the voice data.

16. A device having instructions stored which, when executed by a processor, result in the processor performing operations comprising:

storing, in a database, voice data according to user emotions;

identifying, from user speech, a user emotion;

identifying, via at least one processor and according to the user emotion, a first portion of the voice data, wherein the first portion of the voice data comprises a first emotional content for a first speaker;

identifying, via the at least one processor and according to the user emotion, a second portion of the voice data, wherein the second portion of the voice data comprises a second emotional content for a second speaker; and

synthesizing synthesized speech using the first portion of the voice data and the second portion of the voice data.

17. The device of claim 16 , wherein the first portion of the voice data comprises first text-to-speech voice data and wherein the second portion of the voice data comprises second text-to-speech voice data.

18. The device of claim 16 , wherein the synthesized speech using the first portion of the voice data and the second portion of the voice data yields synthesized speech that corresponds to the user emotion.

19. The device of claim 16 , wherein the synthesizing of the speech is performed using selected voice units from the database, wherein the selected voice units comprise a first voice unit from the first portion of the voice data, and comprise a second voice unit from the second portion of the voice data.

20. The device of claim 19 , wherein the synthesized speech is synthesized without parameterizing the first portion of the voice data and without parameterizing the second portion of the voice data.

Assignments (8)
RELEASE (REEL 052935 / FRAME 0584) Recorded Jan 2, 2025
From: WELLS FARGO BANK, NATIONAL ASSOCIATION
To: CERENCE OPERATING COMPANY
Reel/Frame 069797/0818 →
CORRECTIVE ASSIGNMENT TO CORRECT THE REPLACE THE CONVEYANCE DOCUMENT WITH THE NEW ASSIGNMENT PREVIOUSLY RECORDED AT REEL: 050836 FRAME: 0191. ASSIGNOR(S) HEREBY CONFIRMS THE ASSIGNMENT. Recorded Apr 19, 2022
From: NUANCE COMMUNICATIONS, INC.
To: CERENCE OPERATING COMPANY
Reel/Frame 059804/0186 →
SECURITY AGREEMENT Recorded Jun 15, 2020
From: CERENCE OPERATING COMPANY
To: WELLS FARGO BANK, N.A.
Reel/Frame 052935/0584 →
RELEASE OF SECURITY INTEREST Recorded Jun 12, 2020
From: BARCLAYS BANK PLC
To: CERENCE OPERATING COMPANY
Reel/Frame 052927/0335 →
SECURITY AGREEMENT Recorded Nov 7, 2019
From: CERENCE OPERATING COMPANY
To: BARCLAYS BANK PLC
Reel/Frame 050953/0133 →
CORRECTIVE ASSIGNMENT TO CORRECT THE ASSIGNEE NAME PREVIOUSLY RECORDED AT REEL: 050836 FRAME: 0191. ASSIGNOR(S) HEREBY CONFIRMS THE INTELLECTUAL PROPERTY AGREEMENT. Recorded Oct 29, 2019
From: NUANCE COMMUNICATIONS, INC.
To: CERENCE OPERATING COMPANY
Reel/Frame 050871/0001 →
INTELLECTUAL PROPERTY AGREEMENT Recorded Oct 23, 2019
From: NUANCE COMMUNICATIONS, INC.
To: CERENCE INC.
Reel/Frame 050836/0191 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jan 26, 2017
From: AT&T INTELLECTUAL PROPERTY I, L.P.
To: NUANCE COMMUNICATIONS, INC.
Reel/Frame 041504/0952 →