IP Library Granted Patent US 11,741,941
Granted Patent B2
US 11,741,941 · App. 17/341,082 · Granted Aug 29, 2023

Configurable neural speech synthesis

Inventor: Andrew Richards (Toulouse, FR)
Assignee: SoundHound, Inc
G10L13/047G06F3/167G06N3/04G06N3/084G10L13/033G10L13/08G10L15/26G06F3/04847
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,741,941
App. No.
17/341,082
Granted
Aug 29, 2023
Kind
B2
Abstract

A discriminator trained on labeled samples of speech can compute probabilities of voice properties. A speech synthesis generative neural network that takes in text and continuous scale values of voice properties is trained to synthesize speech audio that the discriminator will infer as matching the values of the input voice properties. Voice parameters can include speaker voice parameters, accents, and attitudes, among others. Training can be done by transfer learning from an existing neural speech synthesis model or such a model can be trained with a loss function that considers speech and parameter values. A graphical user interface can allow voice designers for products to synthesize speech with a desired voice or generate a speech synthesis engine with frozen voice parameters. A vector of parameters can be used for comparison to previously registered voices in databases such as ones for trademark registration.

Claims (51)

1. A computerized method of synthesizing speech audio, the computerized method comprising:

receiving a string of text and at least one voice property value with a perceptible meaning, wherein the at least one voice property value comprises a voice property vector;

reading at least one stored voice property vector from a brand database;

computing a distance between the at least one stored voice property vector and the voice property vector to generate a computed distance;

generating code for execution by a computer, the code implementing a neural speech synthesis model, wherein a node in a hidden layer includes, in its summation, a constant term derived from a product of the at least one voice property value and a weight learned from a training process; and

outputting the code, wherein synthesizing speech audio corresponding to the string of text using the neural speech synthesis model that conditions a sound of speech audio on the at least one voice property value generates synthesized speech audio wherein the sound of the synthesized speech audio perceptually relates to the at least one voice property value, and wherein the code implements a speech synthesis function of a speech synthesizer.

2. The computerized method of claim 1 , wherein the at least one voice property value includes at least one of a gender voice property, an age voice property, an accent voice property, a timbre voice property, or an attitude voice property.

3. The computerized method of claim 1 , further comprising:

enabling download of the synthesized speech audio.

4. The computerized method of claim 1 , further comprising:

enabling playback of the synthesized speech audio.

5. The computerized method of claim 1 , wherein the string of text is associated with at least one text tag.

6. The computerized method of claim 1 , wherein the string of text indicates dynamically configurable voice parameter values.

7. The computerized method of claim 1 , further comprising:

providing a graphical user interface that includes one of a text input field or a voice property value input field.

8. The computerized method of claim 1 , wherein the code is in a binary format.

9. The computerized method of claim 1 , further comprising:

determining that the computed distance satisfies a threshold distance; and

generating an error message.

10. The computerized method of claim 1 , further comprising:

determining that the computed distance fails to satisfy a threshold distance; and

storing the at least one voice property value in the brand database.

11. A non-transitory computer readable storage medium storing instructions that, when executed by at least one processor of a computing system, causes the computing system to:

receive a string of text and at least one voice property value with a perceptible meaning, wherein the at least one voice property value comprises a voice property vector;

read at least one stored voice property vector from a brand database;

compute a distance between the at least one stored voice property vector and the voice property vector to generate a computed distance;

generate code for execution by a computer, the code implementing a neural speech synthesis model, wherein a node in a hidden layer includes, in its summation, a constant term derived from a product of the at least one voice property value and a weight learned from a training process; and

output the code, wherein synthesizing speech audio corresponding to the string of text using the neural speech synthesis model that conditions a sound of speech audio on the at least one voice property value generates synthesized speech audio wherein the sound of the synthesized speech audio perceptually relates to the at least one voice property value, and wherein the code implements a speech synthesis function of a speech synthesizer.

12. The non-transitory computer readable storage medium of claim 11 , wherein the at least one voice property value includes at least one of a gender voice property, an age voice property, an accent voice property, a timbre voice property, or an attitude voice property.

13. The non-transitory computer readable storage medium of claim 11 , wherein the instructions, when executed by the at least one processor, further enables the computing system to:

enable download of the synthesized speech audio.

14. The non-transitory computer readable storage medium of claim 11 , wherein the instructions, when executed by the at least one processor, further enables the computing system to:

enable playback of the synthesized speech audio.

15. The non-transitory computer readable storage medium of claim 11 , wherein the string of text indicates dynamically configurable voice parameter values.

16. The non-transitory computer readable storage medium of claim 11 , wherein the instructions, when executed by the at least one processor, further enables the computing system to:

provide a graphical user interface that includes one of a text input field or a voice property value input field.

17. The non-transitory computer readable storage medium of claim 11 , wherein wherein the code is in a binary format.

18. The non-transitory computer readable storage medium of claim 11 , wherein the instructions, when executed by the at least one processor, further enables the computing system to:

determine that the computed distance satisfies a threshold distance; and

generate an error message.

19. The non-transitory computer readable storage medium of claim 11 , wherein the instructions, when executed by the at least one processor, further enables the computing system to:

determine that the computed distance fails to satisfy a threshold distance; and

store the at least one voice property value in the brand database.

20. A computing system for synthesizing speech audio, comprising:

a processor; and

a memory device including instructions that, when executed by the processor, enables the computing system to:

receive a string of text and at least one voice property value with a perceptible meaning, wherein the at least one voice property value comprises a voice property vector,

read at least one stored voice property vector from a brand database,

compute a distance between the at least one stored voice property vector and the voice property vector to generate a computed distance,

generate code for execution by a computer, the code implementing a neural speech synthesis model, wherein a node in a hidden layer includes, in its summation, a constant term derived from a product of the at least one voice property value and a weight learned from a training process, and

output the code, wherein synthesizing speech audio corresponding to the string of text using the neural speech synthesis model that conditions a sound of speech audio on the at least one voice property value generates synthesized speech audio wherein the sound of the synthesized speech audio perceptually relates to the at least one voice property value, and wherein the code implements a speech synthesis function of a speech synthesizer.

Assignments (7)
TERMINATION AND RELEASE OF SECURITY INTEREST IN PATENTS Recorded Dec 3, 2024
From: MONROE CAPITAL MANAGEMENT ADVISORS, LLC, AS COLLATERAL AGENT
To: SOUNDHOUND, INC.
Reel/Frame 069480/0312 →
SECURITY INTEREST Recorded Aug 9, 2024
From: SOUNDHOUND, INC.
To: MONROE CAPITAL MANAGEMENT ADVISORS, LLC, AS COLLATERAL AGENT
Reel/Frame 068526/0413 →
RELEASE OF SECURITY INTEREST Recorded Jun 11, 2024
From: ACP POST OAK CREDIT II LLC, AS COLLATERAL AGENT
To: SOUNDHOUND, INC.; SOUNDHOUND AI IP, LLC
Reel/Frame 067698/0845 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jun 27, 2023
From: SOUNDHOUND AI IP HOLDING, LLC
To: SOUNDHOUND AI IP, LLC
Reel/Frame 064205/0676 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jun 23, 2023
From: SOUNDHOUND, INC.
To: SOUNDHOUND AI IP HOLDING, LLC
Reel/Frame 064083/0484 →
SECURITY INTEREST Recorded Apr 17, 2023
From: SOUNDHOUND, INC.; SOUNDHOUND AI IP, LLC
To: ACP POST OAK CREDIT II LLC
Reel/Frame 063349/0355 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jun 9, 2021
From: RICHARDS, ANDREW
To: SOUNDHOUND, INC.
Reel/Frame 056490/0016 →
Cited By (2)
US 12,387,619 US 12,475,878