IP Library › Granted Patent US 12,051,428
Granted Patent B1
US 12,051,428 · App. 17/739,642 · Granted Jul 30, 2024

System and methods for generating realistic waveforms

Inventor: Michael Petrochuk (Seattle, WA)
Assignee: WellSaid Labs, Inc.
G10L19/02G10L15/16G10L25/30G06N3/08
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,051,428
App. No.
17/739,642
Granted
Jul 30, 2024
Kind
B1
Abstract

Systems, apparatuses, and methods directed to generating waveforms that can be used to produce realistic sounding speech when used in a speech synthesis or text-to-speech application. In some embodiments, this includes training a neural network model that is part of a generative adversarial network (GAN) to generate the waveform from an input spectrogram, while implementing a specific set of processing stages to determine loss. The loss is used with a back-propagation process to update the weights of the model during a training cycle. In one example use case, text is converted to a spectrogram and the spectrogram is converted to a waveform by a trained model. The output of the model may be used to drive a transducer that converts the waveform to audible sound.

Claims (52)

1. A method for generating a waveform for use as an input to an audio transducer, comprising:

training a model to operate on an input in the form of a spectrogram to generate an output in the form of a waveform, wherein the training further comprises

inputting a spectrogram obtained from a source into the model to produce an output waveform corresponding to the source;

transforming the output waveform corresponding to the source into a representation of the output waveform in the frequency domain;

obtaining an actual waveform corresponding to the source;

transforming the actual waveform corresponding to the source into a representation of the actual waveform in the frequency domain;

inputting both the representation of the output waveform corresponding to the source in the frequency domain and the representation of the actual waveform corresponding to the source in the frequency domain into a Discriminator to generate a first loss term corresponding to the discriminator loss;

inputting both the representation of the output waveform corresponding to the source in the frequency domain and the representation of the actual waveform corresponding to the source in the frequency domain into a Lp-Norm loss function to generate a second loss term corresponding to the Lp-Norm loss;

combining the first loss term and the second loss term to produce a total loss term; and

using the total loss term to adjust a characteristic of the model;

inputting a spectrogram generated from a different source into the trained model to generate an output waveform corresponding to the different source; and

using the output waveform corresponding to the different source as the input to the audio transducer.

2. The method of claim 1 , wherein the source is a set of text.

3. The method of claim 2 , wherein the spectrogram obtained from the source is obtained by using a text-to spectrogram model.

4. The method of claim 2 , wherein obtaining the actual waveform corresponding to the source further comprises obtaining an audio waveform from a text-to-speech generator for the same set of text as used for the source.

5. The method of claim 4 , wherein transforming the actual waveform corresponding to the source into a representation of the actual waveform in the frequency domain further comprises using a Short-time Fourier Transform to transform the actual waveform.

6. The method of claim 1 , wherein transforming the output waveform corresponding to the source into a representation of the output waveform in the frequency domain further comprises using a Short-time Fourier Transform to transform the output waveform.

7. The method of claim 1 , wherein the Lp-Norm loss function is a L1 loss function or a L2 loss function.

8. The method of claim 1 , wherein combining the first loss term and the second loss term to produce a total loss term further comprises adding the two loss terms.

9. The method of claim 1 , wherein using the total loss term to adjust a characteristic of the model further comprises using the total loss term as part of a feedback or back-propagation mechanism to adjust the characteristic, and wherein the characteristic is a weight that is part of the model.

10. The method of claim 9 , wherein the model is a set of layers, with each layer containing a plurality of nodes, with a plurality of connections between a set of nodes in a first layer and a set of nodes in a second layer adjacent to the first layer, and with a set of weights with each weight in the set of weights associated with one of the plurality of connections.

11. The method of claim 1 , wherein the model is a Generator that is part of a generative adversarial network, and the Discriminator is part of the generative adversarial network.

12. The method of claim 1 , where in the source is one of an image, time-series data, or sampled waveform data and the spectrogram is a tensor representation of the source.

13. A system for generating a waveform for use as an input to an audio transducer, comprising:

one or more electronic processors operable to execute a set of computer-executable instructions; and

the set of computer-executable instructions, wherein when executed, the instructions cause the one or more electronic processors to

train a model to operate on an input in the form of a spectrogram to generate an output in the form of a waveform, wherein the training further comprises

inputting a spectrogram obtained from a source into the model to produce an output waveform corresponding to the source;

transforming the output waveform corresponding to the source into a representation of the output waveform in the frequency domain;

obtaining an actual waveform corresponding to the source;

transforming the actual waveform corresponding to the source into a representation of the actual waveform in the frequency domain;

inputting both the representation of the output waveform corresponding to the source in the frequency domain and the representation of the actual waveform corresponding to the source in the frequency domain into a Discriminator to generate a first loss term corresponding to the discriminator loss;

inputting both the representation of the output waveform corresponding to the source in the frequency domain and the representation of the actual waveform corresponding to the source in the frequency domain into a Lp-Norm loss function to generate a second loss term corresponding to the Lp-Norm loss;

combining the first loss term and the second loss term to produce a total loss term; and

using the total loss term to adjust a characteristic of the model;

input a spectrogram generated from a different source into the trained model to generate an output waveform corresponding to the different source; and

use the output waveform corresponding to the different source as the input to the audio transducer.

14. The system of claim 13 , wherein the source is a set of text, the spectrogram obtained from the source is obtained by using a text-to spectrogram model and obtaining the actual waveform corresponding to the source further comprises obtaining an audio waveform from a text-to-speech generator for the same set of text as used for the source.

15. The system of claim 13 , wherein transforming the output waveform corresponding to the source into a representation of the output waveform in the frequency domain further comprises using a Short-time Fourier Transform to transform the output waveform and transforming the actual waveform corresponding to the source into a representation of the actual waveform in the frequency domain further comprises using a Short-time Fourier Transform to transform the actual waveform.

16. The system of claim 13 , wherein the Lp-Norm loss function is a L1 loss function or a L2 loss function and combining the first loss term and the second loss term to produce a total loss term further comprises adding the two loss terms.

17. The system of claim 13 , wherein using the total loss term to adjust a characteristic of the model further comprises using the total loss term as part of a feedback or back-propagation mechanism to adjust the characteristic, the characteristic is a weight that is part of the model, and the model is a set of layers, with each layer containing a plurality of nodes, with a plurality of connections between a set of nodes in a first layer and a set of nodes in a second layer adjacent to the first layer, and with a set of weights with each weight in the set of weights associated with one of the plurality of connections.

18. A non-transitory computer-readable medium storing a set of computer-executable instructions that when executed by one or more programmed electronic processors, cause the processors to generate a waveform by:

training a model to operate on an input in the form of a spectrogram to generate an output in the form of a waveform, wherein the training further comprises

inputting a spectrogram obtained from a source into the model to produce an output waveform corresponding to the source;

transforming the output waveform corresponding to the source into a representation of the output waveform in the frequency domain;

obtaining an actual waveform corresponding to the source; transforming the actual waveform corresponding to the source into a representation of the actual waveform in the frequency domain;

inputting both the representation of the output waveform corresponding to the source in the frequency domain and the representation of the actual waveform corresponding to the source in the frequency domain into a Discriminator to generate a first loss term corresponding to the discriminator loss;

inputting both the representation of the output waveform corresponding to the source in the frequency domain and the representation of the actual waveform corresponding to the source in the frequency domain into a Lp-Norm loss function to generate a second loss term corresponding to the Lp-Norm loss;

combining the first loss term and the second loss term to produce a total loss term; and

using the total loss term to adjust a characteristic of the model.

19. The computer-readable medium of claim 18 , wherein the source is a set of text, the spectrogram obtained from the source is obtained by using a text-to spectrogram model and obtaining the actual waveform corresponding to the source further comprises obtaining an audio waveform from a text-to-speech generator for the same set of text as used for the source, and the set of instructions cause the processors to receive a spectrogram generated from a different source as an input to the trained model to generate an output waveform corresponding to the different source and use the output waveform corresponding to the different source as an input to an audio transducer.

20. The computer-readable medium of claim 18 , wherein transforming the output waveform corresponding to the source into a representation of the output waveform in the frequency domain further comprises using a Short-time Fourier Transform to transform the output waveform and transforming the actual waveform corresponding to the source into a representation of the actual waveform in the frequency domain further comprises using a Short-time Fourier Transform to transform the actual waveform.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jul 11, 2022
From: PETROCHUK, MICHAEL
To: WELLSAID LABS, INC.
Reel/Frame 060477/0961 →
Continuity (1)
Provisional Application 63186634 · May 10, 2021
Cited By (4)
US 12,394,186 US 12,475,617 US 12,488,779 US 12,614,542