System and method for automatic alignment of phonetic content for real-time accent conversion
The disclosed technology relates to methods, accent conversion systems, and non-transitory computer readable media for real-time accent conversion. In some examples, a set of phonetic embedding vectors is obtained for phonetic content representing a source accent and obtained from input audio data. A trained machine learning model is applied to the set of phonetic embedding vectors to generate a set of transformed phonetic embedding vectors corresponding to phonetic characteristics of speech data in a target accent. An alignment is determined by maximizing a cosine distance between the set of phonetic embedding vectors and the set of transformed phonetic embedding vectors. The speech data is then aligned to the phonetic content based on the determined alignment to generate output audio data representing the target accent. The disclosed technology transforms phonetic characteristics of a source accent to match the target accent more closely for efficient and seamless accent conversion in real-time applications.
1 . A system, comprising an audio interface, a communication interface, memory having instructions stored thereon, and one or more processors coupled to the memory and configured to execute the instructions to:
receive output audio data representing a target accent and comprising speech data aligned to phonetic content representing a source accent based on a differentiable alignment determined based on first and second phonetic embedding vectors, wherein:
the first phonetic embedding vectors are generated from input audio data and are for the phonetic content; and
the second phonetic embedding vectors are generated based on an application of a trained neural network to the first phonetic embedding vectors and correspond to first phonetic characteristics of the speech data in the target accent;
store the output audio data in the memory; and
output the output audio data from the memory and via the audio interface.
2 . The system of claim 1 , wherein the first phonetic embedding vectors represent second phonetic characteristics of input speech in the input audio data in a numerical format.
3 . The system of claim 1 , wherein the trained neural network comprises an encoder layer configured to encode the first phonetic embedding vectors into a latent representation and a decoder layer configured to decode the latent representation to generate the second phonetic embedding vectors.
4 . The system of claim 1 , wherein the differentiable alignment is determined based a distance between the first phonetic embedding vectors and the second phonetic embedding vectors.
5 . The system of claim 4 , wherein the distance is a cosine distance determined based on a generated dot product of a normalization of the first and second phonetic embedding vectors based on a scaling of the first and second phonetic embedding vectors to have a magnitude of one and a preservation of a relative direction of the first and second phonetic embedding vectors.
6 . The system of claim 5 , wherein the cosine distance is optimized based on an application of a gradient-based optimization algorithm.
7 . The system of claim 1 , wherein the trained neural network is trained to learn a mapping between the first phonetic embedding vectors and the second phonetic embedding vectors using a labeled dataset comprising paired samples of source accent phonetic embedding vectors and corresponding target accent phonetic embedding vectors.
8 . One or more non-transitory computer-readable media having stored thereon instructions comprising executable code that, when executed by one or more systems, is configured to cause the one or more systems to:
receive output audio data representing a target accent and comprising speech data aligned to phonetic content representing a source accent based on a differentiable alignment determined based on first and second phonetic embedding vectors, wherein:
the first phonetic embedding vectors are generated from input audio data and are for the phonetic content; and
the second phonetic embedding vectors are generated based on an application of a trained neural network to the first phonetic embedding vectors and correspond to first phonetic characteristics of the speech data in the target accent; and
output the output audio data via an audio interface.
9 . The one or more non-transitory computer-readable media of claim 8 , wherein the first phonetic embedding vectors encode one or more of phonetic features, patterns, phonemes, pronunciation, intonation, speech sounds, or phonetic units present in input speech in the input audio data.
10 . The one or more non-transitory computer-readable media of claim 8 , wherein the output audio data is further generated based on an alignment of first frames of the speech data with corresponding second frames of the phonetic content.
11 . The one or more non-transitory computer-readable media of claim 8 , wherein the output audio data is further generated based on an application of one or more techniques comprising prosody modeling, intonation adjustment, or accent-specific acoustic modeling.
12 . The one or more non-transitory computer-readable media of claim 8 , wherein the differentiable alignment is determined based on a distance between the first phonetic embedding vectors and the second phonetic embedding vectors.
13 . The one or more non-transitory computer-readable media of claim 8 , wherein the output audio data preserves linguistic content of the input audio data.
14 . A method, comprising:
receiving output audio data representing a target accent and comprising speech data aligned to phonetic content representing a source accent based on a differentiable alignment determined based on first and second phonetic embedding vectors, wherein:
the first phonetic embedding vectors are generated from input audio data and are for the phonetic content; and
the second phonetic embedding vectors are generated based on an application of a trained neural network to the first phonetic embedding vectors and correspond to first phonetic characteristics of the speech data in the target accent; and
outputting the output audio data via an audio interface, wherein the output audio data represents an accent-converted version of the input audio data.
15 . The method of claim 14 , wherein the first and second phonetic embedding vectors are pre-processed based on an application of one or more dimensionality reduction techniques.
16 . The method of claim 14 , wherein the first phonetic embedding vectors represent second phonetic characteristics of input speech in the input audio data in a numerical format.
17 . The method of claim 14 , wherein the trained neural network comprises an encoder layer configured to encode the first phonetic embedding vectors into a latent representation and a decoder layer configured to decode the latent representation to generate the second phonetic embedding vectors.
18 . The method of claim 14 , wherein the differentiable alignment is determined based on a cosine distance between the first phonetic embedding vectors and the second phonetic embedding vectors.
19 . The method of claim 14 , wherein the trained neural network is trained to learn a mapping between the first phonetic embedding vectors and the second phonetic embedding vectors using a labeled dataset comprising paired samples of source accent phonetic embedding vectors and corresponding target accent phonetic embedding vectors.
20 . The method of claim 14 , wherein the output audio data preserves linguistic content of the input audio data.