IP Library › Patent Application 19206502
Patent Application
App. No. 19/206,502

SYSTEM AND METHOD FOR STYLE EXTRACTION IN SPEECH SYNTHESIS USING NEURAL NETWORKS

Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US None
App. No.
19/206,502
Abstract

The disclosed technology relates to methods, speech processing systems, and non-transitory computer readable media for style extraction in speech synthesis. In some examples, one or more content elements and one or more non-content elements are extracted from input audio data obtained via an audio interface and corresponding to input speech. The one or more non-content elements comprise style elements comprising at least an input pitch. A trained autoencoder is applied to encode the input pitch in a latent representation comprising a low-dimensional vector and combine the one or more content elements and the one or more non-content elements based on the low-dimensional vector to generate a new representation of the input speech. Output audio data is then generated and provided based on the new representation of the input speech. The output audio data comprises a pitch-consistent reconstruction of the input speech.

Claims (33)

1 . A speech processing system, comprising memory having instructions stored thereon, and one or more processors coupled to the memory and configured to execute the instructions to:

using augmentations to separate input audio data into one or more content elements and one or more non-content elements, wherein the augmentations simulate degraded speech characteristics;

apply a trained autoencoder to encode input pitch of the non-content elements in a latent representation comprising a low-dimensional vector, wherein the latent representation excludes the content elements; and

provide output audio data based on a new representation of the input speech generated by combing the content elements and the non-content elements based on the low-dimensional vector.

2 . The speech processing system of claim 1 , wherein the augmentations comprise one or more of background noise, air conditioner sounds, fan sounds, environmental sounds masked data, microphone pops, smooth speech, or convolving speech.

3 . The speech processing system of claim 1 , wherein the input audio data comprises a plurality of frames and the processors are further configured to execute the instructions to convert the low-dimensional vector to a low-dimensional representation of an output pitch on a frame-by-frame basis in accordance with the frames of the input audio data.

4 . The speech processing system of claim 3 , wherein the processors are further configured to execute the instructions to apply the trained autoencoder to combine the content elements and the non-content elements further based on the low-dimensional representation of the output pitch.

5 . The speech processing system of claim 1 , wherein the processors are further configured to execute the instructions to apply one or more automated speech recognition techniques to extract the content elements.

6 . The speech processing system of claim 1 , wherein the excluded content elements comprise audio related to the content of what is being said by a speaker of the input speech.

7 . The speech processing system of claim 1 , wherein the processors are further configured to execute the instructions to apply an autoencoding deconstruction-reconstruction technique to the content elements before generating the output audio data.

8 . A method implemented by a speech processing system and comprising:

extracting, from input audio data, one or more content elements and one or more non-content elements, wherein the non-content elements comprise at least an input pitch;

applying a neural network to encode the input pitch extracted from the input audio data without the content elements in a low-dimensional vector, wherein the excluded content elements comprise audio related to the content of what is being said by a speaker of input speech of the input audio data, wherein the content of what is being said comprises phonemes;

combining the content elements and the non-content elements based on the low-dimensional vector to generate a new representation of the input speech; and

providing output audio data generated based on the new representation of the input speech and comprising a pitch-consistent reconstruction of the input speech.

9 . The method of claim 8 , further comprising using augmentations to separate the input speech into the content elements and the non-content elements, wherein the augmentations simulate degraded speech characteristics.

10 . The method of claim 8 , wherein the input audio data comprises a plurality of frames and the method further comprises converting the low-dimensional vector to a low-dimensional representation of an output pitch on a frame-by-frame basis in accordance with the frames of the input audio data.

11 . The method of claim 10 , further comprising applying the neural network to combine the content elements and the non-content elements further based on the low-dimensional representation of the output pitch.

12 . The method of claim 8 , further comprising applying one or more automated speech recognition techniques to extract the content elements.

13 . The method of claim 9 , wherein the augmentations comprise one or more of background noise, air conditioner sounds, fan sounds, environmental sounds masked data, microphone pops, smooth speech, or convolving speech.

14 . The method of claim 8 , further comprising applying an autoencoding deconstruction-reconstruction technique to the content elements before generating the output audio data.

15 . A non-transitory computer-readable medium comprising instructions that, when executed by at least one processor, cause the processor to:

extract, from input audio data corresponding to input speech, one or more content elements and one or more non-content elements comprising using augmentations to facilitate separation of the input speech into the content elements and the non-content elements, wherein the non-content elements comprise at least an input pitch;

apply a neural network to encode the input pitch in a latent representation, wherein the latent representation excludes the content elements and the excluded content elements comprise audio related to the content of what is being said by a speaker of the input speech;

combine the content elements and the non-content elements based on the latent representation to generate a new representation of the input speech; and

provide output audio data generated based on the new representation of the input speech.

16 . The non-transitory computer-readable medium of claim 14 , wherein the latent representation comprises a low-dimensional vector and represents the input pitch in a compressed form.

17 . The non-transitory computer-readable medium of claim 14 , wherein the augmentations simulate degraded speech characteristics and comprise one or more of background noise, masked data, microphone pops, smooth speech, or convolving speech.

18 . The non-transitory computer-readable medium of claim 14 , wherein the input audio data comprises a plurality of frames and the instructions, when executed by the processor further cause the processor to:

convert the latent representation to a low-dimensional representation of an output pitch on a frame-by-frame basis in accordance with the frames of the input audio data; and

apply the neural network to combine the content elements and the non-content elements further based on the low-dimensional representation of the output pitch.

19 . The non-transitory computer-readable medium of claim 14 , wherein the instructions, when executed by the processor further cause the processor to apply one or more automated speech recognition techniques to extract the content elements.

20 . The non-transitory computer-readable medium of claim 14 , wherein the instructions, when executed by the processor further cause the processor to apply an autoencoding deconstruction-reconstruction technique to the content elements before generating the output audio data.

Assignments (1)
SECURITY INTEREST Recorded Apr 17, 2026
From: SANAS.AI INC.
To: CANADIAN IMPERIAL BANK OF COMMERCE
Reel/Frame 074405/0229 →