IP Library › Patent Application 19090292
Patent Application
App. No. 19/090,292

Conformer-based Speech Conversion Model

Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US None
App. No.
19/090,292
Abstract

A method for speech conversion includes receiving, as input to an encoder of a speech conversion model, an input spectrogram corresponding to an utterance, the encoder including a stack of self-attention blocks. The method further includes generating, as output from the encoder, an encoded spectrogram and receiving, as input to a spectrogram decoder of the speech conversion model, the encoded spectrogram generated as output from the encoder. The method further includes generating, as output from the spectrogram decoder, an output spectrogram corresponding to a synthesized speech representation of the utterance.

Claims (38)

1 . A computer-implemented method that when executed on data processing hardware causes the data processing hardware to perform operations comprising:

receiving audio data corresponding to an utterance, the input audio data comprising a sequence of acoustic frames each having a length equal to a first value;

processing, by an encoder comprising a stack of self-attention blocks, the sequence of acoustic frames to output first hidden representations each having a length equal to a second value greater than the first value;

after the first hidden representations are output from the encoder, processing the first hidden representations to provide second hidden representations each having a length equal to a third value greater than the first value; and

processing, using a word piece decoder, the second hidden representations to output a textual representation that corresponds to a transcription of the utterance,

2 . The computer-implemented method of claim 1 , wherein the first value comprises 10 milliseconds.

3 . The computer-implemented method of claim 1 , wherein the stack of self-attention blocks of the encoder comprises a stack of Conformer blocks.

4 . The computer-implemented method of claim 3 , wherein each Conformer block comprises a multi-headed self-attention mechanism.

5 . The computer-implemented method of claim 3 , wherein each Conformer block comprises:

a first half feed-forward layer;

a second half feed-forward layer;

a multi-head self-attention block;

and a convolution layer disposed between the first and second half feed-forward layers.

6 . The computer-implemented method of claim 3 , wherein the encoder further comprises a subsampling layer disposed before the stack of Conformer blocks.

7 . The computer-implemented method of claim 6 , wherein the subsampling layer comprises Convolution Neural Network (CNN) layers, followed by pooling in time to reduce a number of the sequence of acoustic frames being processed by an initial Conformer block in the stack of Conformer blocks.

8 . The computer-implemented method of claim 1 , wherein the stack of self-attention blocks of the encoder comprises a stack of Transformer blocks.

9 . The computer-implemented method of claim 1 , wherein audio data corresponding to the utterance is extracted from input speech spoken by a speaker associated with atypical speech.

10 . The computer-implemented method of claim 9 , wherein the operations further comprise processing, using a spectrogram decoder, the second hidden representations to generate an output spectrogram corresponding to a synthesized speech representation of the utterance comprising a synthesized canonical fluent speech representation of the utterance in a voice of the speaker that spoke the input speech.

11 . A system comprising:

data processing hardware; and

memory hardware in communication with the data processing hardware and storing instructions that when executed on the data processing hardware causes the data processing hardware to perform operations comprising:

receiving audio data corresponding to an utterance, the input audio data comprising a sequence of acoustic frames each having a length equal to a first value;

processing, by an encoder comprising a stack of self-attention blocks, the sequence of acoustic frames to output first hidden representations each having a length equal to a second value greater than the first value;

after the first hidden representations are output from the encoder, processing the first hidden representations to provide second hidden representations each having a length equal to a third value greater than the first value; and

processing, using a word piece decoder, the second hidden representations to output a textual representation that corresponds to a transcription of the utterance,

12 . The system of claim 11 , wherein the first value comprises 10 milliseconds.

13 . The system of claim 11 , wherein the stack of self-attention blocks of the encoder comprises a stack of Conformer blocks.

14 . The system of claim 13 , wherein each Conformer block comprises a multi-headed self-attention mechanism.

15 . The system of claim 13 , wherein each Conformer block comprises:

a first half feed-forward layer;

a second half feed-forward layer;

a multi-head self-attention block;

and a convolution layer disposed between the first and second half feed-forward layers.

16 . The system of claim 13 , wherein the encoder further comprises a subsampling layer disposed before the stack of Conformer blocks.

17 . The system of claim 16 , wherein the subsampling layer comprises Convolution Neural Network (CNN) layers, followed by pooling in time to reduce a number of the sequence of acoustic frames being processed by an initial Conformer block in the stack of Conformer blocks.

18 . The system of claim 11 , wherein the stack of self-attention blocks of the encoder comprises a stack of Transformer blocks.

19 . The system of claim 11 , wherein audio data corresponding to the utterance is extracted from input speech spoken by a speaker associated with atypical speech.

20 . The system of claim 19 , wherein the operations further comprise processing, using a spectrogram decoder, the second hidden representations to generate an output spectrogram corresponding to a synthesized speech representation of the utterance comprising a synthesized canonical fluent speech representation of the utterance in a voice of the speaker that spoke the input speech.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Mar 25, 2025
From: RAMABHADRAN, BHUVANA; CHEN, ZHEHUAI; BIADSY, FADI; MENGIBAR, PEDRO J. MORENO
To: GOOGLE LLC
Reel/Frame 070626/0257 →