IP Library › Granted Patent US 12,562,154
Granted Patent B2
US 12,562,154 · App. 18/184,630 · Granted Feb 24, 2026

Scalable model specialization framework for speech model personalization

Inventors: Fadi Biadsy (Sandyston, NJ); Youzheng Chen (Mountain View, CA); Xia Zhang (Mountain View, CA); Oleg Rybakov (Mountain View, CA); Andrew M. Rosenberg (Brooklyn, NY); Pedro J. Moreno Mengibar (Jersey City, NJ)
Assignee: Google LLC
G10L15/16G10L15/02G10L15/063G10L2015/025
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,562,154
App. No.
18/184,630
Granted
Feb 24, 2026
Kind
B2
Abstract

A method for speech conversion includes obtaining a speech conversion model configured to convert input utterances of human speech directly into corresponding output utterances of synthesized speech. The method further includes receiving a speech conversion request including input audio data corresponding to an utterance spoken by a target speaker associated with atypical speech and a speaker identifier uniquely identifying the target speaker. The method includes activating, using the speaker identifier, a particular sub-model for biasing the speech conversion model to recognize a type of the atypical speech associated with the target speaker identified by the speaker identifier. The method includes converting, using the speech conversion model biased by the activated particular sub-model, the input audio data corresponding to the utterance spoken by the target speaker associated with atypical speech into output audio data corresponding to a synthesized canonical fluent speech representation of the utterance spoken by the target speaker.

Claims (54)

1 . A computer-implemented method when executed on data processing hardware causes the data processing hardware to perform operations comprising:

obtaining a speech conversion model configured to convert input utterances of human speech directly into corresponding output utterances of synthesized speech, the speech conversion model comprising an encoder and a decoder;

receiving a speech conversion request comprising input audio data corresponding to an utterance spoken by a target speaker associated with atypical speech and a speaker identifier uniquely identifying the target speaker;

activating, using the speaker identifier, a particular sub-model for biasing the speech conversion model to recognize a type of the atypical speech associated with the target speaker identified by the speaker identifier; and

converting, using the speech conversion model biased by the activated particular sub-model, the input audio data corresponding to the utterance spoken by the target speaker associated with atypical speech into output audio data corresponding to a synthesized canonical fluent speech representation of the utterance spoken by the target speaker by:

generating, as output from the encoder configured to receive the input audio data as input, encoded audio data, the encoded audio data including a series of vectors; and

generating, as output from the decoder configured to receive the encoded audio data output from the encoder as input, the output audio data corresponding to the synthesized canonical fluent representation of the utterance spoken by the target speaker.

2 . The computer-implemented method of claim 1 , wherein the speech conversion model is:

trained on generalized training data; and

speaker- and domain-independent.

3 . The computer-implemented method of claim 1 , wherein generating the output audio data corresponding to the synthesized canonical fluent representation of the utterance spoken by the target speaker comprises generating the output audio data corresponding to the synthesized canonical fluent representation of the utterance spoken by the target speaker without performing any intermediate text-to-speech conversion on a textual representation corresponding to a transcription of the utterance.

4 . The computer-implemented method of claim 3 , wherein the encoder comprises a stack of self-attention blocks each having a multi-headed self attention mechanism.

5 . The computer-implemented method of claim 4 , wherein the particular sub-model comprises a stack of residual adaptors disposed between each of the self-attention blocks in the stack of self-attention blocks of the encoder.

6 . The computer-implemented method of claim 5 , wherein each residual adaptor comprises a normalization layer, followed by a feed-forward layer with down-projection to a bottleneck dimension and a non-linear activation, and another feed-forward layer with up-projection.

7 . The computer-implemented method of claim 3 , wherein the speech conversion model further comprises a wordpiece decoder configured to:

receive, as input, the encoded audio data from the encoder; and

generate, as output, a textual representation corresponding to a transcription of the utterance.

8 . The computer-implemented method of claim 3 , wherein the speech conversion model further comprises a phoneme decoder configured to:

receive, as input, the encoded audio data from the encoder; and

generate, as output, a phoneme representation of the utterance.

9 . The computer-implemented method of claim 1 , wherein:

the input audio data comprises one of an input spectrogram or an input audio waveform; and

the output audio data comprises one of an output spectrogram or an output audio waveform.

10 . The computer-implemented method of claim 1 , wherein activating the particular sub-model for biasing the speech conversion model comprises:

selecting, from among a plurality of sub-models each associated with a different type of atypical speech, the particular sub-model associated with the type of atypical speech associated with the target speaker; and

loading the particular sub-model into the speech conversion model for biasing the speech conversion model to recognize the type of the atypical speech associated with the target speaker.

11 . A system comprising:

data processing hardware; and

memory hardware in communication with the data processing hardware, the memory hardware storing instructions that when executed on the data processing hardware cause the data processing hardware to perform operations comprising:

obtaining a speech conversion model configured to convert input utterances of human speech directly into corresponding output utterances of synthesized speech, the speech conversion model comprising an encoder and a decoder;

receiving a speech conversion request comprising input audio data corresponding to an utterance spoken by a target speaker associated with atypical speech and a speaker identifier uniquely identifying the target speaker;

activating, using the speaker identifier, a particular sub-model for biasing the speech conversion model to recognize a type of the atypical speech associated with the target speaker identified by the speaker identifier; and

converting, using the speech conversion model biased by the activated particular sub-model, the input audio data corresponding to the utterance spoken by the target speaker associated with atypical speech into output audio data corresponding to a synthesized canonical fluent speech representation of the utterance spoken by the target speaker by:

generating, as output from the encoder configured to receive the input audio data as input, encoded audio data, the encoded audio data including a series of vectors; and

generating, as output from the decoder configured to receive the encoded audio data output from the encoder as input, the output audio data corresponding to the synthesized canonical fluent representation of the utterance spoken by the target speaker.

12 . The system of claim 11 , wherein the speech conversion model is:

trained on generalized training data; and

speaker- and domain-independent.

13 . The system of claim 11 , wherein generating the output audio data corresponding to the synthesized canonical fluent representation of the utterance spoken by the target speaker comprises generating the output audio data corresponding to the synthesized canonical fluent representation of the utterance spoken by the target speaker without performing any intermediate text-to-speech conversion on a textual representation corresponding to a transcription of the utterance.

14 . The system of claim 13 , wherein the encoder comprises a stack of self-attention blocks each having a multi-headed self attention mechanism.

15 . The system of claim 14 , wherein the particular sub-model comprises a stack of residual adaptors disposed between each of the self-attention blocks in the stack of self-attention blocks of the encoder.

16 . The system of claim 15 , wherein each residual adaptor comprises a normalization layer, followed by a feed-forward layer with down-projection to a bottleneck dimension and a non-linear activation, and another feed-forward layer with up-projection.

17 . The system of claim 13 , wherein the speech conversion model further comprises a wordpiece decoder configured to:

receive, as input, the encoded audio data from the encoder; and

generate, as output, a textual representation corresponding to a transcription of the utterance.

18 . The system of claim 13 , wherein the speech conversion model further comprises a phoneme decoder configured to:

receive, as input, the encoded audio data from the encoder; and

generate, as output, a phoneme representation of the utterance.

19 . The system of claim 11 , wherein:

the input audio data comprises one of an input spectrogram or an input audio waveform; and

the output audio data comprises one of an output spectrogram or an output audio waveform.

20 . The system of claim 11 , wherein activating the particular sub-model for biasing the speech conversion model comprises:

selecting, from among a plurality of sub-models each associated with a different type of atypical speech, the particular sub-model associated with the type of atypical speech associated with the target speaker; and

loading the particular sub-model into the speech conversion model for biasing the speech conversion model to recognize the type of the atypical speech associated with the target speaker.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Mar 21, 2023
From: BIADSY, FADI; CHEN, YOUZHENG; ZHANG, XIA; RYBAKOV, OLEG; ROSENBERG, ANDREW M.; MENGIBAR, PEDRO J. MORENO
To: GOOGLE LLC
Reel/Frame 063047/0941 →
Continuity (2)
Provisional Application 63269611 · Mar 18, 2022
Related Publication 20230298574A1 · Sep 21, 2023
References Cited (22)
US 12249336B2 · Li · 2025 [cited by examiner]
US 12367860B2 · Ganong, III · 2025 [cited by examiner]
US 20200380215A1 · Kannan · 2020 [cited by examiner]
US 20210209315A1 · Jia · 2021 [cited by examiner]
US 20230169954A1 · Thomas · 2023 [cited by examiner]
JP 2019528476A · 2019 [cited by applicant]
JP 2021033048A · 2021 [cited by applicant]
WO 2021154563A1 · 2021 [cited by applicant]
WO 2022046526A1 · 2022 [cited by applicant]
Houlsby et al. “Parameter-Efficient Transfer Learning for NLP” Jun. 13, 2019, Proceedings of the 36th International Conference on Machine Learning (Year: 2019). [cited by examiner]
Irie et al. “On the Choice of Modeling Unit for Sequence-to-Sequence Speech Recognition” Jul. 23, 2019, Human Language Technology and Pattern Recognition Group (Year: 2019). [cited by examiner]
Bapna et al. “Simple, Scalable Adaptation for Neural Machine Translation” Sep. 18, 2019 (Year: 2019). [cited by examiner]
Michel et al. “Extreme Adaptation for Personalized Neural Machine Translation” Language Technologies Institute Carnegie Mellon University May 4, 2018 (Year: 2018). [cited by examiner]
International Search Report and Written Opinion for the related Application No. PCT/US2023/064492 dated May 22, 2023, 144 pages. [cited by applicant]
Doshi Rohan et al: “Extending Parrotron: An End-to-End, Speech Conversion and Speech Recognition Model for Atypical Speech”, ICASSP 2021—2021 IEEE International Conference on Acoustics, Speech and Signal Processing (ICA… [cited by applicant]
Fadi Biadsy et al: “Parrotron: An End-to-End Speech-to-Speech Conversion Model and its Applications to Hearing-Impaired Speech and Speech Separation”, arxiv.org, Cornell University Library, 201 Olin Library Cornell Univ… [cited by applicant]
Harvill John et al: “Synthesis of New Words for Improved Dysarthric Speech Recognition on an Expanded Vocabulary”, ICASSP 2021—2021 IEEE International Conference on Acoustics , Speech and Signal Processing (ICASSP), IEE… [cited by applicant]
Chen Chen-Yu et al: “Enhancing Intelligibility of Dysarthric Speech Using Gated Convolutional-Based Voice Conversion System”, Interspeech 2020, Feb. 2, 2020 (Feb. 2, 2020), pp. 4686-4690, XP093046276 , ISCA DO : 10.2143… [cited by applicant]
Berrak Sisman et al: “An overview of Voice Conversion and its Challenges: From Statistical Modeling to Deep Learning”, arxiv.org, Cornell University Library, 201 Olin Library Cornell University Ithaca, NY 14853, Aug. 9,… [cited by applicant]
Berrak Sisman et al: “Ari Overview of Voice Conversion and its Challenges: From Statistical. Modeling to Deep Learning”, arxiv.org, Cornell University Library, 201 Olin Library Cornell University Ithaca, NY 14853, Nov. … [cited by applicant]
Momina Masood et al: “Deepfakes Generation and Detection: State-of-the-art, open challenges, countermeasures, and way forward”, arxiv.org, Cornell University Library, 201 Olin Library Cornell University Ithaca, NY 14853… [cited by applicant]
Japanese Office Action for the related Application No. 2024-555419 dated Oct. 21, 2025. [cited by applicant]