IP Library › Granted Patent US 12,744,029
Granted Patent B2
US 12,744,029 · App. 18/746,809 · Granted Sep 22, 2026

Phonemes and graphemes for neural text-to-speech

Inventors: Ye Jia (Mountain View, CA); Byungha Chun (Tokyo, JP); Yu Zhang (Mountain View, CA); Jonathan Shen (Mountain View, CA); Yonghui Wu (Fremont, CA)
Assignee: Google LLC
G10L13/086G06F40/263G06F40/279G06N3/08G10L13/047
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,744,029
App. No.
18/746,809
Granted
Sep 22, 2026
Kind
B2
Abstract

A method includes receiving a text input including a sequence of words represented as an input encoder embedding. The input encoder embedding includes a plurality of tokens, with the plurality of tokens including a first set of grapheme tokens representing the text input as respective graphemes and a second set of phoneme tokens representing the text input as respective phonemes. The method also includes, for each respective phoneme token of the second set of phoneme tokens: identifying a respective word of the sequence of words corresponding to the respective phoneme token and determining a respective grapheme token representing the respective word of the sequence of words corresponding to the respective phoneme token. The method also includes generating an output encoder embedding based on a relationship between each respective phoneme token and the corresponding grapheme token determined to represent a same respective word as the respective phoneme token.

Claims (54)

1 . A computer-implemented method executed by data processing hardware causes the data processing hardware to perform operations comprising:

receiving, at an encoder of a speech synthesis model, a text input comprising a sequence of words represented as an input encoder embedding, the input encoder embedding comprising set of grapheme tokens representing the text input as respective graphemes and a set of phoneme tokens representing the text input as respective phonemes, wherein the input encoder embedding represents, for each respective word in the sequence of words, respective sub-word level positions for both one or more of the grapheme tokens from the set of grapheme tokens that correspond to the respective word and one or more of the phoneme tokens from the set of phoneme tokens that correspond to the respective word;

for each respective phoneme token of the set of phoneme tokens:

identifying, by the encoder, a respective word of the sequence of words corresponding to the respective phoneme token; and

determining, by the encoder, one or more respective grapheme tokens representing the same respective word of the sequence of words corresponding to the respective phoneme token based on the respective sub-word level position for the respective phoneme token that corresponds to the respective word and the respective sub-word level position for each of the one or more respective grapheme tokens that correspond to the same respective word; and

generating, by the encoder, an output encoder embedding for use by a decoder of the speech synthesis model to synthesize speech from the text input, the output encoder embedding based on a relationship between each respective phoneme token and the one or more respective grapheme tokens determined to represent a same respective word as the respective phoneme token.

2 . The method of claim 1 , wherein the input encoder embedding further represents a combination of:

a segment embedding; and

a position embedding.

3 . The method of claim 2 , wherein the position embedding represents an overall index of position for each grapheme token of the set of grapheme tokens and each phoneme token of the set of phoneme tokens of the input encoder embedding.

4 . The method of claim 1 , wherein the speech synthesis model comprises an attention mechanism in communication with the encoder.

5 . The method of claim 1 , wherein the speech synthesis model comprises a duration-based upsampler in communication with the encoder.

6 . The method of claim 1 , wherein the input encoder embedding further comprises a special token identifying a language of the input text.

7 . The method of claim 1 , wherein the encoder of the speech synthesis model is pre-trained by:

feeding the encoder a plurality of training examples, each training example represented as a sequence of training grapheme tokens corresponding to a training sequence of words and a sequence of training phoneme tokens corresponding to the same training sequence of words;

masking a training phoneme token from the sequence of training phoneme tokens for a respective word from the training sequence of words; and

masking a training grapheme token from the sequence of training phoneme tokens for the respective word from the training sequence of words.

8 . The method of claim 1 , wherein:

the speech synthesis model comprises a multilingual speech synthesis model; and

the encoder of the speech synthesis model is pre-trained using a classification objective to predict a classification token of the input encoder embedding, the classification token comprising a language identifier.

9 . The method of claim 1 , wherein:

the speech synthesis model comprises a multilingual speech synthesis model; and

the output encoder embedding comprises a sequence of encoder tokens, each encoder token comprising language information about the input text.

10 . The method of claim 1 , wherein:

the speech synthesis model comprises a multi-accent speech synthesis model; and

the encoder of the speech synthesis model is pre-trained using a classification objective to predict a classification token, the classification token comprising an accent identifier.

11 . A system comprising:

data processing hardware; and

memory hardware in communication with the data processing hardware, the memory hardware storing instructions that when executed on the data processing hardware cause the data processing hardware to perform operations comprising:

receiving, at an encoder of a speech synthesis model, a text input comprising a sequence of words represented as an input encoder embedding, the input encoder embedding comprising a set of grapheme tokens representing the text input as respective graphemes and a set of phoneme tokens representing the text input as respective phonemes, wherein the input encoder embedding represents, for each respective word in the sequence of words, respective sub-word level positions for both one or more of the grapheme tokens from the set of grapheme tokens that correspond to the respective word and one or more of the phoneme tokens from the set of phoneme tokens that correspond to the respective word;

for each respective phoneme token of the set of phoneme tokens:

identifying, by the encoder, a respective word of the sequence of words corresponding to the respective phoneme token; and

determining, by the encoder, one or more respective grapheme tokens representing the same respective word of the sequence of words corresponding to the respective phoneme token based on the respective sub-word level position for the respective phoneme token that corresponds to the respective word and the respective sub-word level position for each of the one or more respective grapheme tokens that correspond to the same respective word; and

generating, by the encoder, an output encoder embedding for use by a decoder of the speech synthesis model to synthesize speech from the text input, the output encoder embedding based on a relationship between each respective phoneme token and the one or more respective grapheme tokens determined to represent a same respective word as the respective phoneme token.

12 . The system of claim 11 , wherein the input encoder embedding further represents a combination of:

a segment embedding; and

a position embedding.

13 . The system of claim 12 , wherein the position embedding represents an overall index of position for each grapheme token of the set of grapheme tokens and each phoneme token of the set of phoneme tokens of the input encoder embedding.

14 . The system of claim 11 , wherein the speech synthesis model comprises an attention mechanism in communication with the encoder.

15 . The system of claim 11 , wherein the speech synthesis model comprises a duration-based upsampler in communication with the encoder.

16 . The system of claim 11 , wherein the input encoder embedding further comprises a special token identifying a language of the input text.

17 . The system of claim 11 , wherein the encoder of the speech synthesis model is pre-trained by:

feeding the encoder a plurality of training examples, each training example represented as a sequence of training grapheme tokens corresponding to a training sequence of words and a sequence of training phoneme tokens corresponding to the same training sequence of words;

masking a training phoneme token from the sequence of training phoneme tokens for a respective word from the training sequence of words; and

masking a training grapheme token from the sequence of training phoneme tokens for the respective word from the training sequence of words.

18 . The system of claim 11 , wherein:

the speech synthesis model comprises a multilingual speech synthesis model; and

the encoder of the speech synthesis model is pre-trained using a classification objective to predict a classification token of the input encoder embedding, the classification token comprising a language identifier.

19 . The system of claim 11 , wherein:

the speech synthesis model comprises a multilingual speech synthesis model; and

the output encoder embedding comprises a sequence of encoder tokens, each encoder token comprising language information about the input text.

20 . The system of claim 11 , wherein:

the speech synthesis model comprises a multi-accent speech synthesis model; and

the encoder of the speech synthesis model is pre-trained using a classification objective to predict a classification token, the classification token comprising an accent identifier.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jun 18, 2024
From: JIA, YE; CHUN, BYUNGHA; ZHANG, YU; SHEN, JONATHAN; WU, YONGHUI
To: GOOGLE LLC
Reel/Frame 067759/0771 →
Continuity (3)
Continuation 17643684 · Dec 10, 2021
Provisional Application 63166929 · Mar 26, 2021
Related Publication 20240339106A1 · Oct 10, 2024
References Cited (16)
US 20070112569A1 · Wang et al. · 2007 [cited by applicant]
US 20180247636A1 · Arik · 2018 [cited by examiner]
US 20220020355A1 · Ming · 2022 [cited by examiner]
US 20220246136A1 · Yang · 2022 [cited by examiner]
Kyle Kastner, Joao Felipe Santos, Yoshua Bengio, Aaron Courville, “Representation Mixing for TTS Synthesis”, Nov. 24, 2018, arXiv: 1811.07240 (Year: 2018). [cited by examiner]
ST-BERT: Cross-Modal Langauge Model Pre-Training for End-to-End Spoken Language Understanding, Oct. 23, 2020, arXiv: 2010.12283v2 (Year: 2020) (Year: 2020). [cited by examiner]
Kyle Kastner, Joao Felipe Santos, Yoshua Bengio, Aaron Courville, “Representation Mixing for TTS Synthesis”, Nov. 24, 2018, arXiv: 1811.07240 (Year: 2019). [cited by examiner]
ST-BERT: Cross-Modal Langauge Model Pre-Training for End-to-End Spoken Language Understanding, Oct. 23, 2020, arXiv: 2010.12283v2 (Year: 2020). [cited by examiner]
Zhehuai Chen, Mahaveer Jain, Yongqiang Wang, Michael L. Seltzer, Christian Fuegen, “Joint Grapheme and Phoneme Embeddings for Contextual End-to-End ASR”, Interspeech 2019 (Year: 2019). [cited by examiner]
Apr. 5, 2022 Written Opinion (WO) of the International Searching Authority (ISA) and International Search Report (ISR) issued in International Application No. PCT/US2021/062837. [cited by applicant]
Chung Yu-An et al.: “Semi-supervised Training for Improving Data Efficiency in End-to-end Speech Synthesis”, ICASSP 2019—2019 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), IEEE, May … [cited by applicant]
Kastner Kyle et al: “Representation Mixing for TTS Synthesis”, ICASSP 2019—2019 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), IEEE, May 12, 2019 (May 12, 2019), pp. 5906-5910, XP0335… [cited by applicant]
Ye Jia et al: “PnG BERT: Augmented BERT on Phonemes and Graphemes for Neural TTS”, arxiv.org, Cornell University Library, 201 Olin Library Cornell University Ithaca, Ny 14853, Jun. 7, 2021 (Jun. 7, 2021), XP081980096. [cited by applicant]
Jacob Devlin et al. “Bert: Pre-training of deep bidirectional transformers for language understanding.” arXiv preprint arXiv:1810.04805 (2018). [cited by applicant]
Minjeong Kim et al. “St-bert: Cross-modal language model pre-training for end-to-end spoken language understanding.” ICASSP 2021-2021 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP). IE… [cited by applicant]
Chinese Office Action for the related Application No. 202180096396.2 dated May 27, 2026. [cited by applicant]