IP Library › Granted Patent US 12,272,349
Granted Patent B2
US 12,272,349 · App. 18/487,227 · Granted Apr 8, 2025

Attention-based clockwork hierarchical variational encoder

Inventors: Robert Clark (Hertfordshire, GB); Chun-An Chan (Mountain View, CA); Vincent Wan (London, GB)
Assignee: Google LLC
G10L13/10G10L25/30G10L2013/105
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,272,349
App. No.
18/487,227
Granted
Apr 8, 2025
Kind
B2
Abstract

A method for representing an intended prosody in synthesized speech includes receiving a text utterance having at least one word, and selecting an utterance embedding for the text utterance. Each word in the text utterance has at least one syllable and each syllable has at least one phoneme. The utterance embedding represents an intended prosody. For each syllable, using the selected utterance embedding, the method also includes: predicting a duration of the syllable by decoding a prosodic syllable embedding for the syllable based on attention by an attention mechanism to linguistic features of each phoneme of the syllable and generating a plurality of fixed-length predicted frames based on the predicted duration for the syllable.

Claims (66)

1. A computer-implemented method executed on data processing hardware that causes the data processing hardware to perform operations comprising:

receiving a text utterance having at least one word, each word having at least one syllable, each syllable having at least one phoneme;

selecting an utterance embedding for the text utterance, the utterance embedding representing an intended prosody; and

for each corresponding syllable, using the selected utterance embedding:

predicting a duration of the corresponding syllable by decoding a prosodic syllable embedding for the corresponding syllable based on attention by an attention mechanism to linguistic features of each phoneme of an adjacent syllable to the corresponding syllable; and

generating a plurality of fixed-length predicted frames based on the predicted duration for the corresponding syllable.

2. The method of claim 1 , wherein the operations further comprise:

predicting a pitch contour of the corresponding syllable based on the predicted duration of the corresponding syllable, and

wherein the plurality of fixed-length predicted frames comprise fixed-length predicted pitch frames, each fixed-length predicted pitch frame representing part of the predicted pitch contour of the corresponding syllable.

3. The method of claim 1 , wherein the operations further comprise, for each corresponding syllable, using the selected utterance embedding:

predicting an energy contour of the corresponding syllable based on the predicted duration of the corresponding syllable; and

generating a plurality of fixed-length predicted energy frames based on the predicted duration of the corresponding syllable, each fixed-length energy frame representing the predicted energy contour of the corresponding syllable.

4. The method of claim 1 , wherein the plurality of fixed length predicted frames comprise fixed-length predicted spectral frames for the corresponding syllable.

5. The method of claim 1 , wherein a network representing a hierarchical linguistic structure of the text utterance comprises:

a first level including each word of the text utterance;

a second level including each syllable of the text utterance; and

a third level including each fixed-length predicted frame for each syllable of the text utterance.

6. The method of claim 1 , wherein predicting the duration of the corresponding syllable comprises:

for each phoneme associated with the adjacent syllable to the corresponding syllable:

encoding one or more linguistic features of a corresponding phoneme;

inputting the encoded one or more linguistic features into the attention mechanism; and

applying the attention of the attention mechanism to the prosodic syllable embedding.

7. The method of claim 1 , wherein the adjacent syllable to the corresponding syllable comprises one of a preceding adjacent syllable that precedes the corresponding syllable or a subsequent adjacent syllable that is subsequent to the corresponding syllable.

8. The method of claim 7 , wherein predicting the duration of the corresponding syllable by decoding the prosodic syllable embedding for the corresponding syllable is further based on attention by the attention mechanism to linguistic features of each phoneme of another adjacent syllable to the corresponding syllable, the another adjacent syllable to the corresponding syllable comprising the other one of the preceding adjacent syllable or the subsequent adjacent syllable.

9. The method of claim 1 , wherein the operations further comprise:

receiving training data including a plurality of reference audio signals, each reference audio signal comprising a spoken utterance of human speech and having a corresponding prosody; and

training a deep neural network for a prosody model by encoding each reference audio signal into a corresponding fixed-length utterance embedding representing the corresponding prosody of the reference audio signal.

10. The method of claim 9 , wherein the operations further comprise generating the selected utterance embedding by encoding linguistic features for a plurality of linguistic units with a frame-based syllable embedding and a phone feature-based syllable embedding.

11. The method of claim 1 , wherein the utterance embedding comprises a fixed-length numerical vector.

12. The method of claim 1 , wherein the attention of the attention mechanism comprises location-based attention.

13. The method of claim 12 , wherein the location-based attention comprises monotonically shifting, location sensitive attention, the monotonically shifting, location sensitive attention defined by a window of phoneme information for a respective syllable.

14. The method of claim 1 , wherein the attention mechanism comprises a transformer.

15. A system comprising:

data processing hardware; and

memory hardware in communication with the data processing hardware, the memory hardware storing instructions that when executed on the data processing hardware cause the data processing hardware to perform operations comprising:

receiving a text utterance having at least one word, each word having at least one syllable, each syllable having at least one phoneme;

selecting an utterance embedding for the text utterance, the utterance embedding representing an intended prosody; and

for each corresponding syllable, using the selected utterance embedding:

predicting a duration of the corresponding syllable by decoding a prosodic syllable embedding for the corresponding syllable based on attention by an attention mechanism to linguistic features of each phoneme of an adjacent syllable to the corresponding syllable; and

generating a plurality of fixed-length predicted frames based on the predicted duration for the corresponding syllable.

16. The system of claim 15 , wherein the operations further comprise:

predicting a pitch contour of the corresponding syllable based on the predicted duration of the corresponding syllable, and

wherein the plurality of fixed-length predicted frames comprise fixed-length predicted pitch frames, each fixed-length predicted pitch frame representing part of the predicted pitch contour of the corresponding syllable.

17. The system of claim 15 , wherein the operations further comprise, for each corresponding syllable, using the selected utterance embedding:

predicting an energy contour of the corresponding syllable based on the predicted duration of the corresponding syllable; and

generating a plurality of fixed-length predicted energy frames based on the predicted duration of the corresponding syllable, each fixed-length energy frame representing the predicted energy contour of the corresponding syllable.

18. The system of claim 15 , wherein the plurality of fixed length predicted frames comprise fixed-length predicted spectral frames for the corresponding syllable.

19. The system of claim 15 , wherein a network representing a hierarchical linguistic structure of the text utterance comprises:

a first level including each word of the text utterance;

a second level including each syllable of the text utterance; and

a third level including each fixed-length predicted frame for each syllable of the text utterance.

20. The system of claim 15 , wherein predicting the duration of the corresponding syllable comprises:

for each phoneme associated with the adjacent syllable to the corresponding syllable:

encoding one or more linguistic features of a corresponding phoneme;

inputting the encoded one or more linguistic features into the attention mechanism; and

applying the attention of the attention mechanism to the prosodic syllable embedding.

21. The system of claim 15 , wherein the adjacent syllable to the corresponding syllable comprises one of a preceding adjacent syllable that precedes the corresponding syllable or a subsequent adjacent syllable that is subsequent to the corresponding syllable.

22. The system of claim 21 , wherein predicting the duration of the corresponding syllable by decoding the prosodic syllable embedding for the corresponding syllable is further based on attention by the attention mechanism to linguistic features of each phoneme of another adjacent syllable to the corresponding syllable, the another adjacent syllable to the corresponding syllable comprising the other one of the preceding adjacent syllable or the subsequent adjacent syllable.

23. The system of claim 15 , wherein the operations further comprise:

receiving training data including a plurality of reference audio signals, each reference audio signal comprising a spoken utterance of human speech and having a corresponding prosody; and

training a deep neural network for a prosody model by encoding each reference audio signal into a corresponding fixed-length utterance embedding representing the corresponding prosody of the reference audio signal.

24. The system of claim 23 , wherein the operations further comprise generating the selected utterance embedding by encoding linguistic features for a plurality of linguistic units with a frame-based syllable embedding and a phone feature-based syllable embedding.

25. The system of claim 15 , wherein the utterance embedding comprises a fixed-length numerical vector.

26. The system of claim 15 , wherein the attention of the attention mechanism comprises location-based attention.

27. The system of claim 26 , wherein the location-based attention comprises monotonically shifting, location sensitive attention, the monotonically shifting, location sensitive attention defined by a window of phoneme information for a respective syllable.

28. The system of claim 15 , wherein the attention mechanism comprises a transformer.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Oct 16, 2023
From: CLARK, ROBERT; CHAN, CHUN-AN; WAN, VINCENT
To: GOOGLE LLC
Reel/Frame 065227/0371 →
Continuity (2)
Continuation 17756264
Related Publication 20240038214A1 · Feb 1, 2024
References Cited (38)
US 5715372A · Meyers · 1998 [cited by examiner]
US 7340393B2 · Mitsuyoshi · 2008 [cited by examiner]
US 8280013B2 · Bushey · 2012 [cited by examiner]
US 8412528B2 · Fischer et al. · 2013 [cited by applicant]
US 8571849B2 · Bangalore · 2013 [cited by examiner]
US 8886542B2 · Lagadec · 2014 [cited by examiner]
US 11182566B2 · Jaitly · 2021 [cited by examiner]
US 11404045B2 · Choi · 2022 [cited by examiner]
US 11514888B2 · Finkelstein · 2022 [cited by examiner]
US 11640832B2 · Ma · 2023 [cited by examiner]
US 11657800B2 · Kim · 2023 [cited by examiner]
US 11688399B2 · Diamant · 2023 [cited by examiner]
US 11694675B2 · Aoki · 2023 [cited by examiner]
US 11705105B2 · Chae · 2023 [cited by examiner]
US 11715485B2 · Park · 2023 [cited by examiner]
US 11727939B2 · Grancharov · 2023 [cited by examiner]
US 11776544B2 · Kim · 2023 [cited by examiner]
US 11790935B2 · Lee · 2023 [cited by examiner]
US 12080272B2 · Clark · 2024 [cited by examiner]
US 20050154581A1 · Ferrieux · 2005 [cited by examiner]
US 20060235690A1 · Tomasic · 2006 [cited by examiner]
US 20110246076A1 · Su · 2011 [cited by examiner]
US 20120271634A1 · Lenke · 2012 [cited by examiner]
US 20170345411A1 · Raitio · 2017 [cited by examiner]
US 20190348020A1 · Clark · 2019 [cited by examiner]
US 20200074985A1 · Clark · 2020 [cited by examiner]
US 20220051654A1 · Finkelstein · 2022 [cited by examiner]
US 20220139377A1 · Lee · 2022 [cited by examiner]
US 20220415306A1 · Clark · 2022 [cited by examiner]
US 20230064749A1 · Finkelstein · 2023 [cited by examiner]
US 20230230572A1 · Biadsy · 2023 [cited by examiner]
US 20230298577A1 · Yasa · 2023 [cited by examiner]
US 20240038214A1 · Clark · 2024 [cited by examiner]
Park Jungbae et al, “Phonemic-level Duration Control Using Attention Alignment for Natural Speech Synthesis”, ICASSP 2019—2019 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), IEEE, May… [cited by applicant]
Yibin Zheng et al, “BLSTM-CRF Based End-to-End Prosodic Boundary Prediction with Context Sensitive Embeddings in a Text-to-Speech Front-End”, Interspeech 2018, Jan. 1, 2018 (Jan. 1, 2018), p. 47-51, XP055697974. [cited by applicant]
Yuxuan Wang et al, “Tacotron: Towards End-To-End Speech Synthesis”, DOI: 10.21437/ Interspeech.2017-1452 external link Apr. 6, 2017 (Apr. 6, 2017), Retrieved from the Internet: URL:https://arxiv.org/pdf/1703.10135.pdf, … [cited by applicant]
Jun. 17, 2021 Written Opinion (WO) of the International Searching Authority (ISA) and International Search Report (ISR) issued in International Application No. PCT/US2019/065566. [cited by applicant]
Intellectual Property India. Examination report relating to application No. 202227026391, dated Aug. 30, 2022. [cited by applicant]