IP Library Granted Patent US 12,094,447
Granted Patent B2
US 12,094,447 · App. 17/293,404 · Granted Sep 17, 2024

Neural text-to-speech synthesis with multi-level text information

Inventors: Huaiping Ming (Redmond, WA); Lei He (Redmond, WA)
Assignee: Microsoft Technology Licensing, LLC
G10L13/08G06F40/20G06F40/205G06F40/253G06N3/045G06N20/20G10L13/047G10L13/06G10L25/30
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,094,447
App. No.
17/293,404
Granted
Sep 17, 2024
Kind
B2
Abstract

A method and apparatus for generating speech through neural text-to-speech (TTS) synthesis. A text input may be obtained ( 1310 ). Phoneme or character level text information may be generated based on the text input ( 1320 ). Context-sensitive text information may be generated based on the text input ( 1330 ). A text feature may be generated based on the phoneme or character level text information and the context-sensitive text information ( 1340 ). A speech waveform corresponding to the text input may be generated based at least on the text feature ( 1350 ).

Claims (71)

1. A method for generating speech through neural text-to-speech (TTS) synthesis, comprising:

obtaining a text input having a sequence;

generating phoneme or character level text information based on the text input;

generating context-sensitive text information among words based on the text input, the context-sensitive text information having word level text information where generating the context-sensitive text information comprises:

identifying a word sequence from the text input;

up-sampling the word sequence to align with the text input;

generating a word embedding vector sequence of the word sequence and a phoneme vector sequence;

generating sentence level text information having a grammatical parsing information sequence, wherein generating the sentence level text information comprises:

performing grammatical parsing on the text input to obtain a grammatical structure of the text input; and

generating the grammatical parsing information sequence based on the grammatical structure of the text input by:

extracting grammatical parsing information of each word in the text input from the grammatical structure;

up-sampling the grammatical parsing information of each word to align with corresponding phonemes or characters in a phoneme or character sequence of the text input thereby generating a grammatical parsing information sequence; and

combining a phoneme vector sequence, the word embedding vector sequence and the grammatical parsing information sequence;

generating a text feature via a multi-input encoder coupled to receive the phoneme or character level text information and the word embedding vector sequence as inputs;

generating acoustic features from the text feature via a decoder; and

generating a speech waveform corresponding to the text input based at least on the text feature.

2. The method of claim 1 , wherein word embedding vector is generated with a word embedding model that is based on neural machine translation (NMT).

3. The method of claim 1 , wherein the generating the text feature comprises:

generating the text feature based on the phoneme or character level text information and the word level text information.

4. The method of claim 1 , wherein the grammatical parsing information of each word comprises at least one of:

an indication of phrase type of at least one phrase containing a word;

an indication of whether the word is a border of the at least one phrase; and

an indication of relative position of the word in the at least one phrase.

5. The method of claim 1 , wherein the generating the text feature comprises:

generating the text feature based on the phoneme or character level text information and the sentence level text information.

6. The method of claim 1 , wherein the generating the text feature comprises:

generating the text feature based on the phoneme or character level text information, the word level text information and the sentence level text information.

7. The method of claim 1 , wherein the generating the text feature comprises:

generating a first text feature based on the phoneme or character level text information through a first neural network;

generating at least one second text feature based on the word level text information and/or the sentence level text information comprised in the context-sensitive text information through at least one second neural network; and

generating the text feature through concatenating the first text feature with the at least one second text feature.

8. The method of claim 1 , wherein the generating the text feature comprises:

concatenating the phoneme or character level text information with the word level text information and/or the sentence level text information comprised in the context-sensitive text information, to form mixed text information; and

generating the text feature based on the mixed text information through a first neural network.

9. The method of claim 1 , wherein the generating the text feature comprises:

generating at least one compressed representation of the word level text information and/or the sentence level text information comprised in the context-sensitive text information through at least one first neural network;

concatenating the phoneme or character level text information with the at least one compressed representation to form mixed text information; and

generating the text feature based on the mixed text information through a second neural network.

10. An apparatus for generating speech through neural text-to-speech (TTS) synthesis, comprising:

a text input obtaining module, for obtaining a text input having a sequence;

a phoneme or character level text information generating module, for generating phoneme or character level text information based on the text input;

a context-sensitive text information generating module, for generating context-sensitive text information among words based on the text input, the context-sensitive text information having word level text information where generating the context-sensitive text information comprises:

identifying a word sequence from the text input; and

up-sampling the word sequence to align with the text input; generating a word embedding vector sequence of the word sequence and a phoneme vector sequence;

generating sentence level text information having a grammatical parsing information sequence, wherein generating the sentence level text information comprises:

performing grammatical parsing on the text input to obtain a grammatical structure of the text input; and

generating the grammatical parsing information sequence based on the grammatical structure of the text input by:

extracting grammatical parsing information of each word in the text input from the grammatical structure;

up-sampling the grammatical parsing information of each word to align with corresponding phonemes or characters in a phoneme or character sequence of the text input thereby generating a grammatical parsing information sequence; and

combining a phoneme vector sequence, the word embedding vector sequence and the grammatical parsing information sequence;

a text feature generating module comprising a multi-input encoder, for generating a text feature from inputs including the phoneme or character level text information and the word embedding vector sequence as inputs;

an acoustic features generating module comprising a decoder for generating acoustic features from the text feature; and

a speech waveform generating module, for generating a speech waveform corresponding to the text input based at least on the text feature.

11. An apparatus for generating speech through neural text-to-speech (TTS) synthesis, comprising:

at least one processor; and

a memory storing computer-executable instructions that, when executed, cause the at least one processor to:

obtain a text input having a sequence;

generate phoneme or character level text information and context-sensitive text information among words based on the text input, the context-sensitive text information having word level text information where generating the context-sensitive text information comprises:

identifying a word sequence from the text input;

up-sampling the word sequence to align with the text input;

generating a word embedding vector sequence of the word sequence and a phoneme vector sequence;

generating sentence level text information having a grammatical parsing information sequence, wherein generating the sentence level text information comprises:

performing grammatical parsing on the text input to obtain a grammatical structure of the text input; and

generating the grammatical parsing information sequence based on the grammatical structure of the text input by:

extracting grammatical parsing information of each word in the text input from the grammatical structure;

up-sampling the grammatical parsing information of each word to align with corresponding phonemes or characters in a phoneme or character sequence of the text input thereby generating a grammatical parsing information sequence; and

combining a phoneme vector sequence, the word embedding vector sequence and the grammatical parsing information sequence;

generate context-sensitive text information among words based on the text input;

generate a text feature via a multi-input encoder coupled to receive the phoneme or character level text information and the word embedding vector sequence as inputs;

generate acoustic features from the text feature via a decoder, and

generate a speech waveform corresponding to the text input based at least on the text feature.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded May 13, 2021
From: MING, HUAIPING; HE, LEI
To: MICROSOFT TECHNOLOGY LICENSING, LLC
Reel/Frame 056235/0013 →
Continuity (1)
Related Publication 20220020355A1 · Jan 20, 2022