IP Library › Granted Patent US 10,395,654
Granted Patent B2
US 10,395,654 · App. 15/673,574 · Granted Aug 27, 2019

Text normalization based on a data-driven learning network

Inventors: Ladan Golipour (Cupertino, CA); Matthias Neeracher (Zürich, CH); Ramya Rasipuram (Sunnyvale, CA)
Assignee: Apple Inc.
G10L15/22G10L15/16G10L15/1815G10L15/26G10L15/30G10L2013/083
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 10,395,654
App. No.
15/673,574
Granted
Aug 27, 2019
Kind
B2
Abstract

Systems and processes for operating an intelligent automated assistant to perform text-to-speech conversion are provided. An example method includes, at an electronic device having one or more processors, receiving a text corpus comprising unstructured natural language text. The method further includes generating a sequence of normalized text based on the received text corpus; and generating a pronunciation sequence representing the sequence of the normalized text. The method further includes causing an audio output to be provided to the user based on the pronunciation sequence. At least one of the sequence of normalized text and the pronunciation sequence is generated based on a data-driven learning network.

Claims (81)

1. An electronic device comprising:

one or more processors;

memory; and

one or more programs stored in memory, the one or more programs including instructions for:

receiving a text corpus comprising unstructured natural language text;

generating, based on the received text corpus, a sequence of normalized text;

generating a pronunciation sequence representing the sequence of the normalized text; and

causing an audio output to be provided to the user based on the pronunciation sequence, wherein the sequence of normalized text is generated by a first data-driven learning network, wherein the pronunciation sequence is generated based on a second data-driven learning network, and wherein the first data-driven learning network is different from the second data-driven learning network.

2. The electronic device of claim 1 , wherein the unstructured natural language text comprises at least one non-standard word.

3. The electronic device of claim 2 , wherein the at least one non-standard word includes at least one of:

an abbreviation;

at least one of a letter sequence or a letter-and-symbol sequence;

a cardinal number;

an ordinal number;

a number as digits;

a number and an associated address;

a date;

a time;

a telephone number;

a currency;

a symbol;

a punctuation mark; and

at least a silent word or a silent sequence of letters and/or symbols.

4. The electronic device of claim 1 , wherein generating, based on the received text corpus, the sequence of normalized text comprises:

generating one or more tokens based on the text corpus, wherein the tokens include at least one non-standard word;

extracting features associated with the one or more tokens;

classifying the tokens based on the extracted features; and

generating a grapheme sequence of normalized text based on the classified tokens.

5. The electronic device of claim 4 , wherein the features include one or more patterns associated with the one or more tokens.

6. The electronic device of claim 4 , wherein the features include at least one of morphological features and categorical features.

7. The electronic device of claim 4 , wherein the features include at least one of one or more lexical features or one or more semantic features.

8. The electronic device of claim 4 , wherein extracting the features associated with the one or more tokens comprises:

determining features associated with the one or more tokens; and

performing word embedding of the tokens based on the determined features.

9. The electronic device of claim 8 , wherein determining features associated with the one or more tokens comprises determining features of a current token and a predetermined number of neighboring tokens of the current token.

10. The electronic device of claim 8 , wherein performing word embedding of the tokens based on the determined features comprises mapping each token to a feature vector.

11. The electronic device of claim 4 , wherein classifying the tokens based on extracted features comprises:

determining a label for each token;

associating the label and the corresponding token; and

storing the association of the label and the corresponding token.

12. The electronic device of claim 1 , wherein generating the pronunciation sequence representing the sequence of the normalized text comprises:

generating one or more vectors representing encoded grapheme sequences, wherein the grapheme sequences correspond to the sequence of the normalized text;

providing an attention layer, wherein the attention layer indicates relations between one or more vectors representing the encoded grapheme sequences and one or more vectors representing the pronunciation sequence; and

generating the pronunciation sequence based on the attention layer.

13. The electronic device of claim 12 , wherein generating one or more vectors representing encoded grapheme sequences comprises:

generating a character embedding matrix representing features of one or more grapheme sequences corresponding to the sequence of the normalized text; and

generating, based on the character embedding matrix, one or more vectors representing the encoded grapheme sequences.

14. The electronic device of claim 12 , wherein generating the pronunciation sequence based on the attention layer comprises:

processing, based on the attention layer, the one or more vectors representing the encoded grapheme sequences;

generating a phoneme embedding matrix based on the processing result; and

predicting one or more phoneme sequences based on the phoneme embedding matrix.

15. A non-transitory computer-readable storage medium storing one or more programs, the one or more programs comprising instructions, which when executed by one or more processors

of an electronic device, cause the electronic device to:

receive a text corpus comprising unstructured natural language text;

generate, based on the received text corpus, a sequence of normalized text;

generate a pronunciation sequence representing the sequence of the normalized text; and

cause an audio output to be provided to the user based on the pronunciation sequence, wherein the sequence of normalized text is generated by a first data-driven learning network, wherein the pronunciation sequence is generated based on a second data-driven learning network, and wherein the first data-driven learning network is different from the second data-driven learning network.

16. The computer-readable storage medium of claim 15 , wherein generating the sequence of normalized text comprises:

generating one or more tokens based on the text corpus, wherein the tokens include at least one non-standard word;

extracting features associated with the one or more tokens;

classifying the tokens based on the extracted features; and

generating a grapheme sequence of normalized text based on the classified tokens.

17. The computer-readable storage medium of claim 15 , wherein generating the pronunciation sequence representing the sequence of the normalized text comprises:

generating one or more vectors representing encoded grapheme sequences, wherein the grapheme sequences correspond to the sequence of the normalized text;

providing an attention layer, wherein the attention layer indicates relations between one or more vectors representing the encoded grapheme sequences and one or more vectors representing the pronunciation sequence; and

generating the pronunciation sequence based on the attention layer.

18. A method for performing text-to-speech conversion, comprising:

at one or more electronic devices with one or more processors and memory;

receiving a text corpus comprising unstructured natural language text;

generating, based on the received text corpus, a sequence of normalized text;

generating a pronunciation sequence representing the sequence of the normalized text; and

causing an audio output to be provided to the user based on the pronunciation sequence, wherein the sequence of normalized text is generated by a first data-driven learning network, wherein the pronunciation sequence is generated based on a second data-driven learning network, and wherein the first data-driven learning network is different from the second data-driven learning network.

19. The method of claim 18 , wherein generating the sequence of normalized text comprises:

generating one or more tokens based on the text corpus, wherein the tokens include at least one non-standard word;

extracting features associated with the one or more tokens;

classifying the tokens based on the extracted features; and

generating a grapheme sequence of normalized text based on the classified tokens.

20. The method of claim 18 , wherein generating the pronunciation sequence representing the sequence of the normalized text comprises:

generating one or more vectors representing encoded grapheme sequences, wherein the grapheme sequences correspond to the sequence of the normalized text;

providing an attention layer, wherein the attention layer indicates relations between one or more vectors representing the encoded grapheme sequences and one or more vectors representing the pronunciation sequence; and

generating the pronunciation sequence based on the attention layer.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jan 4, 2018
From: GOLIPOUR, LADAN; NEERACHER, MATTHIAS; RASIPURAM, RAMYA
To: APPLE INC.
Reel/Frame 044535/0159 →
Continuity (2)
Provisional Application 62504736 · May 11, 2017
Related Publication 20180330729A1 · Nov 15, 2018
Cited By (2)
US 12,361,213 US 12,412,582