IP Library › Granted Patent US 11,699,430
Granted Patent B2
US 11,699,430 · App. 17/245,048 · Granted Jul 11, 2023

Using speech to text data in training text to speech models

Inventors: Andrew R. Freed (Cary, NC); Vamshi Krishna Thotempudi (San Jose, CA); Sujatha B. Perepa (Durham, NC)
Assignee: International Business Machines Corporation
G10L13/08G06N20/00G10L13/06
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,699,430
App. No.
17/245,048
Granted
Jul 11, 2023
Kind
B2
Abstract

A system and method for providing a text to speech output by receiving user audio data, determining a user region-specific-pronunciation classification according to the audio data, determining text for a response to the user according to the audio data, identifying a portion from the text, where a region specific-pronunciation dictionary includes the portion, and using a phoneme string, from the dictionary selected according to the user region-specific pronunciation classification, for the word in a text to speech output to the user.

Claims (79)

1. A computer implemented method for providing a text to speech output, the method comprising:

receiving user audio data;

determining, by the one or more computer processors, a string of text corresponding to the user audio data;

determining, by the one or more computer processors, a string of phonemes corresponding to the user audio data;

determining, by one or more computer processors, a user region-specific-pronunciation classification according to the string of phonemes;

determining, by the one or more computer processors, text for a response to the user according to the string of text;

identifying, by the one or more computer processors, a portion from the text for a response, wherein a region specific-pronunciation dictionary comprises the portion;

using, by the one or more computer processors, a phoneme string, from the region-specific-pronunciation dictionary selected according to the user region-specific pronunciation classification, for the portion in a text to speech output to the user; and

providing an audio text to speech output including the portion to the user using a system speaker.

2. The computer implemented method according to claim 1 , further comprising:

using, by the one or more computer processors, a default phoneme sequence in the text to speech output to the user for words from the text absent from the region-specific-pronunciation dictionary.

3. The computer implemented method according to claim 1 , further comprising building the region-specific-pronunciation dictionary by:

receiving, by the one or more computer processors, audio data from a plurality of speakers, the audio data comprising domain-specific portions and region-specific pronunciations of the domain-specific portions;

classifying, by the one or more computer processors, the audio data according to a region-specific-pronunciation;

determining, by the one or more computer processors, a most common region-specific pronunciation for a domain-specific portion; and

storing, by the one or more computer processors, the most common region-specific pronunciation for the domain-specific portion as the phoneme string for the domain-specific portion-region specific pronunciation combination.

4. The computer implemented method according to claim 3 , further comprising:

defining, by the one or more computer processors, domain-specific portions.

5. The computer implemented method according to claim 3 , further comprising:

converting, by the one or more computer processors, the audio data to text data; and

scanning, by the one or more computer processors, the text data for domain-specific portions.

6. The computer implemented method according to claim 1 , wherein the portion comprises at least one of a word, an n-gram, and a phrase.

7. The computer implemented method according to claim 1 , further comprising:

determining, by the one or more computer processors, a user text from the audio data:

determining, by the one or more computer processors, a response according to the user text;

scanning, by the one or more computer processors, the response for domain portions; and

matching, by the one or more computer processors, a domain portion with a region-specific pronunciation dictionary entry.

8. A computer program product for providing a text to speech output, the computer program product comprising one or more computer readable storage devices and collectively stored program instructions on the one or more computer readable storage devices, the stored program instructions comprising:

program instructions to receive user audio data;

program instructions to determine a string of text corresponding to the user audio data;

program instructions to determine a string of phonemes corresponding to the user audio data;

program instructions to determine a user region-specific-pronunciation classification according to the string of phonemes;

program instructions to determine text for a response to the user according to the string of text;

program instructions to identify a portion from the text for a response, wherein a region-specific-pronunciation dictionary comprises the portion;

program instructions to use a phoneme string, from the region-specific-pronunciation dictionary selected according to the user region-specific pronunciation classification, for the portion in a text to speech output to the user; and

program instructions to provide an audio text to speech output including the portion, to the user using a system speaker.

9. The computer program product according to claim 8 , the stored program instructions further comprising:

program instructions to use a default phoneme sequence in the text to speech output to the user for words from the text absent from the region-specific-pronunciation dictionary.

10. The computer program product according to claim 8 , the stored program instructions further comprising program instructions to build the region-specific-pronunciation dictionary by:

receiving audio data from a plurality of speakers, the audio data comprising domain-specific portions and region-specific pronunciations of the domain-specific portions;

classifying the audio data according to a region-specific-pronunciation;

determining a most common region-specific pronunciation for a domain-specific portion; and

storing the most common region-specific pronunciation for the domain-specific portion as the phoneme string for the domain-specific portion-region specific pronunciation combination.

11. The computer program product according to claim 10 , the stored program instructions further comprising:

program instructions to define domain-specific portions.

12. The computer program product according to claim 10 , the stored program instructions further comprising:

program instructions to convert the audio data to text data; and

program instructions to scan the text data for domain-specific portions.

13. The computer program product according to claim 8 , wherein the portion comprises at least one of a word, an n-gram, and a phrase.

14. The computer program product according to claim 8 , the stored program instructions further comprising:

program instructions to determine a user text from the audio data:

program instructions to determine a response according to the user text;

program instructions to scan the response for domain portions; and

program instructions to match a domain portion with a region-specific pronunciation dictionary entry.

15. A computer system for providing a text to speech output, the computer system comprising:

one or more computer processors;

one or more computer readable storage devices; and

stored program instructions on the one or more computer readable storage devices for execution by the one or more computer processors, the stored program instructions comprising:

program instructions to receive user audio data;

program instructions to determine a string of text corresponding to the user audio data;

program instructions to determine a string of phonemes corresponding to the user audio data;

program instructions to determine a user region-specific-pronunciation classification according to the string of phonemes;

program instructions to determine text for a response to the user according to the string of text;

program instructions to identify a portion from the text for a response, wherein a region-specific-pronunciation dictionary comprises the portion;

program instructions to use a phoneme string, from the region-specific-pronunciation dictionary selected according to the user region-specific pronunciation classification, for the portion in a text to speech output to the user; and

program instructions to provide an audio text to speech output including the portion, to the user using a system speaker.

16. The computer system according to claim 15 , the stored program instructions further comprising:

program instructions to use a default phoneme sequence in the text to speech output to the user for words from the text absent from the region-specific-pronunciation dictionary.

17. The computer system according to claim 15 , the stored program instructions further comprising program instructions to build the region-specific-pronunciation dictionary by:

receiving audio data from a plurality of speakers, the audio data comprising domain-specific portions and region-specific pronunciations of the domain-specific portions;

classifying the audio data according to a region-specific-pronunciation;

determining a most common region-specific pronunciation for a domain-specific portion; and

storing the most common region-specific pronunciation for the domain-specific portion as the phoneme string for the domain-specific portion-region specific pronunciation combination.

18. The computer system according to claim 17 , the stored program instructions further comprising:

program instructions to define domain-specific portions.

19. The computer system according to claim 17 , the stored program instructions further comprising:

program instructions to convert the audio data to text data; and

program instructions to scan the text data for domain-specific portions.

20. The computer system according to claim 15 , wherein the portion comprises at least one of a word, an n-gram, and a phrase.

Assignments (2)
CORRECTIVE ASSIGNMENT TO CORRECT THE TITLE OF INVENTION INSIDE THE ASSIGNMENT DOCUMENT PREVIOUSLY RECORDED AT REEL: 056091 FRAME: 0603. ASSIGNOR(S) HEREBY CONFIRMS THE ASSIGNMENT. Recorded Jul 19, 2021
From: FREED, ANDREW R.; THOTEMPUDI, VAMSHI KRISHNA; PEREPA, SUJATHA B.
To: INTERNATIONAL BUSINESS MACHINES CORPORATION
Reel/Frame 057066/0074 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Apr 30, 2021
From: FREED, ANDREW R.; THOTEMPUDI, VAMSHI KRISHNA; PEREPA, SUJATHA B.
To: INTERNATIONAL BUSINESS MACHINES CORPORATION
Reel/Frame 056091/0603 →
Continuity (1)
Related Publication 20220351715A1 · Nov 3, 2022