IP Library Granted Patent US 9,852,728
Granted Patent B2
US 9,852,728 · App. 14/733,289 · Granted Dec 26, 2017

Process for improving pronunciation of proper nouns foreign to a target language text-to-speech system

View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 9,852,728
App. No.
14/733,289
Granted
Dec 26, 2017
Kind
B2
Abstract

A system and method configured for use in a text-to-speech (TTS) system is provided. Embodiments may include identifying, using one or more processors, a word or phrase as a named entity and identifying a language of origin associated with the named entity. Embodiments may further include transliterating the named entity to a script associated with the language of origin. If the TTS system is operating in the language of origin, embodiments may include passing the transliterated script to the TTS system. If the TTS system is not operating in the language of origin, embodiments may include generating a phoneme sequence in the language of origin using a grapheme to phoneme (G2P) converter.

Claims (36)

1. A computer-implemented method configured for use in a text-to-speech (TTS) system comprising:

identifying, using one or more processors, a word or phrase as a named entity;

identifying a language of origin associated with the named entity;

transliterating the named entity to a script associated with the language of origin;

if the TTS system is operating in the language of origin, passing the transliterated script to the TTS system; and

if the TTS system is not operating in the language of origin, generating a phoneme sequence in the language of origin using a grapheme to phoneme (G2P) converter.

2. The method of claim 1 , further comprising:

if the TTS system is not operating in the language of origin, mapping the phoneme sequence to a sequence of target language phonemes.

3. The method of claim 2 , wherein mapping includes generating a map of most likely unigram, bigram, and trigram mappings from the phoneme sequence to the sequence of target language phonemes.

4. The method of claim 1 , wherein identifying a word or phrase as a named entity includes one or more of table lookup and contextual analysis.

5. The method of claim 1 , wherein identifying a language of origin associated with the named entity includes one or more of table lookup and shortest distance measures to an existing names database.

6. The method of claim 1 , further comprising:

augmenting a text to speech dictionary based upon, at least in part, the phoneme sequence.

7. The method of claim 6 , wherein the text to speech dictionary is associated with an automatic speech recognition (ASR) system.

8. A non-transitory computer-readable storage medium having stored thereon instructions, which when executed by a processor result in one or more operations configured for use in a text-to-speech (TTS) system, the operations comprising:

identifying, using one or more processors, a word or phrase as a named entity;

identifying a language of origin associated with the named entity;

transliterating the named entity to a script associated with the language of origin;

if the TTS system is operating in the language of origin, passing the transliterated script to the TTS system; and

if the TTS system is not operating in the language of origin, generating a phoneme sequence in the language of origin using a grapheme to phoneme (G2P) converter.

9. The non-transitory computer-readable storage medium of claim 8 , further comprising:

if the TTS system is not operating in the language of origin, mapping the phoneme sequence to a sequence of target language phonemes.

10. The non-transitory computer-readable storage medium of claim 9 , wherein mapping includes generating a map of most likely unigram, bigram, and trigram mappings from the phoneme sequence to the sequence of target language phonemes.

11. The non-transitory computer-readable storage medium of claim 8 , wherein identifying a word or phrase as a named entity includes one or more of table lookup and contextual analysis.

12. The non-transitory computer-readable storage medium of claim 8 , wherein identifying a language of origin associated with the named entity includes one or more of table lookup and shortest distance measures to an existing names database.

13. The non-transitory computer-readable storage medium of claim 8 , further comprising:

augmenting a text to speech dictionary based upon, at least in part, the phoneme sequence.

14. The non-transitory computer-readable storage medium of claim 13 , wherein the text to speech dictionary is associated with an automatic speech recognition (ASR) system.

15. A text to speech system comprising:

one or more processors configured to identify a word or phrase as a named entity, the one or more processors further configured to identify a language of origin associated with the named entity and transliterate the named entity to a script associated with the language of origin, if the TTS system is operating in the language of origin, the one or more processors further configured to pass the transliterated script to the TTS system, and if the TTS system is not operating in the language of origin, the one or more processors further configured to generate a phoneme sequence in the language of origin using a grapheme to phoneme (G2P) converter.

16. The system of claim 15 , wherein if the TTS system is not operating in the language of origin, mapping the phoneme sequence to a sequence of target language phonemes.

17. The system of claim 16 , wherein mapping includes generating a map of most likely unigram, bigram, and trigram mappings from the phoneme sequence to the sequence of target language phonemes.

18. The system of claim 15 , wherein identifying a word or phrase as a named entity includes one or more of table lookup and contextual analysis.

19. The system of claim 15 , wherein identifying a language of origin associated with the named entity includes one or more of table lookup and shortest distance measures to an existing names database.

20. The system of claim 15 , further comprising:

augmenting a text to speech dictionary based upon, at least in part, the phoneme sequence.

Assignments (8)
RELEASE (REEL 052935 / FRAME 0584) Recorded Jan 2, 2025
From: WELLS FARGO BANK, NATIONAL ASSOCIATION
To: CERENCE OPERATING COMPANY
Reel/Frame 069797/0818 →
CORRECTIVE ASSIGNMENT TO CORRECT THE REPLACE THE CONVEYANCE DOCUMENT WITH THE NEW ASSIGNMENT PREVIOUSLY RECORDED AT REEL: 050836 FRAME: 0191. ASSIGNOR(S) HEREBY CONFIRMS THE ASSIGNMENT. Recorded Apr 19, 2022
From: NUANCE COMMUNICATIONS, INC.
To: CERENCE OPERATING COMPANY
Reel/Frame 059804/0186 →
SECURITY AGREEMENT Recorded Jun 15, 2020
From: CERENCE OPERATING COMPANY
To: WELLS FARGO BANK, N.A.
Reel/Frame 052935/0584 →
RELEASE OF SECURITY INTEREST Recorded Jun 12, 2020
From: BARCLAYS BANK PLC
To: CERENCE OPERATING COMPANY
Reel/Frame 052927/0335 →
SECURITY AGREEMENT Recorded Nov 7, 2019
From: CERENCE OPERATING COMPANY
To: BARCLAYS BANK PLC
Reel/Frame 050953/0133 →
CORRECTIVE ASSIGNMENT TO CORRECT THE ASSIGNEE NAME PREVIOUSLY RECORDED AT REEL: 050836 FRAME: 0191. ASSIGNOR(S) HEREBY CONFIRMS THE INTELLECTUAL PROPERTY AGREEMENT. Recorded Oct 29, 2019
From: NUANCE COMMUNICATIONS, INC.
To: CERENCE OPERATING COMPANY
Reel/Frame 050871/0001 →
INTELLECTUAL PROPERTY AGREEMENT Recorded Oct 23, 2019
From: NUANCE COMMUNICATIONS, INC.
To: CERENCE INC.
Reel/Frame 050836/0191 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jun 8, 2015
From: SINGH, ANURAG RATAN; XU, YIFANG; SANCHEZ QUIJAS, IVAN A.; GODAVARTI, MAHESH
To: NUANCE COMMUNICATIONS, INC.
Reel/Frame 035856/0797 →