IP Library Granted Patent US 12,380,875
Granted Patent B2
US 12,380,875 · App. 17/842,986 · Granted Aug 5, 2025

Speech synthesis with foreign fragments

Inventors: Corinne Bos-Plachez (Baisieux, FR); Vito Quinci (Turin, IT); Alina Lenhardt (Ulm, DE); Benjamin Vincent Marcel Picart (Merelbeke, BE); Martine Marguerite Staessen (Wervik, BE); Athos Toniolo (Turin, IT)
Assignee: Cerence Operating Company
G10L13/086G06F40/263G06F40/279G10L13/047G10L13/06
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,380,875
App. No.
17/842,986
Granted
Aug 5, 2025
Kind
B2
Abstract

A method for synthesizing speech from a textual input includes receiving the textual input, the textual input including native words in a native language and foreign words in a foreign language, and processing the textual input to determine a phonetic representation of the textual input. The processing includes determining a native phonetic representation of the of the native words, and determining a nativized phonetic representation of the foreign words. Determining the nativized phonetic representation includes forming a foreign phonetic representation of the foreign words using a foreign phoneme set, and mapping the foreign phonetic representation to the nativized phonetic representation according to a model of a native speaker's pronunciation of foreign words.

Claims (53)

1. A method for synthesizing speech from a textual input, the method comprising:

receiving the textual input, the textual input including native words in a native language and foreign words in a foreign language;

processing the textual input to determine a phonetic representation of the textual input, the processing including:

determining a native phonetic representation of the of the native words, and

determining a nativized phonetic representation of the foreign words, determining the nativized phonetic representation including:

forming a foreign phonetic representation of the foreign words using a foreign phoneme set, and

mapping the foreign phonetic representation to the nativized phonetic representation according to a model of a native speaker's pronunciation of foreign words;

providing a combination of the native phonetic representation of the native words and the nativized phonetic representation of the foreign words to a waveform synthesizer for synthesis of a speech waveform; and

providing the speech waveform for acoustic presentation to a user by a loudspeaker;

wherein the nativized phonetic representation comprises a combination of phonemes from a native phoneme set and phonemes from the foreign phoneme set.

2. The method of claim 1 wherein the native phonetic representation comprises phonemes from a native phoneme set.

3. The method of claim 1 wherein the mapping of the foreign phonetic representation to the nativized phonetic representation uses contextual information associated with the foreign phonetic representation, and wherein the contextual information includes a grapheme representation of the foreign text associated with the foreign phonetic representation.

4. The method of claim 3 wherein the contextual information further includes an alignment of graphemes in the grapheme representation of the foreign text to phonemes in the foreign phonetic representation.

5. The method of claim 3 wherein the contextual information includes location information of phonemes in the foreign phonetic representation.

6. The method of claim 1 wherein the model of the native speaker's pronunciation of foreign words is based on training data comprising foreign textual phrases and phonetic transcriptions of a native speaker's pronunciation of the foreign textual phrases.

7. The method of claim 6 wherein at least some of the native speaker's pronunciations of foreign textual phrases are mispronunciations.

8. The method of claim 1 further comprising synthesizing the speech waveform based on the combination of the native phonetic representation of the of the native words and the nativized phonetic representation of the foreign words using the waveform synthesizer.

9. The method of claim 1 further comprising configuring a neural network according to the model of a native speaker's pronunciation of foreign words.

10. The method of claim 1 wherein mapping the foreign phonetic representation to the nativized phonetic representation according to the model of a native speaker's pronunciation of foreign words includes applying one or more mapping rules.

11. The method of claim 1 further comprising identifying the foreign words in the textual input.

12. A system for synthesizing speech from a textual input, the system comprising:

an input for receiving the textual input, the textual input including native words in a native language and foreign words in a foreign language;

one or more processors configured to process the textual input to determine a phonetic representation of the textual input, the processing including:

determining a native phonetic representation of the of the native words, and

determining a nativized phonetic representation of the foreign words, determining the nativized phonetic representation including:

forming a foreign phonetic representation of the foreign words using a foreign phoneme set, and

mapping the foreign phonetic representation to the nativized phonetic representation according to a model of a native speaker's pronunciation of foreign words;

providing a combination of the native phonetic representation of the native words and the nativized phonetic representation of the foreign words to a waveform synthesizer for synthesis of a speech waveform; and

providing the speech waveform for acoustic presentation to a user by a loudspeaker;

wherein the nativized phonetic representation comprises a combination of phonemes from a native phoneme set and phonemes from the foreign phoneme set.

13. Software stored in a non-transitory form on a computer-readable medium, the software including instructions for causing a computing system to synthesize speech from a textual input including:

receive the textual input, the textual input including native words in a native language and foreign words in a foreign language;

process the textual input to determine a phonetic representation of the textual input, the processing including:

determining a native phonetic representation of the of the native words, and

determining a nativized phonetic representation of the foreign words, determining the nativized phonetic representation including:

forming a foreign phonetic representation of the foreign words using a foreign phoneme set, and

mapping the foreign phonetic representation to the nativized phonetic representation according to a model of a native speaker's pronunciation of foreign words;

providing a combination of the native phonetic representation of the native words and the nativized phonetic representation of the foreign words to a waveform synthesizer for synthesis of a speech waveform; and

providing the speech waveform for acoustic presentation to a user by a loudspeaker;

wherein the nativized phonetic representation comprises a combination of phonemes from a native phoneme set and further comprises phonemes from the foreign phoneme set a second set of phonemes different from the native phoneme set.

14. A method for synthesizing speech from a textual input, the method comprising:

receiving the textual input, the textual input including native words in a native language and foreign words in a foreign language;

processing the textual input to determine a phonetic representation of the textual input, the processing including

determining a native phonetic representation of the of the native words, and

determining a nativized phonetic representation of the foreign words including

forming a foreign phonetic representation of the foreign words using a foreign phoneme set, and

mapping the foreign phonetic representation to the nativized phonetic representation according to a model of a native speaker's pronunciation of foreign words,

providing a combination of the native phonetic representation of the native words and the nativized phonetic representation of the foreign words to a waveform synthesizer for synthesis of a speech waveform; and

providing the speech waveform for acoustic presentation to a user by a loudspeaker;

wherein the nativized phonetic representation comprises a combination of phonemes from a native phoneme set and phonemes from the foreign phoneme set, the mapping of the foreign phonetic representation to the nativized phonetic representation uses contextual information associated with the foreign phonetic representation, and the contextual information includes a grapheme representation of the foreign text associated with the foreign phonetic representation.

15. The method of claim 1 , wherein the model is a trained neural network sequence-to-sequence model, and the mapping of the foreign phonetic representation comprises providing said representation to an input of the neural network model, and obtaining the nativized phonetic representation as an output of the neural network model.

16. The method of claim 6 , wherein the model is a trained neural network sequence-to- sequence model, and the method further comprises training said model using the training data.

17. The method of claim 1 , wherein the mapping of the foreign phonetic representation is performed according to user preference for a degree of nativization.

Assignments (3)
RELEASE (REEL 067417 / FRAME 0303) Recorded Jan 2, 2025
From: WELLS FARGO BANK, NATIONAL ASSOCIATION
To: CERENCE OPERATING COMPANY
Reel/Frame 069797/0422 →
SECURITY AGREEMENT Recorded Apr 15, 2024
From: CERENCE OPERATING COMPANY
To: WELLS FARGO BANK, N.A., AS COLLATERAL AGENT
Reel/Frame 067417/0303 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jul 15, 2022
From: BOS-PLACHEZ, CORINNE; QUINCI, VITO; LENHARDT, ALINA; PICART, BENJAMIN VINCENT MARCEL; STAESSEN, MARTINE MARGUERITE; TONIOLO, ATHOS
To: CERENCE OPERATING COMPANY
Reel/Frame 060514/0818 →