IP Library › Granted Patent US 11,094,311
Granted Patent B2
US 11,094,311 · App. 16/411,930 · Granted Aug 17, 2021

Speech synthesizing devices and methods for mimicking voices of public figures

Inventors: Brant Candelore (Escondido, CA); Mahyar Nejat (San Diego, CA)
Assignee: Sony Corporation
G10L13/033G10L13/00G10L13/047
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,094,311
App. No.
16/411,930
Granted
Aug 17, 2021
Kind
B2
Abstract

Speech synthesizing devices and methods are disclosed for mimicking the voices of public figures. A text-to-speech deep neural network (DNN) can be used to do so, with the DNN being trained using publicly available audio recordings of a given public figure speaking as well as text corresponding to the words that are spoken by the public figure in the audio recordings. The DNN may then be used to produce various audio outputs in the voice of the public figure.

Claims (33)

1. An apparatus, comprising:

at least one computer memory that is not a transitory signal and that comprises instructions executable by at least one processor to:

extract recorded speech of a celebrity from at least one piece of content that is publicly available;

analyze the recorded speech of the celebrity;

based on the analysis, configure an artificial intelligence model that can mimic the voice of the celebrity to output additional speech in the voice of the celebrity;

identify text from a television channel guide presented on a television display;

using the artificial intelligence model, convert the text from the television channel guide to speech in the voice of the celebrity to render an audible signal; and

play the audible signal on a playback device.

2. The apparatus of claim 1 , wherein the instructions are executable to:

analyze the recorded speech to train at least one neural network to mimic the voice of the celebrity, the artificial intelligence model comprising the at least one neural network.

3. The apparatus of claim 2 , wherein the at least one neural network is at least in part trained unsupervised.

4. The apparatus of claim 3 , wherein the at least one neural network is trained unsupervised at least in part using text that indicates words spoken by the celebrity in the recorded speech.

5. The apparatus of claim 4 , wherein the text is associated with closed captioning data corresponding to the recorded speech.

6. The apparatus of claim 4 , wherein the at least one neural network is trained unsupervised at least in part using the recorded speech of the celebrity.

7. The apparatus of claim 6 , wherein the recorded speech of the celebrity is extracted based on identification of the recorded speech as not including speech from other speakers during one or more segments of the recorded speech.

8. The apparatus of claim 2 , wherein the neural network creates a model of the celebrity which may be shared with other devices with text-to-speech engines.

9. The apparatus of claim 2 , wherein the at least one neural network is trained at least in part as supervised by a human, the at least one processor receiving an indication from the human that the recorded speech is that of the celebrity.

10. The apparatus of claim 1 , wherein the recorded speech is extracted from one or more of: a movie, a television show, other publicly available audio video (AV) content, a publicly available audio recording.

11. The apparatus of claim 1 , wherein the additional speech is output using text-to-speech software and text accessible to the at least one processor.

12. The apparatus of claim 1 , comprising the at least one processor, and comprising at least one speaker through which the additional speech is output.

13. A method, comprising:

analyzing, using a device, words spoken by a public figure;

based on the analysis, configuring a speech synthesizer to duplicate the public figure's voice for producing audio corresponding to text accessible to the device; and

producing the audio as a response to a user's query to a digital assistant.

14. The method of claim 13 , wherein the speech synthesizer uses a text-to-speech system to duplicate the public figure's voice.

15. The method of claim 14 , wherein the speech synthesizer is configured to employ a deep neural network (DNN) to produce the audio in the voice of the public figure, the DNN trained to the public figure's voice, the DNN establishing at least part of the text-to-speech system.

16. The method of claim 15 , comprising:

training the DNN using one or more audio recordings of the words spoken by the public figure and using text indicating the words spoken by the public figure.

17. An apparatus, comprising:

at least one computer readable storage medium that is not a transitory signal, the at least one computer readable storage medium comprising instructions executable by at least one processor to:

use a trained deep neural network (DNN) to produce a representation of a public figure's voice as speaking audio corresponding to first text that is either presented on an electronic display, second text from Closed Captioning, or that is to be used by a digital assistant as part of a response to a query, the trained DNN being trained using both audio of words spoken by the public figure and second text corresponding to the words, the first text being different from the second text.

18. The apparatus of claim 17 , wherein the apparatus is embodied in a server, and wherein the server executes the instructions.

19. The apparatus of claim 17 , wherein the apparatus is embodied as a consumer electronics device, and wherein the consumer electronic device comprises the electronics display.

Assignments (3)
CORRECTIVE ASSIGNMENT TO CORRECT THE APPLICATION NUMBER FROM 15/411,930 TO 16/411,930 PREVIOUSLY RECORDED ON REEL 049394 FRAME 0686. ASSIGNOR(S) HEREBY CONFIRMS THE ASSIGNMENT. Recorded Aug 20, 2019
From: CANDELORE, BRANT; NEJAT, MAHYAR
To: SONY CORPORATION
Reel/Frame 050110/0145 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jul 18, 2019
From: CANDELORE, BRANT; NEJAT, MAHYAR
To: SONY CORPORATION
Reel/Frame 049795/0576 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jun 6, 2019
From: CANDELORE, BRANT; NEJAT, MAHYAR
To: SONY CORPORATION
Reel/Frame 049394/0686 →
Continuity (1)
Related Publication 20200365136A1 · Nov 19, 2020
Cited By (26)
US 12,197,712 US 12,197,817 US 12,200,297 US 12,204,932 US 12,211,498 US 12,211,502 US 12,216,894 US 12,219,314 US 12,236,952 US 12,260,234 US 12,277,954 US 12,293,763 US 12,301,635 US 12,333,404 US 12,361,943 US 12,367,879 US 12,380,281 US 12,386,434 US 12,386,491 US 12,431,128 US 12,477,470 US 12,556,890 US 12,608,171 US 12,613,730 US 12,619,452 US 12,748,568