IP Library Granted Patent US 11,948,555
Granted Patent B2
US 11,948,555 · App. 17/441,220 · Granted Apr 2, 2024

Method and system for content internationalization and localization

Inventors: Mark Christie (Auckley, GB); Gerald Chao (Los Angeles, CA)
Assignee: NEP SUPERSHOOTERS L.P.
G10L15/07G10L15/005G10L15/063G10L15/083H04N21/2335H04N21/23418H04N21/234345H04N21/8106
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,948,555
App. No.
17/441,220
Granted
Apr 2, 2024
Kind
B2
Abstract

A method of processing a video file to generate a modified video file, the modified video file including a translated audio content of the video file, the method comprising: receiving the video file; accessing a facial model or a speech model for a specific speaker, wherein the facial model maps speech to facial expressions, and the speech model maps text to speech; receiving a reference content for the originating video file for the specific speaker; generating modified audio content for the specific speaker and/or modified facial expression for the specific speaker; and modifying the video file in accordance with the modified content and/or the modified expression to generate the modified video file.

Claims (55)

1. A method of processing an original video file to generate a modified video file, the modified video file including a translated audio content of the original video file, the method comprising:

receiving the original video file for processing;

receiving a second video file of a different speaker than the speaker in the original video file, wherein the second video file is a video of a different speaker stating speech expressions;

accessing a model associating facial characteristics of the speaker in the original video file with portions of speech for those portions of replaced audio content; and

replacing facial expressions of the speaker in the original video file with facial expressions according to the video of the different speaker, on determination of a facial expression of the speaker matching a facial expression in the model.

2. The method of claim 1 , further comprising generating a model for use in processing a video file to generate a version of the original video file with translated audio content, the method further comprising:

identifying a specific speaker in the original video file;

obtaining speech samples of the identified specific speaker;

converting each speech sample into a portion of text; and

storing an association, for at least one speaker in the original video file, of speech sample to text.

3. The method of claim 2 further comprising the step of training the model, via at least one machine learning algorithm, to associate at least one speaker's voice with each speech sample of the at least one speaker.

4. The method of claim 3 wherein each speech sample of the speaker is spoken text.

5. The method of claim 1 , wherein the modified video file includes a modified audio content of the original video file, the method further comprising:

processing the received video file in dependence on the model created according to the method of claim 1 for the at least one speaker.

6. The method of claim 1 , further comprising generating the model for use in processing a video file to generate a version of the original video file with translated audio content, the method comprising:

identifying a specific speaker in the original video file;

determining an appearance of the specific speaker in the original video file;

obtaining speech samples of the identified speaker; and

storing an association, for at least one speaker in the original video file, of the speaker appearance for each speech sample of the at least one speaker.

7. The method of claim 6 further comprising the step of training the model, via at least one machine learning algorithm, to associate at least one speaker's appearance to each speech sample of the speaker.

8. The method according to claim 6 wherein the step of determining an appearance of the specific speaker comprises capturing a facial expression of the specific speaker.

9. The method of claim 1 , wherein the modified video file includes a modified audio content of the video file, the method comprising:

processing the received video file in dependence on the model created according to the method of claim 6 .

10. The method of claim 1 , the method further comprising:

accessing a facial model for a specific speaker, wherein the facial model maps speech to facial expressions;

receiving a reference content for the original video file for the specific speaker;

generating modified facial expression for the specific speaker; and

modifying the video file in accordance with the modified expression to generate the modified video file.

11. A method of claim 1 , the method further comprising:

receiving a translated dialogues in text format of the original video file for a speaker in the video file;

accessing a model associating speech of the speaker with portions of text; and

replacing audio content in the video file with generated speech or translated dialog in accordance with the received model.

12. The method of claim 1 , the method further comprising:

receiving a translated dialogues in text format of the original video file for a speaker in the original video file;

receiving a model associating speech of the speaker in the original video file with portions of text;

replacing audio content in the original video file with generated speech in accordance with the received model;

accessing a model associating facial characteristics of the speaker in the original video file with portions of speech for those portions of replaced audio content; and

replacing facial characteristics of the speaker in the original video file in accordance with the received model.

13. The method of claim 1 further comprising receiving a dubbed speech file for a speaker in the original video file spoken by a voice actor.

14. The method of claim 13 wherein the replaced audio content is of the voice actor.

15. The method of claim 1 wherein the replaced audio content is of a voice actor.

16. The method of claim 1 , the method further comprising:

receiving a translated dialogue in text format of the audio in the original video file for the speaker in the video file;

receiving a model associating speech of the speaker in the original video file with portions of text;

replacing audio content in the original video file with translated dialog in accordance with the received model;

receiving the model associating facial characteristics of the speaker in the original video file with portions of speech for those portions of replaced audio content; and

replacing facial characteristics of the speaker in the original video file in accordance with the received model.

17. The method of claim 16 further comprising the step of receiving a dubbed speech file for a speaker in the original video file spoken by a voice speaker.

18. The method of claim 1 , further comprising:

accessing a speech model for a specific speaker, wherein the speech model maps text to speech;

receiving a reference content for the original video file for the specific speaker;

generating modified audio content for the specific speaker; and

modifying the video file in accordance with the modified content and to generate the modified video file.

19. The method of claim 1 further comprising generating a model of facial characteristics of the different speaker stating speech expressions.

20. The method of claim 1 wherein the speech in the video file is translated and matched to speech in the model of the different speaker, wherein the facial characteristics associated with the different speaker for that speech is used to replace the facial expressions of the speaker in the original video file with a facial expression of that speaker which matches the facial expression of the different speaker.

Assignments (4)
SECURITY INTEREST Recorded Oct 31, 2025
From: NEP SUPERSHOOTERS, LP
To: BARCLAYS BANK PLC, AS COLLATERAL AGENT
Reel/Frame 072752/0067 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Dec 6, 2023
From: CHRISTIE, MARK; CHAO, GERALD
To: PIKSEL, INC.
Reel/Frame 065785/0054 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded May 15, 2023
From: PJR HOLDING COMPANY LLC
To: NEP SUPERSHOOTERS L.P.
Reel/Frame 063637/0532 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jun 6, 2022
From: PIKSEL, INC.
To: PRJ HOLDING COMPANY, LLC
Reel/Frame 060703/0956 →
Continuity (2)
Provisional Application 62821274 · Mar 20, 2019
Related Publication 20220172709A1 · Jun 2, 2022
Cited By (1)
US 12,229,313