IP Library Granted Patent US 12,505,859
Granted Patent B2
US 12,505,859 · App. 17/877,561 · Granted Dec 23, 2025

System and method for generating video in target language

Inventors: Paloma de Juan (Newy York, NY); Alex J. Shaw (New York, NY); Eric M. Dodds (Berkeley, CA); Benjamin J. Culpepper (Berkeley, CA); Kapil Raj Thadani (New York, NY); Lakshmi V. Kesiraju (San Jose, CA); Praveen Mareedu (Jersey City, NJ); Sanika Shirwadkar (Milpitas, CA); Xingyue Zhou (Mountain View, CA); Yueh-Ning Ku (Sunnyvale, CA)
Assignee: Yahoo Assets LLC
G11B27/031G06F40/58G06V20/40G06V40/161G06V40/172G06V40/50G10L13/08G10L15/26G10L21/043G10L21/055G10L25/57
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,505,859
App. No.
17/877,561
Granted
Dec 23, 2025
Kind
B2
Abstract

One or more computing devices, systems, and/or methods for generating a video in a target language are provided. In an example, a first video, in which a first speaker speaks in a first language, is identified. A translated transcript in a second language is determined. The translated transcript is indicative of a translation of speech spoken by the first speaker in the first video. Based upon the translated transcript and a speaker profile associated with a second speaker, first audio, including an auditory representation of the translated transcript being spoken in a voice of the second speaker, is generated. Based upon the first video and the first audio, a second video, in which mouth movements of the first speaker are aligned with speech of the auditory representation of the first audio, is generated.

Claims (98)

1 . A method, comprising:

identifying a first video;

determining a transcript indicative of speech spoken by a speaker in the first video;

translating the transcript from a first language to a second language to generate a translated transcript in the second language;

determining a speaker profile associated with the speaker, the speaker profile generated based upon second speech of the speaker from a first source of audio identified as being associated with the speaker during analysis of a first internet resource and third speech of the speaker from a second source of audio identified as being associated with the speaker during analysis of a second internet resource;

generating, based upon the translated transcript and the speaker profile associated with the speaker, first audio comprising an auditory representation of the translated transcript being spoken in a voice of the speaker; and

generating, based upon the first video and the first audio, a second video in which mouth movements of the speaker are aligned with speech of the auditory representation of the first audio, wherein generating the second video comprises aligning each audio segment of a plurality of audio segments of the second video with a corresponding video segment of the second video.

2 . The method of claim 1 , comprising:

presenting the second video via a client device, wherein the presenting the second video comprises:

displaying the second video via a graphical user interface; and

outputting the first audio via a loudspeaker.

3 . The method of claim 1 , wherein:

the transcript comprises a plurality of text segments;

a first text segment of the plurality of text segments is associated with a first video segment of the first video;

a second text segment of the plurality of text segments is associated with a second video segment of the first video;

the translated transcript comprises a plurality of translated text segments;

a first translated text segment of the plurality of translated text segments is a translation of the first text segment; and

a second translated text segment of the plurality of translated text segments is a translation of the second text segment.

4 . The method of claim 3 , wherein:

the first text segment corresponds to one or more first sentences; and

the second text segment corresponds to one or more second sentences.

5 . The method of claim 3 , wherein:

the first audio comprises a plurality of audio segments; and

the generating the first audio comprises:

generating a first audio segment of the plurality of audio segments of the first audio based upon the first translated text segment and a duration of time of the first video segment associated with the first text segment, wherein:

the first audio segment comprises an auditory representation, of the first translated text segment, in the voice of the speaker; and

a duration of time of the first audio segment matches the duration of time of the first video segment associated with the first text segment; and

generating a second audio segment of the plurality of audio segments of the first audio based upon the second translated text segment and a duration of time of the second video segment associated with the second text segment, wherein:

the second audio segment comprises an auditory representation, of the second translated text segment, in the voice of the speaker; and

a duration of time of the second audio segment matches the duration of time of the second video segment associated with the second text segment.

6 . The method of claim 5 , wherein at least one of:

the generating the first audio segment comprises:

generating, based upon the first translated text segment, a third audio segment comprising an auditory representation, of the first translated text segment, in the voice of the speaker, wherein a duration of time of the third audio segment is longer than the duration of time of the first video segment associated with the first text segment; and

shortening the third audio segment to generate the first audio segment of the plurality of audio segments of the first audio; or

the generating the second audio segment comprises:

generating, based upon the second translated text segment, a fourth audio segment comprising an auditory representation, of the second translated text segment, in the voice of the speaker, wherein a duration of time of the fourth audio segment is shorter than the duration of time of the second video segment associated with the second text segment; and

lengthening the fourth audio segment to generate the second audio segment of the plurality of audio segments of the first audio.

7 . The method of claim 6 , wherein at least one of:

the shortening the third audio segment comprises increasing a speed of audio of the third audio segment; or

the lengthening the fourth audio segment comprises padding the fourth audio segment.

8 . The method of claim 1 , wherein the generating the second video comprises:

identifying a face of the speaker in the first video; and

modifying, based upon the first audio, pixels of the first video to generate the second video in which mouth movements, of the speaker, are aligned with the speech of the first audio, wherein the pixels of the first video that are modified correspond to a portion, of the face, comprising a mouth of the face.

9 . The method of claim 8 , wherein:

the generating the second video comprises aligning a first audio segment of the plurality of audio segments of the second video with a corresponding first video segment of the second video by lengthening the first audio segment by padding the first audio segment with silence that was not in a corresponding first audio segment in the first video.

10 . The method of claim 8 , comprising:

generating a face profile associated with the speaker based upon one or more images comprising a face of the speaker, wherein the identifying the face of the speaker in the first video comprises performing, based upon the face profile, facial recognition on the first video.

11 . The method of claim 1 , comprising:

generating the speaker profile associated with the speaker based upon audio of the first video.

12 . The method of claim 1 , wherein:

the speaker profile comprises a vector representation generated based upon the second speech of the speaker.

13 . A computing device comprising:

a processor; and

memory comprising processor-executable instructions that when executed by the processor cause performance of operations, the operations comprising:

identifying a first video;

determining a transcript indicative of speech spoken by a speaker in the first video;

translating the transcript from a first language to a second language to generate a translated transcript in the second language;

determining a speaker profile associated with the speaker, the speaker profile generated based upon second speech of the speaker from a first source of audio identified as being associated with the speaker during analysis of a first internet resource and third speech of the speaker from a second source of audio identified as being associated with the speaker during analysis of a second internet resource;

generating, based upon the translated transcript and the speaker profile associated with the speaker, first audio comprising an auditory representation of the translated transcript being spoken in a voice of the speaker; and

generating, based upon the first video and the first audio, a second video in which mouth movements of the speaker are aligned with speech of the auditory representation of the first audio, wherein generating the second video comprises aligning each audio segment of a plurality of audio segments of the second video with a corresponding video segment of the second video.

14 . The computing device of claim 13 , the operations comprising:

presenting the second video via a client device, wherein the presenting the second video comprises:

displaying the second video via a graphical user interface; and

outputting the first audio via a loudspeaker.

15 . The computing device of claim 13 , wherein:

the transcript comprises a plurality of text segments;

a first text segment of the plurality of text segments is associated with a first video segment of the first video;

a second text segment of the plurality of text segments is associated with a second video segment of the first video;

the translated transcript comprises a plurality of translated text segments;

a first translated text segment of the plurality of translated text segments is a translation of the first text segment; and

a second translated text segment of the plurality of translated text segments is a translation of the second text segment.

16 . The computing device of claim 15 , wherein:

the first audio comprises a plurality of audio segments; and

the generating the first audio comprises:

generating a first audio segment of the plurality of audio segments of the first audio based upon the first translated text segment and a duration of time of the first video segment associated with the first text segment, wherein:

the first audio segment comprises an auditory representation, of the first translated text segment, in the voice of the speaker; and

a duration of time of the first audio segment matches the duration of time of the first video segment associated with the first text segment; and

generating a second audio segment of the plurality of audio segments of the first audio based upon the second translated text segment and a duration of time of the second video segment associated with the second text segment, wherein:

the second audio segment comprises an auditory representation, of the second translated text segment, in the voice of the speaker; and

a duration of time of the second audio segment matches the duration of time of the second video segment associated with the second text segment.

17 . The computing device of claim 16 , wherein at least one of:

the generating the first audio segment comprises:

generating, based upon the first translated text segment, a third audio segment comprising an auditory representation, of the first translated text segment, in the voice of the speaker, wherein a duration of time of the third audio segment is longer than the duration of time of the first video segment associated with the first text segment; and

shortening the third audio segment to generate the first audio segment of the plurality of audio segments of the first audio; or

the generating the second audio segment comprises:

generating, based upon the second translated text segment, a fourth audio segment comprising an auditory representation, of the second translated text segment, in the voice of the speaker, wherein a duration of time of the fourth audio segment is shorter than the duration of time of the second video segment associated with the second text segment; and

lengthening the fourth audio segment to generate the second audio segment of the plurality of audio segments of the first audio.

18 . The computing device of claim 16 , wherein the generating the second video comprises:

generating a third video segment, of the second video, based upon the first audio segment of the plurality of audio segments of the first audio and the first video segment of the first video; and

generating a fourth video segment, of the second video, based upon the second audio segment of the plurality of audio segments of the first audio and the second video segment of the first video.

19 . A non-transitory machine readable medium having stored thereon processor-executable instructions that when executed cause performance of operations, the operations comprising:

identifying a first video in which a first speaker speaks in a first language;

determining a translated transcript, in a second language, indicative of a translation of speech spoken by the first speaker in the first video;

determining a speaker profile associated with a second speaker, the speaker profile generated based upon second speech of the second speaker from a first source of audio identified as being associated with the second speaker during analysis of a first internet resource and third speech of the second speaker from a second source of audio identified as being associated with the second speaker during analysis of a second internet resource;

generating, based upon the translated transcript and the speaker profile associated with the second speaker, first audio comprising an auditory representation of the translated transcript being spoken in a voice of the second speaker; and

generating, based upon the first video and the first audio, a second video in which mouth movements of the first speaker are aligned with speech of the auditory representation of the first audio.

20 . The non-transitory machine readable medium of claim 19 , wherein:

the second speaker is the same as the first speaker.

Assignments (2)
SUPPLEMENTAL PATENT SECURITY AGREEMENT Recorded Sep 17, 2025
From: YAHOO ASSETS LLC
To: ROYAL BANK OF CANADA, AS COLLATERAL AGENT
Reel/Frame 072915/0540 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jul 29, 2022
From: DE JUAN, PALOMA; SHAW, ALEX J.; DODDS, ERIC M.; CULPEPPER, BENJAMIN J.; THADANI, KAPIL RAJ; KESIRAJU, LAKSHMI V.; MAREEDU, PRAVEEN; SHIRWADKAR, SANIKA; ZHOU, XINGYUE; KU, YUEH-NING
To: YAHOO ASSETS LLC
Reel/Frame 060676/0185 →
Continuity (1)
Related Publication 20240038271A1 · Feb 1, 2024
References Cited (15)
US 8620139B2 · Li · 2013 [cited by examiner]
US 10887672B1 · Wu · 2021 [cited by examiner]
US 11043230B1 · Riding · 2021 [cited by examiner]
US 11195507B2 · Kumar · 2021 [cited by examiner]
US 20190244623A1 · Hall · 2019 [cited by examiner]
US 20210065712A1 · Holm · 2021 [cited by examiner]
US 20210136200A1 · Li · 2021 [cited by examiner]
US 20210398541A1 · Aher · 2021 [cited by examiner]
US 20230290332A1 · Jawahar · 2023 [cited by examiner]
US 20230325611A1 · Garg · 2023 [cited by examiner]
Prajwal, K.R., et al.: “A Lip Sync Expert Is All You Need for Speech to Lip Generation in the Wild”, MM '20 Oct. 12-16, 2020, Seattle, WA USA, https://arxiv.org/pdf/2003.00418.pdf, 10 pages. [cited by applicant]
Prajwal K R., et al.: “Towards Automatic Face-to-Face Translation”, MM '19, Oct. 21-25, 2019, Nice, France, https://arxiv.org/pdf/2003.00418.pdf, 9 pages. [cited by applicant]
Edresson, Casanova et al, Your TTS: Towards Zero-Shot Multip-Speaker TTS and Zero Shot Voice Conversion for Everyone, Dec. 2021, https://github.com/Edresson/YourTTS/; 6 pages, Retrieved on Oct. 17, 2022. [cited by applicant]
Prajwal, K.R., MukhopadhyahRudrabha, Namboodiri, Vinay P, Jawahar, C.V.: “A Lip Sync Expert is All you Need for Speech to Lip Generation in the Wild”, https://github.com/Rudrabha/Wav2Lip, ACM Multimedia 2020, 7 pages, R… [cited by applicant]
Helsinki—“NLP Language Technology Research Group at the University of Helsinki”, https://huggingface.co/Helsinki-NLP, 122 pages, Retrieved on Oct. 17, 2022. [cited by applicant]