IP Library Granted Patent US 12,236,935
Granted Patent B2
US 12,236,935 · App. 17/931,026 · Granted Feb 25, 2025

Generating dubbed audio from a video-based source

Inventors: Andrew R. Levine (New York, NY); Buddhika Kottahachchi (San Mateo, CA); Christopher Davie (Queens, NY); Kulumani Sriram (Danville, CA); Richard James Potts (Mountain View, CA); Sasakthi S. Abeysinghe (Santa Clara, CA)
Assignee: GOOGLE LLC
G10L13/02G06F40/58G10L13/086
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,236,935
App. No.
17/931,026
Granted
Feb 25, 2025
Kind
B2
Abstract

The present disclosure relates to generating and adjusting translated audio from a video-based source. The method includes receiving video data and corresponding audio data in a first language; generating a translated preliminary transcript in a second language; aligning timing windows of portions of the translated preliminary transcript with corresponding segments of the audio data; determining portions of the translated aligned transcript in the second language that exceed a timing window range of the corresponding segments of the audio data in the first language to generate flagged transcript portions; transmitting the original transcript, the translated aligned transcript, and the first speech dub to a first device, the generated flagged transcript portions included in the original transcript and the translated aligned transcript; receiving, from the first device, a modified original transcript; and generating, based on the modified original transcript, a second speech dub in the second language.

Claims (61)

1. A method of dubbing a video, comprising:

receiving video data and corresponding audio data in a first language;

generating, based on the audio data and an original transcript in the first language, a translated preliminary transcript in a second language;

based on the video data in the first language, aligning timing windows of portions of the translated preliminary transcript with corresponding segments of the audio data in the first language to generate a translated aligned transcript;

based on the timing windows of the portions of the translated preliminary transcript and timing windows of the corresponding segments of the audio data in the first language, determining portions of the translated aligned transcript in the second language that exceed a timing window range of the corresponding segments of the audio data in the first language to generate flagged transcript portions;

based on the translated aligned transcript, generating a first speech dub in the second language and combining the first speech dub with the video data to generate a first dubbed video;

transmitting the original transcript, the translated aligned transcript, and the first speech dub to a first device, the generated flagged transcript portions included in the original transcript and the translated aligned transcript;

receiving, from the first device, a modified original transcript; and

generating, based on the modified original transcript, a second speech dub in the second language.

2. The method of claim 1 , further comprising

combining the second speech dub in the second language with the video data excluding the audio data in the first language to generate dubbed a second dubbed video; and

outputting the second dubbed video via a user device.

3. The method of claim 1 , wherein the transmitting the original transcript, the translated aligned transcript, and the first speech dub to the first device further comprises transmitting the original transcript, the translated aligned transcript, and the first speech dub to the first device to be displayed.

4. The method of claim 3 , wherein the transmitting the original transcript, the translated aligned transcript, and the first speech dub to the first device further comprises displaying the original transcript with the flagged transcript portions and the translated aligned transcript with the flagged transcript portions at a corresponding location in the translated aligned transcript as the original transcript.

5. The method of claim 4 , wherein

the flagged transcript portions include text corresponding to portions of the first speech dub that have a timing adjustment applied, and

the transmitting the original transcript, the translated aligned transcript, and the first speech dub to the first device further comprises applying a formatting to the text corresponding to portions of the first speech dub that have a timing adjustment applied.

6. The method of claim 5 , wherein the formatting applied to the text is based on an amount of the timing adjustment applied, the formatting being more visible as the amount of the timing adjustment applied increases.

7. The method of claim 1 , wherein the determining the portions of the translated aligned transcript in the second language that do not fit within the timing window range of the corresponding segments of the audio data further comprises applying a timing adjustment to the portions of the translated aligned transcript determined to not fit within the timing window range of the corresponding segments of the audio data, the timing adjustment not exceeding a predetermined maximum speedup rate.

8. The method of claim 7 , wherein neighboring portions of the translated aligned transcript determined to exceed the timing window range of the corresponding segments of the audio data are merged together.

9. The method of claim 1 , further comprising

assigning a confidence score to words in the original transcript;

analyzing frames of the video data at times corresponding to words in the original transcript with a low confidence score to detect relevant text, symbols, and mouth movements of a human in the frames of the video data;

generating a replacement word based on the detected relevant text, symbols, and mouth movements of the human; and

replacing the word having the low confidence score with the replacement word.

10. The method of claim 1 , further comprising, after the determining portions of the translated aligned transcript in the second language that exceed a timing window range of the corresponding segments of the audio data in the first language to generate flagged transcript portions, automatically re-aligning timing windows, merging timing windows, and adjusting a speed of the flagged transcript portions while maintaining a pitch of the first speech dub.

11. A non-transitory computer-readable storage medium for storing computer-readable instructions that, when executed by a computer, cause the computer to perform a method, the method comprising:

receiving video data and corresponding audio data in a first language;

generating, based on the audio data and an original transcript in the first language, a translated preliminary transcript in a second language;

based on the video data in the first language, aligning timing windows of portions of the translated preliminary transcript with corresponding segments of the audio data in the first language to generate a translated aligned transcript;

based on the timing windows of the portions of the translated preliminary transcript and timing windows of the corresponding segments of the audio data in the first language, determining portions of the translated aligned transcript in the second language that exceed a timing window range of the corresponding segments of the audio data in the first language to generate flagged transcript portions;

based on the translated aligned transcript, generating a first speech dub in the second language and combining the first speech dub with the video data to generate a first dubbed video;

transmitting the original transcript, the translated aligned transcript, and the first speech dub to a first device, the generated flagged transcript portions included in the original transcript and the translated aligned transcript;

receiving, from the first device, a modified original transcript; and

generating, based on the modified original transcript, a second speech dub in the second language.

12. The non-transitory computer-readable storage medium of claim 11 , further comprising

combining the second speech dub in the second language with the video data excluding the audio data in the first language to generate a second dubbed video; and

outputting the second dubbed video via a user device.

13. The non-transitory computer-readable storage medium of claim 11 , wherein the transmitting the original transcript, the translated aligned transcript, and the first speech dub to the first device further comprises transmitting the original transcript, the translated aligned transcript, and the first speech dub to the first device to be displayed.

14. The non-transitory computer-readable storage medium of claim 13 , wherein the transmitting the original transcript, the translated aligned transcript, and the first speech dub to the first device further comprises displaying the original transcript with the flagged transcript portions and the translated aligned transcript with the flagged transcript portions at a corresponding location in the translated aligned transcript as the original transcript.

15. The non-transitory computer-readable storage medium of claim 14 , wherein

the flagged transcript portions include text corresponding to portions of the first speech dub that have a timing adjustment applied, and

the transmitting the original transcript, the translated aligned transcript, and the first speech dub to the first device further comprises applying a formatting to the text corresponding to portions of the first speech dub that have a timing adjustment applied.

16. The non-transitory computer-readable storage medium of claim 15 , wherein the formatting applied to the text is based on an amount of the timing adjustment applied, the formatting being more visible as the amount of the timing adjustment applied increases.

17. The non-transitory computer-readable storage medium of claim 11 , wherein the determining the portions of the translated aligned transcript in the second language that do not fit within the timing window range of the corresponding segments of the audio data further comprises applying a timing adjustment to the portions of the translated aligned transcript determined to not fit within the timing window range of the corresponding segments of the audio data, the timing adjustment not exceeding a predetermined maximum speedup rate.

18. The non-transitory computer-readable storage medium of claim 11 , wherein neighboring portions of the translated aligned transcript determined to exceed the timing window range of the corresponding segments of the audio data are merged together.

19. The non-transitory computer-readable storage medium of claim 11 , further comprising

assigning a confidence score to words in the original transcript;

analyzing frames of the video data at times corresponding to words in the original transcript with a low confidence score to detect relevant text, symbols, and mouth movements of a human in the frames of the video data;

generating a replacement word based on the detected relevant text, symbols, and mouth movements of the human; and

replacing the word having the low confidence score with the replacement word.

20. An apparatus for dubbing a video, comprising:

processing circuitry configured to

receive video data and corresponding audio data in a first language;

generate, based on the audio data and an original transcript in the first language, a translated preliminary transcript in a second language;

based on the video data in the first language, align timing windows of portions of the translated preliminary transcript with corresponding segments of the audio data in the first language to generate a translated aligned transcript;

based on the timing windows of the portions of the translated preliminary transcript and timing windows of the corresponding segments of the audio data in the first language, determine portions of the translated aligned transcript in the second language that exceed a timing window range of the corresponding segments of the audio data in the first language to generate flagged transcript portions;

based on the translated aligned transcript, generate a first speech dub in the second language and combine the first speech dub with the video data to generate a first dubbed video;

transmit the original transcript, the translated aligned transcript, and the first speech dub to a first device, the generated flagged transcript portions included in the original transcript and the translated aligned transcript;

receive, from the first device, a modified original transcript; and

generate, based on the modified original transcript, a second speech dub in the second language.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Sep 9, 2022
From: LEVINE, ANDREW R.; KOTTAHACHCHI, BUDDHIKA; DAVIE, CHRISTOPHER; SRIRAM, KULUMANI; POTTS, RICHARD JAMES; ABEYSINGHE, SASAKTHI S.
To: GOOGLE LLC
Reel/Frame 061051/0601 →
Continuity (1)
Related Publication 20240087557A1 · Mar 14, 2024
References Cited (10)
US 20080092047A1 · Fealkoff · 2008 [cited by examiner]
US 20140039871A1 · Crawford · 2014 [cited by examiner]
US 20200118582A1 · Garland · 2020 [cited by examiner]
US 20230025800A1 · Werfelli · 2023 [cited by examiner]
US 20240038271A1 · de Juan · 2024 [cited by examiner]
“Automatically dub your videos to 20+ foreign languages.”, Maestra, downloaded Sep. 9, 2022, https://maestrahelp.zendesk.com/hc/en-us/articles/360053132452-Automatically-dub-your-videos-to-20-foreign-languages. [cited by applicant]
“Automatically add subtitles to your videos, and translate automatically to foreign languages.”, Maestra, downloaded Sep. 9, 2022, https://maestrahelp.zendesk.com/hc/en-us/articles/360053373931-Automatically-add-subtitl… [cited by applicant]
“Automatically transcribe your audio & video files to text online.”, Maestra, downloaded Sep. 9, 2022, https://maestrahelp.zendesk.com/hc/en-us/articles/360052713852-Automatically-transcribe-your-audio-video-files-to-te… [cited by applicant]
“Automatic Video Dubbing and Voiceover Services”, Maestra, downloaded Sep. 9, 2022, https://maestrasuite.com/video-dubber#how-it-works. [cited by applicant]
“Add Captions to your videos automatically, in just minutes!”, Maestra, downloaded Sep. 9, 2022, https://maestrahelp.zendesk.com/hc/en-us/articles/360053575651-Add-Captions-to-your-videos-automatically-in-just-minutes-. [cited by applicant]