IP Library Granted Patent US 11,159,597
Granted Patent B2
US 11,159,597 · App. 16/777,097 · Granted Oct 26, 2021

Systems and methods for artificial dubbing

Inventors: Ben Avi Ingel (Binyamina, IL); Ron Zass (Kiryat Tivon, IL)
Assignee: VIDUBLY LTD
H04L65/605G10L13/00G10L17/00
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,159,597
App. No.
16/777,097
Granted
Oct 26, 2021
Kind
B2
Abstract

Methods, systems, and computer-readable media for artificially generating a revoiced media stream are provided. In one implementation, a system may receive a media stream including an individual with particular voice speaking in an origin language. The system may obtain a transcript of the media stream including utterances spoken in the origin language and translate the transcript to a target language. The translated transcript may include a set of words in the target language for each of at least some of the utterances spoken in the origin language. The system may analyze the media stream to determine a voice profile for the individual. Thereafter, the system may determine a synthesized voice for a virtual entity intended to dub the individual that is similar to the particular voice. Then, the system may generate a revoiced media stream in which the translated transcript in the target language is spoken by the virtual entity.

Claims (73)

1. A computer program product for artificially generating revoiced media streams, the computer program product embodied in a non-transitory computer-readable medium and including instructions for causing at least one processor to execute a method comprising:

receiving a media stream including an individual speaking in an origin language and sounds from a sound-emanating object, wherein the individual is associated with particular voice;

obtaining a transcript of the media stream including utterances spoken in the origin language;

translating the transcript of the media stream to a target language, wherein the translated transcript includes a set of words in the target language for each of at least some of the utterances spoken in the origin language;

analyzing the media stream to determine a voice profile for the individual, wherein the voice profile includes characteristics of the particular voice;

determining a synthesized voice for a virtual entity intended to dub the individual, wherein the synthesized voice has characteristics identical to the characteristics of the particular voice;

determining auditory relationship between the individual and the sound-emanating object, wherein the auditory relationship is indicative of a ratio of volume levels between the utterances spoken in the origin language by the individual and the sounds from the sound-emanating object; and

generating a revoiced media stream in which the translated transcript in the target language is spoken by the virtual entity, and wherein in the revoiced media stream a ratio of volume levels between utterances spoken by the virtual entity in the target language and the sounds from the sound-emanating object is substantially identical to the ratio of volume levels between utterances spoken by the individual in the original language and the sounds from the sound-emanating object.

2. The computer program product of claim 1 , wherein the media stream including a plurality of first individuals speaking in a primary language and at least one second individual speaking in a secondary language, and the method further includes:

using determined voice profiles for the plurality of first individuals to artificially generate a revoiced media stream in which a plurality of virtual entities associated with the plurality of first individuals speak the target language and at least one virtual entity associated the at least one second individual speaks the secondary language.

3. The computer program product of claim 1 , wherein the media stream including a first individual speaking in a first origin language and a second individual speaking in a second origin language, and the method further includes:

using determined voice profiles for the first and second individual to artificially generate a revoiced media stream in which virtual entities associated with both the first individual and the second individuals speak the target language.

4. The computer program product of claim 1 , wherein the media stream including at least one individual speaking in a first origin language with an accent in a second language, and the method further includes:

determining a desired level of accent in the second language to introduce in a dubbed version of the received media stream; and

using determined at least one voice profile for the at least one individual to artificially generate a revoiced media stream in which at least one virtual entity associated with the at least one individual speaks the target language with an accent in the second language at the desired level.

5. The computer program product of claim 1 , wherein the media stream including at a first individual and a second individual speaking the origin language, and the method further includes:

based on at least one rule for revising transcripts of media streams, automatically revising a first part of the transcript associated with the first individual and avoid from revising a second part of the transcript associated with the second individual; and

using determined voice profiles for the first and second individuals to artificially generate a revoiced media stream in which a first virtual entity associated with the first individual speaks the revised first part of the transcript and a second virtual entity associated with the second individual speaks the second unrevised part of the transcript.

6. The computer program product of claim 1 , wherein the media stream is destined to a particular user, and the method further includes:

based on a determined user category indicative of a desired vocabulary for the particular user, revising the transcript of the media stream; and

using determined voice profile for the individual to artificially generate a revoiced media stream in which the virtual entity associated with the individual speaks the revised transcript in the target language.

7. The computer program product of claim 1 , wherein the media stream is destined to a particular user, and the method further includes:

translating the transcript of the media stream to the target language based on received preferred language characteristics; and

using determined voice profile for the individual to artificially generate a revoiced media stream in which the virtual entity associated with the individual speaks in the target language according to the preferred language characteristics of the particular user.

8. The computer program product of claim 1 , wherein the media stream is destined to a particular user, and the method further includes:

determining a preferred target language for the particular user; and

using the determined voice profile for the individual to artificially generate a revoiced media stream in which the translated transcript is spoken by the virtual entity associated with the individual in the preferred target language.

9. The computer program product of claim 1 , wherein the method further includes:

analyzing the transcript to determine a set of language characteristics for the individual; and

using the determined voice profile for the individual to artificially generate a revoiced media stream in which the translated transcript is spoken by the virtual entity associated with the individual, wherein the transcript is translated to the target language based on the determined set of language characteristics.

10. The computer program product of claim 1 , wherein the method further includes:

analyzing the transcript to determine that the individual discussed a subject likely to be unfamiliar with users associated with the target language; and

using the determined voice profile for the individual to artificially generate a revoiced media stream in which the translated transcript is spoken by the virtual entity associated with the individual, wherein the revoiced media stream provides explanation to the subject discussed by the individual in the origin language.

11. The computer program product of claim 1 , wherein the media stream is destined to a particular user, and the method further includes:

analyzing the transcript to determine that the individual in the received media stream discussed a subject likely to be unfamiliar with the particular user; and

using the determined voice profile for the individual to artificially generate a revoiced media stream in which the translated transcript is spoken by the virtual entity associated with the individual, wherein the revoiced media stream provides the determined explanation to the subject discussed by the at least one individual in the origin language.

12. The computer program product of claim 1 , wherein the method further includes:

analyzing the transcript to determine that an original name of a character in the received media stream is likely to cause antagonism with users that speak the target language; and

using the determined voice profile for the individual to artificially generate a revoiced media stream in which the translated transcript is spoken by the virtual entity associated with the individual and the character has a substitute name.

13. The computer program product of claim 1 , wherein the method further includes:

determining that the transcript includes a first utterance that rhymes with a second utterance; and

using the determined voice profile for the individual to artificially generate a revoiced media stream in which the translated transcript is spoken by the virtual entity associated with the individual, wherein the transcript is translated in a manner that at least partially preserves the rhymes of the transcript in the origin language.

14. The computer program product of claim 1 , wherein the voice profile is indicative of a ratio of volume levels between different utterances spoken by the individual in the origin language, and the method further includes:

determining metadata information for the translated transcript, wherein the metadata information includes desired volume levels for different words; and

using the determined voice profile for the individual to artificially generate a revoiced media stream in which the translated transcript is spoken by the virtual entity associated with the individual, wherein a ratio of the volume levels between utterances spoken by the virtual entity in the target language are substantially identical to the ratio of volume levels between different utterances spoken by the individual in the origin language.

15. The computer program product of claim 1 , wherein the media stream including at a first individual and a second individual speaking the origin language, and the method further includes:

analyzing the media stream to determine voice profiles for the first individual and the second individual, wherein the voice profiles are indicative of a ratio of volume levels between utterances spoken by each individual as they were recorded in the media stream; and

using the determined voice profiles for the first individual and the second individual to artificially generate a revoiced media stream in which the translated transcript is spoken by a first virtual entity associated with the first individual and a second virtual entity associated with the second individual, wherein a ratio of the volume levels between utterances spoken by the first virtual entity and the second virtual entity in the target language are substantially identical to the ratio of volume levels between utterances spoken by the first individual and the second individual in the origin language.

16. The computer program product of claim 1 , wherein the method further includes:

determining timing differences between the original language and the target language, wherein the timing differences represent time discrepancy between saying the utterances in a target language and saying the utterances in the original language; and

using determined voice profile for the individual to artificially generate a revoiced media stream in which the translated transcript is spoken by the virtual entity associated with the individual in a manner than accounts for the determined timing differences between the original language and the target language.

17. The computer program product of claim 1 , wherein the method further includes:

analyzing the media stream to determine a set of voice parameters of the individual and visual data; and

using a voice profile for the individual determined based on the set of voice parameters and visual data to artificially generate a revoiced media stream in which the translated transcript is spoken by the virtual entity associated with the individual.

18. The computer program product of claim 1 , wherein the method further includes:

analyzing the media stream to determine visual data; and

using the voice profile for the individual to artificially generate a revoiced media stream in which the translated transcript is spoken by the virtual entity associated with the individual, wherein the translation of the transcript to the target language is based on the visual data.

19. The computer program product of claim 1 , wherein the method further includes:

analyzing the media stream to determine visual data that includes text written in the origin language; and

using the voice profile for the individual to artificially generate a revoiced media stream in which the translated transcript is spoken by the virtual entity associated with the individual, wherein the revoiced media stream provides a translation to the text written in the origin language.

20. A system for artificially generating revoiced media streams, the system comprising:

at least one processing device configured to:

receive a media stream including an individual speaking in an origin language, wherein the individual is associated with particular voice;

obtain a transcript of the media stream including utterances spoken in the origin language;

translate the transcript of the media stream to a target language, wherein the translated transcript includes a set of words in the target language for each of at least some of the utterances spoken in the origin language;

analyze the media stream to determine a voice profile for the individual, wherein the voice profile includes characteristics of the particular voice;

determine a synthesized voice for a virtual entity intended to dub the individual, wherein the synthesized voice has characteristics identical to the characteristics of the particular voice;

determine auditory relationship between the individual and the sound-emanating object, wherein the auditory relationship is indicative of a ratio of volume levels between the utterances spoken in the origin language by the individual and the sounds from the sound-emanating object; and

generate a revoiced media stream in which the translated transcript in the target language is spoken by the virtual entity, and wherein in the revoiced media stream a ratio of volume levels between utterances spoken by the virtual entity in the target language and the sounds from the sound-emanating object is substantially identical to the ratio of volume levels between utterances spoken by the individual in the original language and the sounds from the sound-emanating object.

21. The system of claim 20 , wherein the media stream including a plurality of first individuals speaking in a primary language and at least one second individual speaking in a secondary language, and the at least one processing device is further configured to:

use determined voice profiles for the plurality of first individuals to artificially generate a revoiced media stream in which a plurality of virtual entities associated with the plurality of first individuals speak the target language and at least one virtual entity associated the at least one second individual speaks the secondary language.

22. The system of claim 20 , wherein the media stream including a first individual speaking in a first origin language and a second individual speaking in a second origin language, and the method further includes:

use determined voice profiles for the first and second individual to artificially generate a revoiced media stream in which virtual entities associated with both the first individual and the second individuals speak the target language.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Aug 22, 2021
From: INGEL, BEN AVI; ZASS, RON
To: VIDUBLY LTD
Reel/Frame 057247/0908 →
Continuity (4)
Provisional Application 62822856 · Mar 23, 2019
Provisional Application 62816137 · Mar 10, 2019
Provisional Application 62799970 · Feb 1, 2019
Related Publication 20200169591A1 · May 28, 2020
Cited By (16)
US 12,242,826 US 12,279,022 US 12,279,023 US 12,282,755 US 12,380,736 US 12,412,559 US 12,520,014 US 12,591,419 US 12,596,535 US 12,596,883 US 12,597,291 US 12,632,231 US 12,632,667 US 12,639,052 US 12,682,522 US 12,688,352