IP Library Granted Patent US 12,010,399
Granted Patent B2
US 12,010,399 · App. 18/097,900 · Granted Jun 11, 2024

Generating revoiced media streams in a virtual reality

Inventors: Ben Avi Ingel (Binyamina, IL); Ron Zass (Kiryat Tivon, IL)
H04N21/8126G06F40/58G10L13/00G10L13/033G10L13/086G10L13/10H04N21/2668H04N21/458H04N21/4755
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,010,399
App. No.
18/097,900
Granted
Jun 11, 2024
Kind
B2
Abstract

Methods, systems, and computer-readable media for generating videos with characters indicating regions of images are provided. For example, an image containing a first region may be received. At least one characteristic of a character may be obtained. A script containing a first segment of the script may be received. The first segment of the script may be related to the first region of the image. The at least one characteristic of a character and the script may be used to generate a video of the character presenting the script and at least part of the image, where the character visually indicates the first region of the image while presenting the first segment of the script.

Claims (41)

1. A computer program product for generating a revoiced media stream in a virtual reality system, the computer program product embodied in a non-transitory computer-readable medium and including instructions for causing at least one processor to execute a method comprising:

receiving a media stream from an individual speaking in an origin language, wherein the individual is associated with a particular voice;

obtaining a transcript of the media stream in the origin language;

translating the transcript of the media stream to a target language, wherein the translated transcript includes at least one word in the target language for each word spoken in the origin language;

analyzing the media stream to determine a voice profile for the individual, wherein the voice profile corresponds with the particular voice of the individual;

determining at least one characteristic of a personalized avatar that represents the individual;

determining a synthesized voice for the personalized avatar based on the voice profile, wherein the synthesized voice sounds substantially identical to the particular voice; and

enabling the virtual reality system to generate a revoiced media stream that includes a visualization of the personalized avatar speaking the translated transcript in the target language using the synthesized voice.

2. The computer program product of claim 1 , wherein the method further includes: determining a desired level of origin language accent to introduce in the synthesized voice of the personalized avatar; and enabling the virtual reality system to generate a revoiced media stream that includes a visualization of the personalized avatar that speaks in the target language with the desired level of origin language accent.

3. The computer program product of claim 1 , wherein the method further includes: based on at least one rule for revising transcripts of media streams, automatically revising a first part of the transcript and avoid from revising a second part of the transcript; and enabling the virtual reality system to generate a revoiced media stream that includes a visualization of the personalized avatar that speaks the first revised part the translated transcript in the target language and the second unrevised part the translated transcript in the target language.

4. The computer program product of claim 1 , wherein the method further includes: based on a determined user category indicative of a vocabulary level of associated with the individual, revising the transcript of the media stream; and enabling the virtual reality system to generate the revoiced media stream that includes the visualization of the personalized avatar that speaks the revised transcript in the target language.

5. The computer program product of claim 1 , wherein the method further includes: translating the transcript of the media stream to the target language based on preferred language characteristics of the individual; and enabling the virtual reality system to generate the revoiced media stream that includes the visualization of the personalized avatar that speaks the translated the transcript in the target language.

6. The computer program product of claim 1 , wherein the method further includes: analyzing the transcript to determine a set of language characteristics associated with the individual; translating the transcript to the target language based on the determined set of language characteristics; and enabling the virtual reality system to generate the revoiced media stream that includes the visualization of the personalized avatar that speaks the translated the transcript in the target language.

7. The computer program product of claim 1 , wherein the method further includes: based on at least one rule for translating transcripts of media streams, automatically translating a first part of the transcript to the target language and avoid from translating a second part of the transcript to the target language; and enabling the virtual reality system to generate the revoiced media stream that includes the visualization of the personalized avatar that speaks the first part of the transcript in the target language and the second part of the transcript in the origin language.

8. The computer program product of claim 1 , wherein the method further includes: analyzing the transcript to determine that the individual discusses a subject likely to be unfamiliar with at least one individual that would listen to the personalized avatar; and enabling the virtual reality system to provide an explanation in the target language to the subject discussed by the individual in the origin language.

9. The computer program product of claim 1 , wherein the method further includes: determining metadata information for the translated transcript, wherein the metadata information includes desired volume levels for different words; and enabling the virtual reality system to generate the revoiced media stream in which a ratio of the volume levels between words spoken by the personalized avatar in the target language is substantially identical to the ratio of volume levels between different words spoken by the individual in the origin language.

10. The computer program product of claim 1 , wherein the received media stream is associated with a real-time conversation between the individual and at least one other individual, and the method further includes processing visual data to identify text written in the origin language; determining relevancy of the identified text to the particular user; and providing a translation in the target language for the identified text, when the content of the identified text is determined to be relevant.

11. The computer program product of claim 1 , wherein the personalized avatar that speaks in the target language is a realistic avatar or a semi-realistic avatar associated with a depiction of the individual that speaks in the origin language.

12. The computer program product of claim 1 , wherein the method further includes causing the personalized avatar to visually point to an object while presenting in the target language a segment of the translated transcript that relates to the object.

13. The computer program product of claim 12 , wherein the object is a graphic presentation of a weather forecast, a graphic presentation of a calendar event, or a graphic presentation of a past event associated with the individual.

14. The computer program product of claim 1 , wherein the method further includes: receiving user selection from the individual to determine the target language for the personalized avatar; and enabling the virtual reality system to generate a revoiced media stream that includes a visualization of the personalized avatar that speaks the selected target language.

15. The computer program product of claim 1 , wherein the received media stream is associated with a real-time conversation between the individual and at least one other individual, and the method further includes determining a preferred target language for the personalized avatar based on an identity at least one other individual; and enabling the virtual reality system to generate a revoiced media stream that includes a visualization of the personalized avatar that speaks in the preferred target language.

16. The computer program product of claim 1 , wherein the method further includes selecting the at least one characteristic of the personalized avatar based on a profile of the individual associated, at least in part, on a geographical location associated with the individual.

17. The computer program product of claim 1 , wherein the method further includes receiving user selection from the individual to enable selective manipulation of the at least one visual characteristic of the personalized avatar.

18. The computer program product of claim 1 , wherein the method further includes receiving user selection from the individual to enable selective manipulation of the at least one voice characteristic of the personalized avatar.

19. A method for artificially generating a revoiced media stream in a virtual reality system, the method comprising:

receiving a media stream from an individual speaking in an origin language, wherein the individual is associated with a particular voice;

obtaining a transcript of the media stream in the origin language;

translating the transcript of the media stream to a target language, wherein the translated transcript includes at least one word in the target language for each word spoken in the origin language;

analyzing the media stream to determine a voice profile for the individual, wherein the voice profile corresponds with the particular voice of the individual;

determining at least one characteristic of a personalized avatar that represents the individual;

determining a synthesized voice for the personalized avatar based on the voice profile, wherein the synthesized voice sounds substantially identical to the particular voice; and

enabling the virtual reality system to generate a revoiced media stream that includes a visualization of the personalized avatar speaking the translated transcript in the target language using the synthesized voice.

20. A virtual reality system for artificially generating a revoiced media stream, the system comprising at least one processing device configured to:

receive a media stream from an individual speaking in an origin language, wherein the individual is associated with a particular voice;

obtain a transcript of the media stream in the origin language;

translate the transcript of the media stream to a target language, wherein the translated transcript includes at least one word in the target language for each word spoken in the origin language;

analyze the media stream to determine a voice profile for the individual, wherein the voice profile corresponds with the particular voice of the individual;

determine at least one characteristic of a personalized avatar that represents the individual;

determine a synthesized voice for the personalized avatar based on the voice profile, wherein the synthesized voice sounds substantially identical to the particular voice; and

generate a revoiced media stream that includes a visualization of the personalized avatar speaking the translated transcript in the target language using the synthesized voice.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Apr 25, 2026
From: INGEL, BEN AVI; ZASS, RON
To: VIDUBLY LTD
Reel/Frame 074473/0916 →
Continuity (5)
Continuation 17460644 · Aug 30, 2021
Continuation 16813984 · Mar 10, 2020
Provisional Application 62822856 · Mar 23, 2019
Provisional Application 62816137 · Mar 10, 2019
Related Publication 20230156294A1 · May 18, 2023
Cited By (6)
US 12,279,023 US 12,380,736 US 12,520,014 US 12,597,291 US 12,738,099 US 12,739,479