IP Library Granted Patent US 11,551,664
Granted Patent B2
US 11,551,664 · App. 17/737,546 · Granted Jan 10, 2023

Audio and video translator

Inventors: Rijul Gupta (Oakland, CA); Emma Brown (Portland, OR)
Assignee: Deep Media Inc.
G10L13/08G06F40/103G06F40/42G06T7/20G10L15/04G10L17/00G06T2207/30201
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,551,664
App. No.
17/737,546
Granted
Jan 10, 2023
Kind
B2
Abstract

A system and method for translating audio, and video when desired. The translations include synthetic media and data generated using AI systems. Through unique processors and generators executing a unique sequence of steps, the system and method produces more accurate translations that can account for various speech characteristics (e.g., emotion, pacing, idioms, sarcasm, jokes, tone, phonemes, etc.). These speech characteristics are identified in the input media and synthetically incorporated into the translated outputs to mirror the characteristics in the input media. Some embodiments further include systems and methods that manipulate the input video such that the speakers' faces and/or lips appear as if they are natively speaking the generated audio.

Claims (40)

1. A method for generating a synthetic media with translated speech corresponding to an input media file, comprising:

digitally acquiring the input media file, wherein the input media file includes input audio in a first input language;

acquiring a first output language, wherein the first output language is different from the first input language;

segmenting the input audio into a plurality of vocal segments, wherein each vocal segment in the plurality of vocal segments includes a speaker identification to identify the speaker of each vocal segment;

for each vocal segment in the plurality of vocal segments:

identifying pacing information for each word or phoneme in each vocal segment;

acquiring an input transcription, wherein the input transcription includes text corresponding to the words spoken in each vocal segment;

acquiring input meta information, the meta information including emotion data and tone data, wherein emotion data corresponds to one or more detectable emotions from a list of predetermined emotions;

inputting the input meta information and input transcription into a transcription and meta translation generator, wherein the transcription and meta translation generator is a generative adversarial network generator;

translating the input transcription and input meta information into the first output language based at least on the timing information and the emotion data via the transcription and meta translation generator, such that the translated transcription and meta information include similar emotion and pacing in comparison to the input transcription and input meta information; and

generating the synthetic digital media file having translated audio.

2. The method of claim 1 , wherein the input media file is in a computer-readable format.

3. The method of claim 1 , further including preprocessing input audio to partition one vocal stream from another, reduce background noise, or enhance a quality of the vocal streams.

4. The method of claim 1 , further including preprocessing the input video to capture lip movement tracking data.

5. The method of claim 1 , wherein segmenting the input audio into the plurality of vocal segments and identifying pacing information is performed by a speaker diarization processor configured to receive the input media file as an input.

6. The method of claim 1 , wherein the text transcription is formatted according to the international phonetics alphabet.

7. The method of claim 1 , wherein the input transcription further includes sentiment analysis and tracking data corresponding to anatomical landmarks for the speaker for each vocal segment.

8. The method of claim 1 , wherein acquiring the input transcription of the input audio includes providing the input audio to an artificial intelligence (AI) generator configured to convert the input audio into text.

9. The method of claim 1 , wherein acquiring input meta information includes providing the input audio and the input transcription to an AI meta information processor configured to identify meta information.

10. The method of claim 1 , wherein similar pacing includes less than or equal to a 20% difference.

11. The method of claim 1 , wherein generating translated audio includes providing the translated transcription and meta information to the audio translation generator configured to generate the translated audio.

12. The method of claim 1 , further including stitching the translated audio for each vocal segment back into a single audio file.

13. The method of claim 1 , wherein the input media file includes input video.

14. The method of claim 13 , further including providing the translated audio and the input video to a video sync generator and generating, by the video sync generator, a synced video in which the translated audio syncs with the input video.

15. A method for generating a synthetic media with translated speech corresponding to an input media file, comprising:

acquiring the input media file, wherein the input media file includes input video and input audio in a first input language;

acquiring a first output language, wherein the first output language is different from the first input language;

segmenting the input audio into a plurality of vocal segments, wherein each vocal segment in the plurality of vocal segments includes a speaker identification to identify the speaker of each vocal segment;

for each vocal segment in the plurality of vocal segments:

identifying pacing information for each word or phoneme in each vocal segment;

acquiring an input transcription, wherein the input transcription includes text corresponding to the words spoken in each vocal segment;

acquiring input meta information, the meta information including emotion data and tone data, wherein emotion data corresponds to one or more detectable emotions from a list of predetermined emotions;

translating the input transcription and input meta information into the first output language based at least on the timing information and the emotion data, such that the translated transcription and meta information include similar emotion and pacing in comparison to the input transcription and input meta information;

generating translated audio using translated input transcription and meta information as inputs to an audio translation generator, wherein the audio translation generator is a generative adversarial network generator;

stitching the translated audio for each vocal segment back into a single audio file; and

providing the translated audio and the input video to a video sync generator and generating, by the video sync generator, a synced video in which the translated audio syncs with the input video to create the synthetic media.

16. The method of claim 15 , wherein the input media file is in a computer-readable format.

17. The method of claim 15 , further including preprocessing the input video to capture lip movement tracking data.

18. The method of claim 15 , wherein the text transcription is formatted according to the international phonetics alphabet.

19. The method of claim 15 , wherein generating translated audio includes providing the translated transcription and meta information to an audio translation generator configured to generate the translated audio, wherein the audio translation generator is a generative adversarial network generator.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded May 18, 2022
From: GUPTA, RIJUL; BROWN, EMMA
To: DEEP MEDIA INC.
Reel/Frame 059948/0048 →
Continuity (2)
Provisional Application 63184746 · May 5, 2021
Related Publication 20220358905A1 · Nov 10, 2022