IP Library Granted Patent US 12682885
Granted Patent B2
US 12682885 · App. 18/581,098 · Granted Jul 14, 2026

Audio and video translator

Inventors: Rijul Gupta (Oakland, CA); Emma Brown (Portland, OR)
Assignee: Deep Media Inc.
G10L13/08G06F40/103G06F40/42G06T7/20G10L15/04G10L17/00G06T2207/30201
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12682885
App. No.
18/581,098
Granted
Jul 14, 2026
Kind
B2
Abstract

A system and method for translating audio, and video when desired. The translations include synthetic media and data generated using AI systems. Through unique processors and generators executing a unique sequence of steps, the system and method produces more accurate translations that can account for various speech characteristics (e.g., emotion, pacing, idioms, sarcasm, jokes, tone, phonemes, etc.). These speech characteristics are identified in the input media and synthetically incorporated into the translated outputs to mirror the characteristics in the input media. Some embodiments further include systems and methods that manipulate the input video such that the speakers' faces and/or lips appear as if they are natively speaking the generated audio.

Claims (43)

1 . A system configured to generate a synthetic media with translated speech corresponding to an input media file, comprising:

at least one processor; and

memory including instructions that, when executed by the at least one processor, cause the system to:

digitally acquire the input media file, wherein the input media file includes input audio in a first input language;

acquire an input transcription of the input audio, wherein the input transcription includes text corresponding to the words spoken in the input audio;

acquire input meta information of the input audio, wherein acquiring the meta information includes identifying emotion data for one or more vocal segments of the input media file and emotion data includes labels from a list of predetermined emotions;

input the input transcription and the input meta information with detected emotions for the one or more vocal segments as input data into an artificial intelligence (AI) generator;

translate, by the AI generator, the first input language into a first output language, such that a translated transcription and meta information include similar emotion in comparison to the input transcription and input meta information;

provide the translated transcription and meta information to an audio translation generator which is configured to generate translated audio; and

generate the synthetic digital media having translated audio.

2 . The system of claim 1 , wherein the input media file is in a computer-readable format.

3 . The system of claim 1 , wherein the instructions included in the memory further cause the system to partition one vocal stream from another, reduce background noise, or enhance a quality of the vocal streams.

4 . The system of claim 1 , wherein the instructions included in the memory further cause the system to capture lip movement tracking data from the input media file.

5 . The system of claim 1 , wherein the instructions included in the memory further cause the system to identify pacing information for each word or phoneme in each vocal segment.

6 . The system of claim 5 , wherein segmenting the input audio and identifying pacing information is performed by a speaker diarization processor configured to receive the input media file.

7 . The system of claim 1 , wherein the text of the input transcription is formatted according to the international phonetics alphabet.

8 . The system of claim 1 , wherein the input transcription further includes sentiment analysis and tracking data corresponding to anatomical landmarks for the speaker for each vocal segment.

9 . The system of claim 1 , wherein acquiring the input transcription of the input audio includes providing the input audio to an AI generator configured to convert the input audio into text.

10 . The system of claim 1 , wherein acquiring input meta information includes providing the input audio and the input transcription to an AI meta information processor configured to identify meta information.

11 . The system of claim 1 , wherein the instructions included in the memory further cause the system to stitch the translated audio for each vocal segment back into a single audio file.

12 . The system of claim 1 , wherein the input media file includes input video and the instructions included in the memory further cause the system to provide translated audio and the input video to a video sync generator and generating, by the video sync generator, a synced video in which the translated audio syncs with the input video.

13 . A method for generating a synthetic media with translated speech corresponding to an input media file, comprising:

acquiring the input media file, wherein the input media file includes input video and input audio in a first input language;

acquiring an input transcription, wherein the input transcription includes text corresponding to the words spoken in the input audio;

acquiring input meta information derived from the input media file for one or more vocal segments of the input audio, wherein acquiring the meta information includes identifying tone data and identifying one or more detectable emotions temporally aligned with the one or more vocal segments, selected from a list of predetermined emotions;

inputting the input meta information including the temporally aligned detected emotions for the one or more vocal segments as input data into a meta translation generator;

translating, by the meta translation generator, the input transcription and input meta information into a first output language, such that a translated transcription includes similar emotion in comparison to the input transcription and input meta information; and

generating the synthetic digital media having translated audio by providing the translated input transcription and meta information to an audio translation generator.

14 . The method of claim 13 , further including capturing lip movement tracking data from the input video.

15 . The method of claim 13 , further including formatting the input transcription according to the international phonetics alphabet.

16 . The method of claim 13 , further including identifying pacing information for each word or phoneme in each vocal segment.

17 . The method of claim 13 , further including stitching the translated audio for each vocal segment back into a single audio file.

18 . The method of claim 13 , further including:

providing the translated audio and the input video to a video sync generator; and

generating, by the video sync generator, a synced video in which the translated audio syncs with the input video to create the synthetic media.

19 . A non-transitory computer readable medium for generating a synthetic media with translated speech corresponding to an input media file, comprising instructions stored thereon, that when executed on at least one processor, cause the at least one processor to:

digitally acquire the input media file, wherein the input media file includes input audio in a first input language;

acquire an input transcription, wherein the input transcription includes text corresponding to the words spoken in the input audio;

acquire input meta information corresponding to one or more vocal segments in the input audio, the acquired meta information including emotion data temporally aligned with the one or more vocal segments and selected from a list of predetermined emotions;

input the input meta information including temporally aligned and detected emotions for the one or more vocal segments and input transcription as input data into a transcription and meta translation generator;

translate, by the transcription and meta translation generator, the input transcription and input meta information into a first output language, such that a translated transcription includes similar emotion in comparison to the input transcription and input meta information;

provide the translated transcription and meta information to an audio translation generator which is configured to generate translated audio; and

generate the synthetic digital media having translated audio.