IP Library Granted Patent US 12,452,477
Granted Patent B2
US 12,452,477 · App. 18/674,558 · Granted Oct 21, 2025

Video and audio synchronization with dynamic frame and sample rates

Inventors: Clara Fernandez Labrador (Zurich, CH); Cafer Mertcan Akcay (Zurich, CH); Christopher Richard Schroers (Uster, CH); Joan Massich Vall (Zurich, CH); Scott Labrozzi (Cary, NC); Mitchel Jacobs (Malibu, CA); Katherine Hinsen (Los Angeles, CA); Eitan Abecassis (Raleigh, NC)
Assignee: Disney Enterprises, Inc.
H04N21/4307H04N19/60
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,452,477
App. No.
18/674,558
Granted
Oct 21, 2025
Kind
B2
Abstract

A system includes a hardware processor and a memory storing a video/audio (V/A) synchronizer including video and audio encoders. The hardware processor executes the V/A synchronizer to receive raw video and audio extracted from media content, partition the raw video into video frame patches, partition the raw audio into audio samples, pre-process the video frame patches and the audio samples for encoding. The hardware processor further executes the V/A synchronizer to encode, using the video encoder, the pre-processed video frame patches to provide pre-processed and encoded video frame patches used to provide a latent representation of the raw video, encode, using the audio encoder, the pre-processed audio samples to provide pre-processed and encoded audio samples used to provide a latent representation of the raw audio, and synchronize, using the latent representations of the raw video and the raw audio, the raw audio with the raw video.

Claims (76)

1. A system comprising:

a hardware processor; and

a memory storing a video/audio (V/A) synchronizer including a video encoder and an audio encoder;

the hardware processor configured to execute the V/A synchronizer to:

receive raw video and raw audio extracted from media content;

partition the raw video into a plurality of video frame patches;

partition the raw audio into a plurality of audio samples;

pre-process the plurality of video frame patches for encoding to provide a plurality of pre-processed video frame patches;

pre-process the plurality of audio samples for encoding to provide a plurality of pre-processed audio samples;

encode, using the video encoder, the plurality of pre-processed video frame patches to provide a plurality of pre-processed and encoded video frame patches;

encode, using the audio encoder, the plurality of pre-processed audio samples to provide a plurality of pre-processed and encoded audio samples;

provide, using one or more of the plurality of pre-processed and encoded video frame patches, a latent representation of the raw video;

provide, using the plurality of pre-processed and encoded audio samples, a latent representation of the raw audio; and

synchronize, using the latent representation of the raw video and the latent representation of the raw audio, the raw audio with the raw video.

2. The system of claim 1 , wherein all of the plurality of pre-processed and encoded video frame patches are used to provide the latent representation of the raw video.

3. The system of claim 1 , wherein at least one of the plurality of pre-processed and encoded video frame patches is not used to provide the latent representation of the raw video, and wherein the at least one of the plurality of pre-processed and encoded video frame patches is omitted from use randomly or based on attention.

4. The system of claim 1 , wherein the raw video and the raw audio are not transformed from original media specifications of the media content.

5. The system of claim 1 , wherein:

to pre-process the plurality of video frame patches, the hardware processor is further configured to execute the V/A synchronizer to project each of the plurality of video frame patches onto a respective video token to provide a plurality of tokenized video frame patches;

to pre-process the plurality of audio samples, the hardware processor is further configured to execute the V/A synchronizer to project each of the plurality of audio samples onto a respective audio token to provide a plurality of tokenized audio samples; or

a combination thereof.

6. The system of claim 5 , wherein:

to pre-process the plurality of video frame patches, the hardware processor is further configured to execute the V/A synchronizer to concatenate the plurality of tokenized video frame patches with a learnable video modality token;

to pre-process the plurality of audio samples, the hardware processor is further configured to execute the V/A synchronizer to concatenate the plurality of tokenized audio samples with a learnable audio modality token; or

a combination thereof.

7. The system of claim 5 , wherein:

to pre-process the plurality of video frame patches, the hardware processor is further configured to execute the V/A synchronizer to apply time-aware positional encoding to the plurality of tokenized video frame patches;

to pre-process the plurality of audio samples, the hardware processor is further configured to execute the V/A synchronizer to apply time-aware positional encoding to the plurality of tokenized audio samples; or

a combination thereof.

8. The system of claim 1 , wherein:

the video encoder comprises a first Transformer trained to encode video;

the audio encoder comprises a second Transformer trained to encode audio; or

a combination thereof.

9. The system of claim 1 , wherein a number of video frames included in the raw video varies based on an original frame rate of the media content, and wherein a number of audio samples included in the raw audio varies based on an original sample rate of the media content.

10. The system of claim 1 , wherein:

to synchronize the raw audio with the raw video, the hardware processor is further configured to execute the V/A synchronizer to compare the latent representation of the raw video with the latent representation of the raw audio through a contrastive loss.

11. The system of claim 1 , wherein the V/A synchronizer does not include a convolutional neural network.

12. The system of claim 1 , wherein synchronizing the raw audio with the raw video provides a first V/A synchronized media segment, and wherein the hardware processor is further configured to execute the V/A synchronizer to:

synchronize at least a second raw audio segment of the media content with at least a second raw video segment of the media content to provide at least a second V/A synchronized media segment; and

assess, based on the first V/A synchronized media segment and the at least the second V/A synchronized media segment, a V/A synchronization status of the media content as a whole.

13. A method for use by a system including a hardware processor and a memory storing a video/audio (V/A) synchronizer including a video encoder and an audio encoder, the method comprising:

receiving, by the V/A synchronizer executed by the hardware processor, raw video and raw audio extracted from media content;

partitioning, by the V/A synchronizer executed by the hardware processor, the raw video into a plurality of video frame patches;

partitioning, by the V/A synchronizer executed by the hardware processor, the raw audio into a plurality of audio samples;

pre-processing, by the V/A synchronizer executed by the hardware processor, the plurality of video frame patches for encoding to provide a plurality of pre-processed video frame patches;

pre-processing, by the V/A synchronizer executed by the hardware processor, the plurality of audio samples for encoding to provide a plurality of pre-processed audio samples;

encoding, by the V/A synchronizer executed by the hardware processor and using the video encoder, the plurality of pre-processed video frame patches to provide a plurality of pre-processed and encoded video frames;

encoding, by the V/A synchronizer executed by the hardware processor and using the audio encoder, the plurality of pre-processed audio samples to provide a plurality of pre-processed and encoded audio samples;

providing, by the V/A synchronizer executed by the hardware processor and using one or more of the plurality of pre-processed and encoded video frame patches, a latent representation of the raw video;

providing, by the V/A synchronizer executed by the hardware processor and using the plurality of pre-processed and encoded audio samples, a latent representation of the raw audio; and

synchronizing, by the V/A synchronizer executed by the hardware processor and using the latent representation of the raw video and the latent representation of the raw audio, the raw audio with the raw video.

14. The method of claim 13 , wherein all of the plurality of pre-processed and encoded video frame patches are used to provide the latent representation of the raw video.

15. The method of claim 13 , wherein at least one of the plurality of pre-processed and encoded video frame patches is not used to provide the latent representation of the raw video, and wherein the at least one of the plurality of video frame patches is omitted from use randomly or based on attention.

16. The method of claim 13 , wherein the raw video and the raw audio are not transformed from original media specifications of the media content.

17. The method of claim 13 , further comprising:

pre-processing the plurality of video frame patches includes projecting, by the V/A synchronizer executed by the hardware processor, each of the plurality of video frame patches onto a respective video token to provide a plurality of tokenized video frame patches;

pre-processing the plurality of audio samples by projecting, by the V/A synchronizer executed by the hardware processor, each of the plurality of audio samples onto a respective audio token to provide a plurality of tokenized audio samples; or

a combination thereof.

18. The method of claim 17 , further comprising:

pre-processing the plurality of video frame patches by concatenating, by the V/A synchronizer executed by the hardware processor, the plurality of tokenized video frame patches with a learnable video modality token; and

pre-processing the plurality of audio samples by concatenating, by the V/A synchronizer executed by the hardware processor, the plurality of tokenized audio samples with a learnable audio modality token; or

a combination thereof.

19. The method of claim 17 , further comprising:

pre-processing the plurality of video frame patches by applying, by the V/A synchronizer executed by the hardware processor, time-aware positional encoding to the plurality of tokenized video frame patches; and

pre-processing the plurality of audio samples by applying, by the V/A synchronizer executed by the hardware processor, time-aware positional encoding to the plurality of tokenized audio samples; or

a combination thereof.

20. The method of claim 13 , wherein:

the video encoder comprises a first Transformer trained to encode video;

the audio encoder comprises a second Transformer trained to encode audio; or

a combination thereof.

21. The method of claim 13 , wherein how many video frames are included in the raw video varies based on an original frame rate of the media content, and wherein how many audio samples are included in the raw audio varies based on an original sample rate of the media content.

22. The method of claim 13 , wherein synchronizing the raw audio with the raw video comprises comparing the latent representation of the raw video with the latent representation of raw audio through a contrastive loss.

23. The method of claim 13 , wherein the V/A synchronizer does not include a convolutional neural network.

24. The method of claim 13 , wherein synchronizing the raw audio with the raw video provides a first V/A synchronized media segment, the method further comprising:

synchronizing, by the V/A synchronizer executed by the hardware processor, at least a second raw audio segment of the media content with a respective at least a second raw video segment of the media content to provide at least a second V/A synchronized media segment; and

assessing, by the V/A synchronizer executed by the hardware processor based on the first V/A synchronized media segment and the at least the second V/A synchronized media segment, a V/A synchronization status of the media content as a whole.

Assignments (3)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded May 28, 2024
From: LABROZZI, SCOTT; JACOBS, MITCHEL; HINSEN, KATHERINE; ABECASSIS, EITAN
To: DISNEY ENTERPRISES, INC.
Reel/Frame 067539/0396 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded May 28, 2024
From: FERNANDEZ LABRADOR, CLARA; MERTCAN AKCAY, CAFER; SCHROERS, CHRISTOPHER RICHARD; MASSICH VALL, JOAN
To: THE WALT DISNEY COMPANY (SWITZERLAND) GMBH
Reel/Frame 067539/0629 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded May 28, 2024
From: THE WALT DISNEY COMPANY (SWITZERLAND) GMBH
To: DISNEY ENTERPRISES, INC.
Reel/Frame 067539/0875 →
Continuity (2)
Provisional Application 63521604 · Jun 16, 2023
Related Publication 20240422380A1 · Dec 19, 2024
References Cited (26)
US 7657829B2 · Panabaker · 2010 [cited by examiner]
US 11601698B2 · Gramo · 2023 [cited by examiner]
US 20070153125A1 · Cooper · 2007 [cited by examiner]
US 20100079605A1 · Wang · 2010 [cited by examiner]
US 20130141643A1 · Carson · 2013 [cited by examiner]
US 20150062353A1 · Dalal · 2015 [cited by examiner]
US 20190037018A1 · Scurrell · 2019 [cited by examiner]
US 20210219012A1 · Maurice · 2021 [cited by examiner]
US 20240129580A1 · Collins · 2024 [cited by examiner]
US 20240177740A1 · Danielson · 2024 [cited by examiner]
US 20240292044A1 · Zhang · 2024 [cited by examiner]
US 20250054518A1 · Björkman · 2025 [cited by examiner]
Honglie Chen, Weidi Xie, Triantafyllos Afouras, Arsha Nagrani, Andrea Vedaldi, Andrew Zisserman “Audio-Visual Synchronisation in the Wild” ArXiv abs/2112.04432 2021 24 Pgs. [cited by applicant]
Venkatesh S. Kadandale, Juan F. Montesinos, Gloria Haro “VocaLiST: An Audio-Visual Synchronisation Model for Lips and Voices” Interspeech 2022 Sep. 18-22, 2022, Incheon, Korea 5 Pgs. [cited by applicant]
K R Prajwal, Vinay P. Namboodiri, Rudrabha Mukhopadhyay, C V Jawahar “A lip Sync Expert Is all You Need for Speech to Lip Generation in the Wild” Oral Session A2: Emerging Multimedia Applications MM'20, Oct. 12-16, 2020… [cited by applicant]
Yasheng Sun, hang Zhou, Kaisiyuan Wang, Qianyi Wu, Zhibin Hong, Jingtuo Liu, Errui Ding, Jingdong Wang, Ziwei Liu, hideki Koike “Masked Lip-Sync Prediction by Audio-Visual Contextual Wxploitation in transformers” 2022 A… [cited by applicant]
Kun Cheng, Xiaodong Cun, Yong Zhang, Menghan Xia, Fei Yin, Mingrui Zhu, Xuan Wang, Jue Wang, Nannan Wang “VideoReTalking: Audio-based Lip Synchronization for Talking Head Video Editing in the Wild” 2022 Association of C… [cited by applicant]
Goranka Zoric, Igor S. Pandzic “A Real-Time Lip Sync System using a Genetic Algorithm for Automatic Neural Network Configuration” Proceedings of the 2005 IEEE International Conference on Multimedia and Expo, ICME 2005, … [cited by applicant]
Barrett E. Koster, Robert D. Rodman and Donald bitzer “Automated Lip-Sync: Direct Translation of Speech-Sound to Mouth-Shape” Proceedings of 1994 Asilomar Conference on Signals, Systems and Computers, Oct. 31, 1994 4 Pg… [cited by applicant]
Joon Son Chung and Andrew Zisserman “Out of Time: automated lip sync in the wild” ACCV Workshops Nov. 20, 2016 13 Pgs. [cited by applicant]
Sucharu Aggarwal, Alka Jindal “Comprehensive Overview of various Lip Synchronization Techniques” 2008 International Symposium on Biometrics and Security Technologies, Islamabad, Pakistan, 2008 6 Pgs. [cited by applicant]
Namrata Dave “Feature Extraction Methods LPC, PLP, and MFCC in Speech Recognition” International Journal for Advance Research in Engineering and Technology vol. 1, Issue VI, Jul. 2013 5 Pgs. [cited by applicant]
Etienne Marcheret, Gerasimos Potamianos, Josef Vopicka, Vaibhava Goel “Detecting Audio-Visual Synchrony Using Deep Neural Networks” Interspeech 2015 5 Pgs. [cited by applicant]
Relja Arandjelovic and Andrew Zisserman “Objects that Sound” ECCV 2018 17 Pgs. [cited by applicant]
John L. Lewis “Automated Lip-sync: Background and Techniques” Tie Journal of Visualization and Computer Animation vol. 2: (1991) 5 Pgs. [cited by applicant]
Avijit Vajpayee Zhikang Zhang, Abhinav Jain, Vimal Bhat “A Simple and Efficient method for Dubbed Audio Sync Detection using Comprehensive Sensing” WACV 2023 8 Pgs. [cited by applicant]