IP Library › Granted Patent US 12,334,118
Granted Patent B2
US 12,334,118 · App. 16/687,209 · Granted Jun 17, 2025

Techniques for identifying synchronization errors in media titles

Inventors: Rohit Puri (Campbell, CA); Naji Khosravan (Orlando, FL); Shervin Ardeshir Behrostaghi (Campbell, CA)
Assignee: NETFLIX, INC.
G11B27/36G06F18/2433G06N3/045G06N3/08G06V10/764G06V10/82G06V20/41G06V20/46G10L25/30G10L25/57
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,334,118
App. No.
16/687,209
Granted
Jun 17, 2025
Kind
B2
Abstract

A neural network system that is trained to identify one or more portions of a media title where synchronization errors are likely to be present. The neural network system is trained based on a first set of media titles where synchronization errors are present and a second set of media titles where synchronization errors are absent. The second set of media titles can be generated by introducing synchronization errors into a set of media titles that otherwise lack synchronization errors. Via training, the neural network system learns to identify specific visual features included in one or more video frames and corresponding audio features that should be played back in synchrony with the associated visual features. Accordingly, when presented with a media title that includes synchronization errors, the neural network can indicate the specific frames where synchronization errors are likely to be present.

Claims (51)

1. A neural network system implemented by one or more computers,

wherein the neural network system identifies one or more blocks in a media clip that include audio data that is misaligned with corresponding video data,

wherein the neural network system comprises:

a convolutional subnetwork that generates a plurality of feature maps by generating, for each block included in a plurality of blocks in the media clip, a corresponding feature map derived from both one or more audio features and also one or more video features of the block in the media clip;

an attention module that:

computes a first set of data based on the plurality of feature maps;

executes one or more convolution operations to compute a plurality of confidence values corresponding to the plurality of blocks in the media clip based on the first set of data;

generates a plurality of weight values corresponding to the plurality of blocks based on the plurality of confidence values; and

computes a weighted average based on the plurality of weight values to generate a global feature vector; and

an output layer that identifies, based on the global feature vector from the attention module, a first block included in the plurality of blocks that includes first audio data that is misaligned with corresponding first video data.

2. The neural network system of claim 1 , wherein the convolutional subnetwork comprises:

a plurality of three-dimensional convolutional networks that share a first set of weights, wherein each three-dimensional convolutional network included in the plurality of three-dimensional convolutional networks includes one or more audio feature extraction layers, one or more video feature extraction layers, and one or more joint audio/video extraction layers.

3. The neural network system of claim 1 , wherein the attention module comprises:

a convolution layer that performs the one or more convolution operations on the first set of data to generate the plurality of confidence values;

a softmax layer that generates the plurality of weight values based on the plurality of confidence values, wherein the softmax layer performs a normalization operation on the plurality of confidence values to generate the plurality of weight values; and

a weighted averaging layer that generates the global feature vector for the media clip based on the plurality of weight values, wherein the global feature vector indicates that the first block in the plurality of blocks includes the first audio data that is misaligned with the corresponding first video data.

4. The neural network system of claim 3 , wherein the attention module further comprises:

a global pooling layer that generates a different feature vector for each feature map included in the plurality of feature maps to produce the first set of data, wherein the convolution layer computes each confidence value included in the plurality of confidence values based on a corresponding feature vector, and wherein the global feature vector indicates a first region within the first block in the plurality of blocks where the first audio data is misaligned with the corresponding first video data.

5. The neural network system of claim 1 , wherein the output layer includes one or more fully connected layers that perform a binary classification based on the plurality of weight values to generate a first classification for the first block, wherein the first classification indicates that the first audio data is misaligned with the corresponding first video data.

6. The neural network system of claim 1 , wherein a first feature map corresponding to the first block of the plurality of blocks includes a first joint audio/visual feature that is derived from a first audio feature and a first video feature.

7. The neural network system of claim 6 , wherein the first audio feature corresponds to a first sound that is played during playback of the media clip, the first video feature corresponds to a first event that is depicted during playback of the media clip, and the first sound is played back in conjunction with the first event in the absence of misalignment between audio data and corresponding video data.

8. The neural network system of claim 6 , wherein the first audio feature corresponds to a first sound that is played during playback of the media clip, the first video feature corresponds to a first event that is depicted during playback of the media clip, and the first sound is not played back in conjunction with the first event in the presence of misalignment between audio data and corresponding video data.

9. The neural network system of claim 6 , wherein the first video feature corresponds to a first intersection between a first set of pixels and a second set of pixels, and the first audio feature corresponds to a sound associated with the first intersection.

10. The neural network system of claim 1 , wherein at least one of the convolutional subnetwork, the attention module, or the output layer is trained based on a first set of media clips that do not include misalignment between audio data and corresponding video data and a second set of media clips that include misalignment between audio data and corresponding video data.

11. The neural network system of claim 1 , wherein the attention module comprises a softmax layer that performs a normalization operation on the plurality of confidence values to generate the plurality of weight values.

12. A computer-implemented method, comprising:

identifying, via a neural network, one or more blocks in a media clip that include audio data that is misaligned with corresponding video data, wherein the neural network is configured to:

generate a plurality of feature maps by generating, for each block included in a plurality of blocks in the media clip, a corresponding feature map derived from both one or more audio features and also one or more video features of the block in the media clip,

compute a first set of data based on the plurality of feature maps,

execute one or more convolution operations to compute a plurality of confidence values corresponding to the plurality of blocks in the media clip,

generate a plurality of weight values corresponding to the plurality of blocks based on the plurality of confidence values,

compute a weighted average based on the plurality of weight values to generate a global feature vector, and

identify, based on the global feature vector, a first block included in the plurality of blocks that includes first audio data that is misaligned with corresponding first video data.

13. The computer-implemented method of claim 12 , further comprising:

performing the one or more convolution operations on the first set of data to generate the plurality of confidence values;

generating the plurality of weight values based on the plurality of confidence values computed, wherein a softmax layer performs a normalization operation on the plurality of confidence values to generate the plurality of weight values; and

generating the global feature vector for the media clip based on the plurality of weight values generated, wherein the global feature vector indicates that the first block in the plurality of blocks includes the first audio data that is misaligned with the corresponding first video data.

14. The computer-implemented method of claim 13 , further comprising generating, via global average pooling, a different feature vector for each feature map included in the plurality of feature maps to produce the first set of data, wherein computing each confidence value included in the plurality of confidence values is performed based on a corresponding feature vector, and wherein the global feature vector indicates a first region within the first block in the plurality of blocks where the first audio data is misaligned with the corresponding first video data.

15. The computer-implemented method of claim 12 , wherein a first feature map corresponding to the first block of the plurality of blocks includes a first joint audio/visual feature that is derived from a first audio feature and a first video feature.

16. The computer-implemented method of claim 15 , wherein the first audio feature corresponds to a first sound that is played during playback of the media clip, the first video feature corresponds to a first event that is depicted during playback of the media clip, and the first sound is played back in conjunction with the first event in the absence of misalignment between audio data and corresponding video data.

17. A non-transitory computer-readable medium storing program instructions that, when executed by a processor, cause the processor to perform the steps of:

generating a plurality of feature maps by generating, for each block included in a plurality of blocks in a media clip, a corresponding feature map derived from both one or more audio features and also one or more video features of the block in the media clip;

computing a first set of data based on the plurality of feature maps;

executing one or more convolution operations to compute a plurality of confidence values corresponding to the plurality of blocks in the media clip;

generating a plurality of weight values corresponding to the plurality of blocks based on the plurality of confidence values;

computing a weighted average based on the plurality of weight values to generate a global feature vector; and

identifying, based on the global feature vector, a first block included in the plurality of blocks that includes first audio data that is misaligned with corresponding first video data.

18. The non-transitory computer-readable medium of claim 17 , wherein a first audio feature corresponds to a first sound that is played during playback of the media clip, a first video feature corresponds to a first event that is depicted during playback of the media clip, and the first sound is played back in conjunction with the first event in the absence of misalignment between audio data and corresponding video data.

19. The non-transitory computer-readable medium of claim 18 , wherein the first audio feature corresponds to a first sound that is played during playback of the media clip, the first video feature corresponds to a first event that is depicted during playback of the media clip, and the first sound is not played back in conjunction with the first event in the presence of misalignment between audio data and corresponding video data.

20. The non-transitory computer-readable medium of claim 18 , wherein the first video feature corresponds to a first intersection between a first set of pixels and a second set of pixels, and the first audio feature corresponds to a sound associated with the first intersection.

21. The non-transitory computer-readable medium of claim 17 , wherein the program instructions further comprise program instructions that, when executed by the processor, cause the processor to perform a step of training a neural network to identify the first block based on a first set of media clips that do not include misalignment between audio data and corresponding video data and a second set of media clips that include misalignment between audio data and corresponding video data.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Nov 21, 2019
From: PURI, ROHIT; KHOSRAVAN, NAJI; BEHROSTAGHI, SHERVIN ARDESHIR
To: NETFLIX, INC.
Reel/Frame 051073/0287 →
Continuity (2)
Provisional Application 62769515 · Nov 19, 2018
Related Publication 20200160889A1 · May 21, 2020
References Cited (37)
US 5430485A · Lankford et al. · 1995 [cited by applicant]
US 6122668A · Teng et al. · 2000 [cited by applicant]
US 9836853B1 · Medioni · 2017 [cited by examiner]
US 20080193016A1 · Lim · 2008 [cited by examiner]
US 20120033949A1 · Lu et al. · 2012 [cited by applicant]
US 20130089204A1 · Kumar · 2013 [cited by examiner]
US 20140116231A1 · Leflore · 2014 [cited by examiner]
US 20160314789A1 · Marcheret et al. · 2016 [cited by applicant]
US 20170177972A1 · Cricri · 2017 [cited by examiner]
US 20170178346A1 · Ferro et al. · 2017 [cited by applicant]
US 20180018970A1 · Heyl · 2018 [cited by examiner]
US 20190236136A1 · Sigal · 2019 [cited by examiner]
US 20200051254A1 · Habibian · 2020 [cited by examiner]
US 20200076988A1 · Aides · 2020 [cited by examiner]
US 20200293884A1 · Zhang · 2020 [cited by examiner]
WO 2016100814A1 · 2016 [cited by applicant]
International Search Report for application No. PCT/US2019/062240 dated Feb. 14, 2020. [cited by applicant]
Morgado et al., “Self-Supervised Generation of Spatial Audio for 360 Video”, Cornell University Library, 201, Olin Library Cornell University Ithaca, NY14853, Sep. 7, 2018, pp. 1-10. [cited by applicant]
Abolghasemi et al., “Pay attention!—Robustifying a Deep Visuomotor Policy through Task-Focused Attention”, arXiv preprint arXiv:1809.10093, DOI 10.1109/CVPR.2019.00438, IEEE/CVF Conference on Computer Vision and Pattern… [cited by applicant]
Bahdanau et al., “Neural Machine Translation by Jointly Learning to Align and Translate”, In International Conference on Learning Representations, arXiv:1409.0473, 2015, pp. 1-15. [cited by applicant]
Chung et al., “Lip Reading Sentences in the Wild”, In CVPR, 2017, pp. 3444-3453. [cited by applicant]
Chung et al., “Out of time: automated lip sync in the wild”, In Asian Conference on Computer Vision, 2016, pp. 251-263. [cited by applicant]
Gemmeke et al., “Audio Set: An Ontology and Human-Labeled Dataset for Audio Events”, In Acoustics, Speech and Signal Processing (ICASSP), 2017 IEEE International Conference on, 2017, pp. 776-780. [cited by applicant]
Liu et al., “RankIQA: Learning from Rankings for No-reference Image Quality Assessment”, Computer Vision and Pattern Recognition, https://arxiv.org/abs/1707.08347, 2017, pp. 1040-1049. [cited by applicant]
Liu et al., “Leveraging Unlabeled Data for Crowd Counting by Learning to Rank”, 2018 IEEE/CVF Conference on Computer Vision and Pattern Recognition, arXiv preprint arXiv:1803.03095, DOI 10.1109/CVPR.2018.00799, 2018, pp… [cited by applicant]
Luong et al., “Effective Approaches to Attention-based Neural Machine Translation”, arXiv preprint arXiv:1508.04025, 2015, 34 pages. [cited by applicant]
Marcheret et al., “Detecting Audio-Visual Synchrony Using Deep Neural Networks”, In Sixteenth Annual conference of the International Speech Communication Association, Sep. 6-10, 2015, pp. 548-552. [cited by applicant]
Mazaheri et al., “Video Fill in the Blank using LR/RL LSTMs with Spatial-Temporal Attentions”, 2017 IEEE International Conference on Computer Vision (ICCV), 10.1109/ICCV2017.157, 2017, pp. 1407-1416. [cited by applicant]
Noroozi et al., “Unsupervised Learning of Visual Representations by Solving Jigsaw Puzzles”, In European Conference on Computer Vision, arXiv:1603.0924, 2017, pp. 69-84. [cited by applicant]
Owens et al., “Audio-Visual Scene Analysis with Self-Supervised Multisensory Features”, arXiv preprint arXiv:1804.03641, 2018, pp. 1-18. [cited by applicant]
Sharma et al., “Action Recognition using Visual Attention”, arXiv preprint arXiv:1511.04119, 2015, pp. 1-6. [cited by applicant]
Sukhbaatar et al., “End-To-End Memory Networks”, In Advances in neural information processing systems, arXiv:1503.08895, 2015, pp. 2440-2448. [cited by applicant]
Vaswani et al., “Attention Is All You Need”, In Advances in Neural Information Processing Systems, arXiv:1706.03762, 2017, pp. 5998-6008. [cited by applicant]
Weston et al., “Memory Networks”, arXiv:1410.3916, ICLR 2015, 2014, pp. 1-15. [cited by applicant]
Xu et al., “Show, Attend and Tell: Neural Image Caption Generation with Visual Attention”, In International Conference on Machine Learning, vol. 37, 2015, pp. 2048-2057. [cited by applicant]
Yang et al., “Stacked Attention Networks for Image Question Answering”, In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, DOI 10.1109/CVPR.2016.10, 2016, pp. 21-29. [cited by applicant]
Zang et al., Attention-Based Temporal Weighted Convolutional Neural Network for Action Recognition, In IFIP International Conference, https://doi.org/10.1007/978-3-319-92007-8_9, 2018, pp. 97-108. [cited by applicant]