IP Library › Granted Patent US 12,732,657
Granted Patent B2
US 12,732,657 · App. 18/822,424 · Granted Sep 8, 2026

Determining video provenance utilizing deep learning

Inventors: Alexander Black (London, GB); Van Tu Bui (Guildford, GB); John Collomosse (Woking, GB); Simon Jenni (Hagendorf, CH); Viswanathan Swaminathan (Hagendorf, CH)
Assignees: Adobe Inc.; University of Surrey
H04N21/4341G06F16/732G06F16/7867H04N21/84H04N21/8456
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,732,657
App. No.
18/822,424
Granted
Sep 8, 2026
Kind
B2
Abstract

The present disclosure relates to systems, methods, and non-transitory computer readable media that utilize deep learning to map query videos to known videos so as to identify a provenance of the query video or identify editorial manipulations of the query video relative to a known video. For example, the video comparison system includes a deep video comparator model that generates and compares visual and audio descriptors utilizing codewords and an inverse index. The deep video comparator model is robust and ignores discrepancies due to benign transformations that commonly occur during electronic video distribution.

Claims (50)

1 . A computer-implemented method comprising:

receiving a query to determine provenance information indicating a known video stored in a database for a query video, the query video comprising a modification from the known video; and

determining, from a plurality of known videos, the known video that has been manipulated to generate the query video by:

sub-dividing the query video into visual segments and audio segments;

generating visual descriptors for the visual segments of the query video utilizing a visual neural network encoder;

generating audio descriptors for the audio segments of the query video utilizing an audio neural network encoder; and

determining that video segments from the known video are similar to the query video based on the visual descriptors and audio descriptors.

2 . The computer-implemented method of claim 1 , further comprising generating one or more visual indicators identifying visual changes in the query video relative to the known video.

3 . The computer-implemented method of claim 2 , further comprising displaying the query video with the one or more visual indicators overlaid on frames of the query video to indicate locations of the visual changes in the query video relative to the known video.

4 . The computer-implemented method of claim 2 , further comprising classifying one or more changes to the query video relative to the known video as benign changes or editorial changes.

5 . The computer-implemented method of claim 4 , wherein generating the one or more visual indicators comprises generating the one or more visual indicators for the editorial changes.

6 . The computer-implemented method of claim 5 , wherein generating the one or more visual indicators comprises ignoring the benign changes.

7 . The computer-implemented method of claim 2 , further comprising generating a heat map indicating locations of the visual changes in the query video.

8 . The computer-implemented method of claim 7 , wherein generating the heat map indicating locations of the visual changes in the query video comprises:

extracting one or more feature maps from the query video;

extracting one or more feature maps from the known video; and

generating the heat map from a combination of the one or more feature maps from the query video and the one or more feature maps from the known video.

9 . A non-transitory computer readable medium comprising instructions that, when executed by at least one processor, cause the at least one processor to perform operations comprising:

receiving a query to determine provenance information indicating a known video stored in a database for a query video, the query video comprising a modification from the known video; and

determining, from a plurality of known videos, the known video that has been manipulated to generate the query video by:

sub-dividing the query video into visual segments and audio segments;

generating visual descriptors for the visual segments of the query video utilizing a visual neural network encoder;

generating audio descriptors for the audio segments of the query video utilizing an audio neural network encoder; and

determining that video segments from the known video are similar to the query video based on the visual descriptors and audio descriptors.

10 . The non-transitory computer readable medium of claim 9 , wherein the operations further comprise generating one or more visual indicators identifying visual changes in the query video relative to the known video.

11 . The non-transitory computer readable medium of claim 10 , wherein the operations further comprise displaying the query video with the one or more visual indicators overlaid on frames of the query video to indicate locations of the visual changes in the query video relative to the known video.

12 . The non-transitory computer readable medium of claim 9 , further comprising instructions that, when executed by the at least one processor, cause the at least one processor to determine the known video based on the query video, wherein the query video comprises at least one of the following modifications relative to the known video: warping, blurring, a modification of a manifest, or removal of metadata identifying the known video.

13 . The non-transitory computer readable medium of claim 9 , wherein determining that the video segments from the known video are similar to the query video based on the visual descriptors and the audio descriptors comprises:

mapping the visual descriptors and the audio descriptors to codewords; and

identifying the video segments based on the mapped codewords.

14 . The non-transitory computer readable medium of claim 13 , wherein the operations further comprise fusing the visual descriptors and audio descriptors prior to mapping the visual descriptors and audio descriptors to the codewords.

15 . The non-transitory computer readable medium of claim 13 , wherein mapping the visual descriptors and the audio descriptors to the codewords comprises:

mapping the visual descriptors to visual codewords; and

mapping the audio descriptors to audio codewords.

16 . The non-transitory computer readable medium of claim 13 , wherein:

the operations further comprise generating unified audio-visual embeddings from corresponding visual and audio descriptors utilizing a fully connected neural network layer; and

mapping the visual descriptors and audio descriptors to the codewords comprises mapping unified audio-visual embeddings to a codebook.

17 . A system comprising:

one or more memory devices; and

one or more processors coupled to the one or more memory devices, wherein the one or more processors cause the system to perform operations comprising:

receiving a query to determine provenance information indicating a known video stored in a database for a query video, the query video comprising a modification from the known video;

determining, from a plurality of known videos, the known video that has been manipulated to generate the query video;

generating one or more visual indicators identifying visual changes in the query video relative to the known video; and

displaying the query video with the one or more visual indicators overlaid on frames of the query video to indicate locations of the visual changes in the query video relative to the known video.

18 . The system of claim 17 , wherein the operations further comprise:

extracting one or more feature maps from the query video;

extracting one or more feature maps from the known video; and

generating a heat map from a combination of the one or more feature maps from the query video and the one or more feature maps from the known video, wherein the heat map indicates the locations of the visual changes in the query video relative to the known video.

19 . The system of claim 18 , wherein the operations further comprise classifying one or more changes to the query video relative to the known video as benign changes or editorial changes from the combination of the one or more feature maps from the query video and the one or more feature maps from the known video utilizing one or more neural network layers.

20 . The system of claim 19 , wherein generating the one or more visual indicators comprises ignoring the benign changes.

Assignments (3)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Sep 11, 2024
From: COLLOMOSSE, JOHN; JENNI, SIMON; SWAMINATHAN, VISWANATHAN
To: ADOBE INC.
Reel/Frame 068558/0728 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Sep 11, 2024
From: BUI, TU
To: UNIVERSITY OF SURREY
Reel/Frame 068558/0788 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Sep 11, 2024
From: BLACK, ALEXANDER
To: UNIVERSITY OF SURREY
Reel/Frame 068558/0872 →
Continuity (2)
Continuation 17822573 · Aug 26, 2022
Related Publication 20240430515A1 · Dec 26, 2024
References Cited (140)
US 7813822B1 · Hoffberg · 2010 [cited by examiner]
US 9570107B2 · Boiman · 2017 [cited by examiner]
US 10825227B2 · Amer · 2020 [cited by examiner]
US 11601442B2 · Sekar et al. · 2023 [cited by applicant]
US 12400104B1 · Passban · 2025 [cited by examiner]
US 20080159622A1 · Agnihotri · 2008 [cited by examiner]
US 20190197187A1 · Zhang · 2019 [cited by examiner]
US 20200169591A1 · Ingel · 2020 [cited by examiner]
US 20200213680A1 · Ingel · 2020 [cited by examiner]
US 20220027732A1 · Arslan et al. · 2022 [cited by applicant]
US 20220215205A1 · Swaminathan et al. · 2022 [cited by applicant]
US 20220300740A1 · Sahu · 2022 [cited by examiner]
US 20230306721A1 · Karlinsky et al. · 2023 [cited by applicant]
US 20240273852A1 · Dong · 2024 [cited by applicant]
CA 3034323A1 · 2018 [cited by applicant]
CN 112966127B · 2022 [cited by applicant]
WO 2022005653A1 · 2022 [cited by applicant]
Georgios Tzimiropoulos et al., (hereinafter Tzimiropoulos) “A Transfer Learning Approach to Heatmap Regression for Action Unit Intensity Estimation”; © 2020 IEEE (Year: 2020). [cited by examiner]
Alexander Black; “Deep Image Comparator: Learning to Visualize Editorial changes”, CVSSP 2021, University of Surrey (Year: 2021). [cited by examiner]
Lorenzo De Donato; “Deep Learning for Audio detection and Video Analysis in Railway Applications”, Scuola Politecnica e delle Scienze, 2019/2020 (Year: 2020). [cited by examiner]
Bolei Zhou “Interpretable Representation Learning for Visual Intelligence”; © Massachusetts Institute of Technology, Jun. 2018. (Year: 2018). [cited by examiner]
U.S. Appl. No. 17/804,376, Sep. 9, 2025, Notice of Allowance. [cited by applicant]
Wang, Junke, et al. “Objectformer for image manipulation detection and localization.” Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. Mar. 29, 2022. (Year: 2022). [cited by applicant]
U.S. Appl. No. 17/804,376, May 14, 2025, Office Action. [cited by applicant]
A. Bharati, D. Moreira, P.J. Flynn, A. de Rezende Rocha, K.W. Bowyer, and W.J. Scheirer. 2021. Transformation-Aware Embeddings for Image Provenance. IEEE Trans. Info. Forensics and Sec. 16 (2021), 2493-2507. [cited by applicant]
A. Black, T. Bui, H. Jin, V. Swaminathan, and J. Collomosse. 2021. Deep Image Comparator: Learning To Visualize Editorial Change. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVP… [cited by applicant]
A. Ilyas, L. Engstrom, A. Athalye, and J. Lin. Black-box adversarial attacks with limited queries and information. In ICML, 2018. [cited by applicant]
A. Krizhevsky, I. Sutskever, and G. Hinton. Imagenet classification with deep convolutional neural networks. Communications of the ACM, 60(6):84-90, 2017. [cited by applicant]
A. Gordo, J. Almazan, J. Revaud, and D. Larlus. Deep image retrieval: Learning global representations for image search. In Proc. ECCV, pp. 241-257, 2016. [cited by applicant]
Aleksander Madry, Aleksandar Makelov, Ludwig Schmidt, Dimitris Tsipras, and Adrian Vladu. Towards deep learning models resistant to adversarial attacks. In ICLR, 2018. [cited by applicant]
Alex Tamkin, Mike Wu, and Noah Goodman. Viewmaker networks: Learning views for unsupervised representation learning. In ICLR, 2021. URL https://openreview.net/forum?id=enoVQWLsfyL. [cited by applicant]
Alexey Kurakin, Ian Goodfellow, and Samy Bengio. Adversarial examples in the physical world. arXiv preprint arXiv:1607.02533, 2016. [cited by applicant]
Andreas Geiger, Philip Lenz, Christoph Stiller, and Raquel Urtasun. Vision meets robotics: The kitti dataset. International Journal of Robotics Research (IJRR), 2013. [cited by applicant]
Andrew Rouditchenko et al., “Self-Supervised Audio-Visual Co-Segmentation”; 978-1-5386-4658-8/18, © 2019 IEEE (Year: 2019). [cited by applicant]
Anish Athalye and Ilya Sutskever. Synthesizing robust adversarial examples. arXiv preprint arXiv:1707.07397, 2017. [cited by applicant]
B. Dolhansky, J. Bitton, B. Pflaum, J. Lu, R. Howes, M. Wang, and C. C. Ferrer. The deepfake detection challenge (DFDC) dataset. CoRR, abs/2006.07397, 2020. [cited by applicant]
Brian Dolhansky and Cristian Canton Ferrer. Adversarial collision attacks on image hashing functions. CVPR Workshop on Adversarial Machine Learning, 2021. [cited by applicant]
C. Jacobs, A. Finkelstein, and D. Salesin. 1995. Fast multiresolution image querying. In Proc. ACM SIGGRAPH. ACM, 277-286. [cited by applicant]
C. Szegedy, W. Liu, Y. Jia, P. Sermanet, S. Reed, D. Anguelov, D. Erhan, V. Vanhoucke, and A. Rabinovich. Going deeper with convolutions. In Proceedings of the IEEE conference on computer vision and pattern recognition,… [cited by applicant]
C. Zauner. Implementation and benchmarking of perceptual image hash functions. Master's thesis, Upper Austria University of Applied Sciences, Hagenberg, 2010. [cited by applicant]
Christian Szegedy, Wojciech Zaremba, Ilya Sutskever, Dumitru Erhan Joan Bruna, Ian Goodfellow, and Rob Fergus. Intriguing properties of neural networks. In ICLR, 2013. [cited by applicant]
Coalition for Content Provenance and Authenticity. 2021. Draft Technical Specification 0.7. Technical Report. C2PA. https://c2pa.org/public-draft/. [cited by applicant]
D. Brian, H. Russ, P. Ben, B. Nicole, and F. C. Canton. 2019. The deepfake detection challenge (dfdc) preview dataset. arXiv preprint arXiv:1910.08854 (2019). [cited by applicant]
D. Hendrycks and T. Dietterich. Benchmarking neural network robustness to common corruptions and perturbations. In Proc. ICLR, 2019. [cited by applicant]
D. Moreira, A. Bharati, J. Brogan, A. Pinto, M. Parowski, K.W. Bowyer, P.J. Flynn, A. Rocha, and W.J. Scheirer. 2018. Image provenance analysis at scale. IEEE Trans. Image Proc. 27, 12 (2018), 6109-6122. [cited by applicant]
D. Profrock, M. Schlauweg, and E. Muller. Content-based watermarking by geometric wrapping and feature-based image segmentation. In Proc. SITIS, pp. 572-581, 2006. [cited by applicant]
D. Tsipras, S. Santurkar, L. Engstrom, A. Turner, and A. Madry. Robustness may be at odds with accuracy. In ICLR, 2019. [cited by applicant]
Dan Hendrycks, Norman Mu, Ekin D. Cubuk, Barret Zoph, Justin Gilmer, and Balaji Lakshminarayanan. Augmix: A simple data processing method to improve robustness and uncertainty. ICLR, 2020. [cited by applicant]
Dan Hendrycks, Steven Basart, Norman Mu, Saurav Kadavath, Frank Wang, Evan Dorundo, Rahul Desai, Tyler Zhu, Samyak Parajuli, Mike Guo, et al. The many faces of robustness: A critical analysis of out-of-distribution gene… [cited by applicant]
Dou Goodman, Hao Xin, Wang Yang, Wu Yuesheng, Xiong Junfeng, and Zhang Huan. Advbox: a toolbox to generate adversarial examples that fool neural networks. arXiv preprint arXiv:2001.05574, 2020. [cited by applicant]
E. J. Humphrey and J. P. Bello. 2012. Rethinking automatic chord recognition with convolutional neural networks. In Proc. Intl. Conf. on Machine Learning and Applications. [cited by applicant]
E. Nguyen, T. Bui, V. Swaminathan, and J. Collomosse. 2021. OSCAR-Net: Objectcentric Scene Graph Attention for Image Attribution. In Proc. ICCV. [cited by applicant]
Eric Wong, Leslie Rice, and J. Zico Kolter. Fast is better than free: Revisiting adversarial training. ICLR, 2020. [cited by applicant]
F. Khelifi and A. Bouridane. Perceptual video hashing for content identification and authentication. IEEE TCSVT, 1(29), 2019. [cited by applicant]
F. Rigaud and M. Radenen. 2016. Singing voice melody transcription using deep neural networks. In Proc. Intl. Conf. on Music Information Retreival (ISMIR). [cited by applicant]
F. Zheng, G. Zhang, and Z. Song. 2001. Deep convolutional neural networks for predominant instrument recognition in polyphonic music. J. Computer Science and Technology 16, 6 (2001), 582-589. [cited by applicant]
G. Tzanetakis and P. Cook. 2002. Musical genre classification of audio signals. IEEE Trans. on Audio and Speech Proc. (2002). [cited by applicant]
Gavin Weiguang Ding, Luyu Wang, and Xiaomeng Jin. AdverTorch v0.1: An adversarial robustness toolbox based on pytorch. arXiv preprint arXiv:1902.07623, 2019. [cited by applicant]
Giorgos Tolias, Filip Radenovic, and Ondrej Chum. Targeted mismatch adversarial attack: Query with a flower to retrieve the tower. In Proceedings of the IEEE/CVF International Conference on Computer Vision, pp. 5037-504… [cited by applicant]
Github.com—ISCC—Specification v1.0.0—Date downloaded Aug. 19, 2022; https://github.com/iscc/iscc-specs/blob/version-1.0/docs/specification.md. [cited by applicant]
H. Chen, W. Xie, A. Vedaldi, and A. Zisserman. 2020b. VGG-Sound: A large scale audio-visual dataset. In Proc. Intl. Conf. on Acoustics, Speech and Signal Processing (ICASSP). [cited by applicant]
H. Lee, P. Pham, Y. Largman, and A. Y. Ng. 2009. Unsupervised feature learning for audio classification using convolutional deep belief networks. In Proc. Advances in Neural Information Processing Systems (NIPS). [cited by applicant]
H. Liu, R. Wang, S. Shan, and X. Chen. Deep supervised hashing for fast image retrieval. In Proc. CVPR, pp. 2064-2072, 2017. [cited by applicant]
H. Shawn, C. Sourish, E. Daniel PW, G. Jort F, J. Aren, M. R. Channing, P. Manoj, P. Devin, S. Rif A, S. Bryan, et al. 2017. CNN architectures for large-scale audio classification. In Proc. Intl. Conf. on Acoustics, Spe… [cited by applicant]
H. Zhu, M. Long, J. Wang, and Y. Cao. Deep hashing network for efficient similarity retrieval. In Proc. AAAI, 2016. [cited by applicant]
Hongyang Zhang, Yaodong Yu, Jiantao Jiao, Eric P. Xing, Laurent El Ghaoui, and Michael I. Jordan. Theoretically principled trade-off between robustness and accuracy. In ICML, 2019. [cited by applicant]
I. Goodfellow, J. Shlens, and C. Szegedy. Explaining and harnessing adversarial examples. arXiv preprint arXiv:1412.6572v3, 2014. [cited by applicant]
IPTC Council. Social media sites photo metadata test results. http://embeddedmetadata.org/ social-media-test-results.php, 2020. [cited by applicant]
J. Aythora et al. Multi-stakeholder media provenance management to counter synthetic media risks in news publishing. In Proc. Intl. Broadcasting Convention (IBC), 2020. [cited by applicant]
J. Buchner. Imagehash. https://pypi.org/ project/ImageHash/, 2021. [cited by applicant]
J. Collomosse, T. Bui, A. Brown, J. Sheridan, A. Green, M. Bell, J. Fawcett, J. Higgins, and O. Thereaux. ARCHANGEL: Trusted archives of digital public docu- ments. In Proc. ACM Doc.Eng, 2018. [cited by applicant]
J. Deng, W. Dong, R. Socher, L.-J. Li, K. Li, and L. Fei-Fei. ImageNet: A Large-Scale Hierarchical Image Database. In Proc. CVPR, 2009. [cited by applicant]
J. Johnson, M. Douze, and H. Jegou. Billion-scale similarity search with gpus. IEEE Transactions on Big Data, 2017. [cited by applicant]
J. Lee, J. Park, K. L. Kim, and J. Nam. 2017. Sample-level deep convolutional neural networks for music auto-tagging using raw waveforms. arXiv preprint arXiv:1703.01789 (2017). [cited by applicant]
J. S. Downie. 2003. Music Information Retrieval. Annual review of information science and technology 37, 1 (2003), 295-340. [cited by applicant]
J. Schluter and S. Bock. 2013. Musical onset detection with convolutional neural networks. In Proc. Intl. Workshop on Machine Learning and Music (MML). [cited by applicant]
J. Wang, H. T. Shen, J. Song, and J.Ji. Hashing for similarity search: A survey. arXiv preprint arXiv:1408.2927, 2014. [cited by applicant]
Jonas Rauber, Wieland Brendel, and Matthias Bethge. Foolbox: A python toolbox to benchmark the robustness of machine learning models. In ICML Reliable Machine Learning in the Wild Workshop, 2017. [cited by applicant]
K. Choi, G. Fazekas, K. Cho, and M. Sandler. 2018. A Tutorial on Deep Learning for Music Information Retrieval. arXiv:1709.04396v2 (2018). [cited by applicant]
K. Choi, G. Fazekas, M. Sandler, and K. Cho. 2017. Convolutional recurrent neural networks for music classification. In Proc. Intl. Conf. on Acoustics, Speech and Signal Processing (ICASSP). [cited by applicant]
K. Hameed, A. Mumtax, and S. Gilani. 2006. Digital image watermarking in the wavelet transform domain. WASET 13 (2006), 86-89. [cited by applicant]
K. He, X. Zhang, S. Ren, and J. Sun. Deep residual learning for image recognition. In Proc. CVPR, pp. 770-778, 2015. [cited by applicant]
Karen Sparck Jones. 1972. A statistical interpretation of term specificity and its application in retrieval. Journal of documentation (1972). [cited by applicant]
Kevin Eykholt, Ivan Evtimov, Earlence Fernandes, Bo Li, Dawn Song, Tadayoshi Kohno, Amir Rahmati, Atul Prakash, and Florian Tramer. Note on attacking object detectors with adversarial stickers. arXiv preprint arXiv:1712… [cited by applicant]
L. Rosenthol, A. Parsons, E. Scouten, J. Aythora, B. MacCormack, P. England, M. Levallee, J. Dotan, et al. 2020. Content Authenticity Initiative (CAI): Setting the Standard for Content Attribution. Technical Report. Ado… [cited by applicant]
L. Yuan, T. Wang, X. Zhang, F. Tay, Z. Jie, W. Liu, and J. Feng. Central similarity quantization for efficient image and video retrieval. In Proc. CVPR, pp. 3083-3092, 2020. [cited by applicant]
L.-C. Yang, S.-Y. Chou, J.-Y. Liu, Y.-H. Yang, and Y.-A. Chen. 2017. Revisiting the problem of audio-based hit song prediction using convolutional neural networks. arXiv preprint arXiv:1704.01280 (2017). [cited by applicant]
Logan Engstrom, Andrew Ilyas, Shibani Santurkar, Dimitris Tsipras, Brandon Tran, and Aleksander Madry. Adversarial robustness as a prior for learned representations. arXiv preprint arXiv:1906.00945, 2019. [cited by applicant]
M. Douze, G. Tolias, E. Pizzi, Z. Papakipos, L. Chanussot, F. Radenovic, T. Jenicek, M. Maximov, L. Leal-Taixé, I. Elezi, O. Chum, and C. C. Ferrer. 2021. The 2021 Image Similarity Dataset and Challenge. CoRR abs/2106.0… [cited by applicant]
M. Huh, A. Liu, A. Owens, and A. Efros. Fighting fake news: Image splice detection via learned self-consistency. In Proc. ECCV, 2018. [cited by applicant]
Maksym Andriushchenko and Nicolas Flammarion. Understanding and improving fast adversarial training. NeurIPS, 2020. [cited by applicant]
Maksym Andriushchenko, Francesco Croce, Nicolas Flammarion, and Matthias Hein. Square attack: a query-efficient black-box adversarial attack via random search. In ECCV, 2020. [cited by applicant]
Marco Melis, Ambra Demontis, Maura Pintor, Angelo Sotgiu, and Battista Biggio. secml: A python library for secure and explainable machine learning. arXiv preprint arXiv:1912.10013, 2019. [cited by applicant]
Maria-Irina Nicolae, Mathieu Sinn, Minh Ngoc Tran, Beat Buesser, Ambrish Rawat, Martin Wistuba, Valentina Zantedeschi, Nathalie Baracaldo, Bryant Chen, Heiko Ludwig, Ian Molloy, and Ben Edwards. Adversarial robustness t… [cited by applicant]
Minseon Kim, Jihoon Tack, and Sung Ju Hwang. Adversarial self-supervised contrastive learning. NeurIPS, 2020. [cited by applicant]
N. Papernot, P. McDaniel, and I. Goodfellow. Transferability in machine learning: from phenomena to black-box attacks using adversarial samples. arXiv preprint arXiv:1605.07277, 2016. [cited by applicant]
N. Yu, L. Davis, and M. Fritz. Attributing fake images to gans: Learning and analyzing gan fingerprints. In IEEE International Conference on Computer Vision (ICCV), 2019. [cited by applicant]
N. Yu, V. Skripniuk, S.r Abdelnabi, and M. Fritz. 2021. Artificial Fingerprinting for Generative Models: Rooting Deepfake Attribution in Training Data. In Proc. Intl. Conf. Computer Vision (ICCV). [cited by applicant]
Nicolas Papernot, Fartash Faghri, Nicholas Carlini, Ian Goodfellow, Reuben Feinman, Alexey Kurakin, Cihang Xie, Yash Sharma, Tom Brown, Aurko Roy, Alexander Matyasko, Vahid Behzadan, Karen Hambardzumyan, Zhishuai Zhang,… [cited by applicant]
P. Devi, M. Venkatesan, and K. Duraiswamy. A fragile watermarking scheme for image authentication with tamper localization using integer wavelet transform. J. Computer Science, 5(11):831-837, 2009. [cited by applicant]
P. Torr and A. Zisserman. Mlesac: A new robust estimator with application to estimating image geometry. Computer Vision Image Understanding (CVIU), 78(1):138-156, 2000. [cited by applicant]
Parsa Saadatpanah, Ali Shafahi, and Tom Goldstein. Adversarial attacks on copyright detection systems. In ICML, 2020. [cited by applicant]
Q-Y. Jiang and W-J. Li. 2018. Asymmetric Deep Supervised Hashing. In AAAI. [cited by applicant]
Q. Li, Z. Sun, R. He, and T. Tan. Deep supervised discrete hashing. In Proc. NeurIPS, pp. 2482-2491, 2017. [cited by applicant]
R. Hadsell, S. Chopra, and Y. LeCun.Dimensionality reduction by learning an invariant mapping. In Proc. CVPR, pp. 1735-1742, 2006. [cited by applicant]
R. Lu, K. Wu, Z. Duan, and C. Zhang. 2017. Deep ranking: triplet matchnet for music metric learning. In Proc. Intl. Conf. on Acoustics, Speech and Signal Processing (ICASSP). [cited by applicant]
S-Y. Wang, O. Wang, A. Owens, R. Zhang, and A. Efros. Detecting photoshopped faces by scripting photoshop. In Proc. ICCV, 2019. [cited by applicant]
S-Y. Wang, O. Wang, R. Zhang, A. Owens, and A. Efros. Cnn-generated images are surprisingly easy to spot . . . for now. In Proc. CVPR, 2020. [cited by applicant]
S. Baba, L. Krekor, T. Arif, and Z. Shaaban. Watermarking scheme for copyright protection of digital images. IJCSNS, 9(4), 2009. [cited by applicant]
S. Dieleman and B. Schrauwen. 2014. End-to-end learning for music audio. In Proc. Intl. Conf. on Acoustics, Speech and Signal Processing (ICASSP). [cited by applicant]
S. Gregory. 2019. Ticks or it didn't happen. Technical Report. Witness.org. [cited by applicant]
S. Heller, L. Rossetto, and H. Schuldt. The PS-Battles Dataset—an Image Collection for Image Manipulation Detection. CoRR, abs/1804.04866, 2018. [cited by applicant]
S. Jenni and P. Favaro. Self-supervised feature learning by learning to spot artifacts. In Proc. CVPR, 2018. [cited by applicant]
S. Sigtia, E. Benetos, and S. Dixon. 2015. An end-to-end neural network for polyphonic music transcription. arXiv preprint arXiv:1508.01774 (2015). [cited by applicant]
S. Su, C. Zhang, K. Han, and Y. Tian. 2018. Greedy hash: Towards fast optimization for accurate hash coding in CNN. In Proc. NeurIPS. 798-807. [cited by applicant]
S. Thys, W. Van Ranst, and T. Goedeme. Fooling automated surveillance cameras: adversarial patches to attack person detection. arXiv preprint arXiv:1904.08653, 2019. [cited by applicant]
S.-M. Moosavi-Dezfooli, A. Fawzi, and P. Frossard. Deepfool: a simple and accurate method to fool deep neural networks. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, pp. 2574-2582, 20… [cited by applicant]
Sepehr Valipour et al., “Recurrent Fully Convolutional Networks for Video Segmentation”: arXiv: 1606.00487v3 [cs.CV] Oct. 31, 2016 (Year: 2016). [cited by applicant]
Shang-Tse Chen, Cory Cornelius, Jason Martin, and Duen Horng Chau. ShapeShifter: Robust physical adversarial attack on faster r-CNN object detector. In Machine Learning and Knowledge Discovery in Databases, pp. 52-68. S… [cited by applicant]
Sven Gowal, Chongli Qin, Jonathan Uesato, Timothy Mann, and Pushmeet Kohli. Uncovering the limits of adversarial training against norm-bounded adversarial examples. arXiv, 2020. [cited by applicant]
T. Bertin-Mahieux, D. P. Ellis, B. Whitman, and P. Lamere. 2011. The million song dataset. In Proc. Intl. Conf of Soceity for Music Information Retrieval. 591-596. [cited by applicant]
T. Brown, D. Mane, A. Roy, M. Abadi, and J. Gilmer. Adversarial patch. arXivpreprintarXiv:1712.09665, 2017. [cited by applicant]
T. Bui, D. Cooper, J. Collomosse, M. Bell, A. Green, J. Sheridan, J. Higgins, A. Das, J. Keller, and O. Thereaux. 2020. Tamper-proofing Video with Hierarchical Attention Autoencoder Hashing on Blockchain. IEEE Trans. Mu… [cited by applicant]
T. Chen, S. Kornblith, M. Norouzi, and G. Hinton. A simple framework for contrastive learning of visual representations. In Proc. ICML, pp. 1597-1607, 2020. [cited by applicant]
T. Grill and J. Schluter. 2015. Music boundary detection using neural networks on spectrograms and self-similarity lag matrices. In Proc. EUSPICO. [cited by applicant]
T. Pan. 2019. Digital-Content-Based Identification: Similarity hashing for content identification in decentralized environments. In Proc. Blockchain for Science. [cited by applicant]
Madimir I Levenshtein et al. 1966. Binary codes capable of correcting deletions, insertions, and reversals. In Soviet physics doklady, vol. 10. Soviet Union, 707-710. [cited by applicant]
W. Li, S. Wang, and W-C. Kang. 2016. Feature learning based deep supervised hashing with pairwise labels. In Proc. IJCAI. 1711-1717. [cited by applicant]
W. Wang, J. Dong, and T. Tan. Tampered region localization of digital color images based on jpeg compression noise. In International Workshop on Digital Watermarking, pp. 120-133. Springer, 2010. [cited by applicant]
X. Zhang, Z. H. Sun, S. Karaman, and S.F. Chang. 2020. Discovering Image Manipulation History by Pairwise Relation and Forensics Tools. IEEE J. Selected Topics in Signal Processing. 14, 5 (2020), 1012-1023. [cited by applicant]
Y. Han, J. Kim, and K. Lee. 2017. Deep convolutional neural networks for predominant instrument recognition in polyphonic music. IEEE Trans. Audio, Speech and Language Processing 25, 1 (2017), 208-221. [cited by applicant]
Y. Li, M-C. Ching, and S. Lyu. In ictu oculi: Exposing ai created fake videos by detecting eye blinking. In Proc. IEEE WIFS, 2018. [cited by applicant]
Y. Li, W. Pei, and J. van Gemert. 2019. Push for Quantization: Deep Fisher Hashing. BMVC (2019). [cited by applicant]
Y. Wu, W. AbdAlmageed, and P. Natarajan. Mantra-net: Manipulation tracing network for detection and localization of image forgeries with anomalous features. In Proc. CVPR, pp. 9543-9552, 2019. [cited by applicant]
Yonglong Tian, Chen Sun, Ben Poole, Dilip Krishnan, Cordelia Schmid, and Phillip Isola. What makes for good views for contrastive learning? NeurIPS, 2020. [cited by applicant]
Z. Cao, M. Long, J. Wang, and P. S. Yu. Hashnet: Deep learning to hash by continuation. In Proc. CVPR, pp. 5608-5617, 2017. [cited by applicant]
Z. Lenyk and J. Park. Microsoft vision model resnet-50 combines web-scale data and multi-task learning to achieve state of the art. https://pypi.org/project/ microsoftvision/, 2021. [cited by applicant]
Z. Teed and J. Deng. Raft: Recurrent all-pairs field transforms for optical flow. In Proc. ECCV, pp. 402-419. Springer, 2020. [cited by applicant]
U.S. Appl. No. 17/822,573, Apr. 19, 2024, Notice of Allowance. [cited by applicant]
U.S. Appl. No. 17/573,041, Sep. 26, 2024, Notice of Allowance. [cited by applicant]