IP Library › Granted Patent US 12,277,501
Granted Patent B2
US 12,277,501 · App. 18/217,745 · Granted Apr 15, 2025

Training a sound effect recommendation network

Inventor: Sudha Krishnamurthy (Foster City, CA)
Assignee: Sony Interactive Entertainment Inc.
G06N3/084G06F16/68G06N3/045G06N20/00G10L15/16
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,277,501
App. No.
18/217,745
Granted
Apr 15, 2025
Kind
B2
Abstract

A Sound effect recommendation network is trained using a machine learning algorithm with a reference image, a positive audio embedding and a negative audio embedding as inputs to train a visual-to-audio correlation neural network to output a smaller distance between the positive audio embedding and the reference image than the negative audio embedding and the reference image. The visual-to-audio correlation neural network is trained to identify one or more visual elements in the reference image and map the one or more visual elements to one or more sound categories or subcategories within an audio database.

Claims (30)

1. A method for training a Sound Effect Recommendation Network, comprising:

a) generating a positive audio embedding from a positive audio signal wherein the positive audio signal is related to a reference image;

b) generating a negative audio embedding from a negative audio signal; and

c) using a machine learning algorithm with the reference image, the positive audio embedding and the negative audio embedding as inputs to train a visual-to-audio correlation neural network to output a smaller distance between the positive audio embedding and the reference image than the negative audio embedding and the reference image, wherein the visual-to-audio correlation neural network is trained to identify one or more visual elements in the reference image and map the one or more visual elements to one or more sound categories or subcategories within an audio database,

wherein a reference audio signal is part of the audio database and wherein the reference audio signal is part of an audio visual sequence having the reference image, and wherein the positive audio signal is the reference audio signal.

2. The method of claim 1 , further comprising determining the negative audio signal using the trained similarity neural network with the audio database wherein the negative signal is not related to the positive audio signal.

3. The method of claim 1 wherein features are extracted from the positive signal, reference signal and negative signal before being used in training with the machine learning algorithm.

4. The method of claim 1 wherein the positive audio signal includes noise.

5. The method of claim 1 wherein the machine learning algorithm used to train the visual-to-audio correlation neural network includes a pairwise loss function.

6. The method of claim 1 wherein the positive signal is part of an audio visual sequence that includes the reference image, wherein the positive signal includes noise signals and wherein the noise signals are other sounds occurring in the audio visual sequence and wherein the negative signal includes noise signals.

7. The method of claim 1 wherein the machine learning algorithm is a self-supervised learning algorithm and wherein the positive, negative and correlated audio are unlabeled or unannotated inputs.

8. A system for training a Sound Effect Recommendation Network, comprising:

a Processor;

a Memory coupled to the Processor;

Non-transitory instructions embedded in the memory that when executed cause the processor to carry out the method for training a sound effect recommendation network comprising;

a) generating a positive audio embedding from a positive audio signal wherein the positive audio signal is related to a reference image;

b) generating a negative audio embedding from a negative audio signal;

c) using a machine learning algorithm with the reference image, the positive audio embedding and the negative audio embedding as inputs to train an image-to-audio correlation neural network to output a smaller distance between the positive audio embedding and the reference image than the negative audio embedding and the reference image, wherein the visual-to-audio correlation neural network is trained to identify one or more visual elements in the reference image and map the one or more visual elements to one or more sound categories or subcategories within an audio database,

wherein a reference audio signal is part of the audio database and wherein the reference audio signal is part of an audio visual sequence having the reference image, and wherein the positive audio signal is the reference audio signal.

9. The system of claim 8 further comprising determining the negative audio signal using the trained similarity neural network with the audio database wherein the negative signal is not related to the reference audio signal.

10. The system of claim 8 wherein the audio features are extracted from the positive signal and negative signal before being used in training with the machine learning algorithm.

11. The system of claim 8 wherein the positive audio signal includes noise.

12. The system of claim 8 wherein the machine learning algorithm used to train the visual-to-audio correlation neural network includes a pairwise loss function.

13. The system of claim 8 wherein the positive signal is part of an audio visual sequence that includes the reference image, wherein positive signal includes noise signals and wherein the noise signals are other sounds occurring in the audio visual sequence and wherein the negative signal includes noise signals.

14. The method of claim 8 wherein the machine learning algorithm is a self-supervised learning algorithm and wherein the positive, negative and correlated audio are unlabeled or unannotated inputs.

15. Non-transitory instructions embedded in a computer readable medium that when executed by a computer cause the computer to carry out a method for training a Sound Recommendation Network comprising:

a) generating a positive audio embedding from a positive audio signal wherein the positive audio signal is related to a reference image;

b) generating a negative audio embedding from a negative audio signal;

c) using a machine learning algorithm with the reference image, the positive audio embedding and a negative audio embedding as inputs to train a visual-to-audio correlation neural network to output a smaller distance between the positive audio embedding and the reference image than the negative audio embedding and the reference image, wherein the visual-to-audio correlation neural network is trained to identify one or more visual elements in the reference image and map the one or more visual elements to one or more sound categories or subcategories within an audio database,

wherein a reference audio signal is part of the audio database and wherein the reference audio signal is part of an audio visual sequence having the reference image, and wherein the positive audio signal is the reference audio signal.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Aug 15, 2023
From: KRISHNAMURTHY, SUDHA
To: SONY INTERACTIVE ENTERTAINMENT INC.
Reel/Frame 064595/0381 →
Continuity (2)
Continuation 16848484 · Apr 14, 2020
Related Publication 20230385646A1 · Nov 30, 2023
References Cited (93)
US 8106284B2 · Suzuki et al. · 2012 [cited by applicant]
US 9094636B1 · Sanders et al. · 2015 [cited by applicant]
US 9113034B2 · Yue · 2015 [cited by applicant]
US 9338420B2 · Xiang · 2016 [cited by applicant]
US 9373320B1 · Lyon et al. · 2016 [cited by applicant]
US 9384214B2 · Slaney et al. · 2016 [cited by applicant]
US 9736580B2 · Cahill et al. · 2017 [cited by applicant]
US 9827496B1 · Zinno · 2017 [cited by applicant]
US 9858967B1 · Nomula et al. · 2018 [cited by applicant]
US 9961403B2 · Kritt et al. · 2018 [cited by applicant]
US 10032457B1 · Liu et al. · 2018 [cited by applicant]
US 10455297B1 · Mahyar et al. · 2019 [cited by applicant]
US 10671854B1 · Mahyar et al. · 2020 [cited by applicant]
US 11381888B2 · Krishnamurthy · 2022 [cited by applicant]
US 11501102B2 · Salamon et al. · 2022 [cited by applicant]
US 11615312B2 · Krishnamurthy · 2023 [cited by applicant]
US 20050175315A1 · Ewing · 2005 [cited by applicant]
US 20060277454A1 · Chen · 2006 [cited by applicant]
US 20070233494A1 · Shen et al. · 2007 [cited by applicant]
US 20110029873A1 · Eseanu et al. · 2011 [cited by applicant]
US 20110075851A1 · Leboeuf et al. · 2011 [cited by applicant]
US 20110173235A1 · Aman et al. · 2011 [cited by applicant]
US 20120033132A1 · Chen et al. · 2012 [cited by applicant]
US 20130073960A1 · Eppolito et al. · 2013 [cited by applicant]
US 20130163949A1 · Tanuma et al. · 2013 [cited by applicant]
US 20140096002A1 · Dey et al. · 2014 [cited by applicant]
US 20140178043A1 · Kritt et al. · 2014 [cited by applicant]
US 20140181668A1 · Kritt et al. · 2014 [cited by applicant]
US 20140228090A1 · Powell · 2014 [cited by applicant]
US 20150234833A1 · Cremer et al. · 2015 [cited by applicant]
US 20160203827A1 · Leff et al. · 2016 [cited by applicant]
US 20160292509A1 · Kaps et al. · 2016 [cited by applicant]
US 20170092001A1 · Anderson · 2017 [cited by applicant]
US 20170228599A1 · De Juan · 2017 [cited by applicant]
US 20180102143A1 · Allison et al. · 2018 [cited by applicant]
US 20180226063A1 · Wood et al. · 2018 [cited by applicant]
US 20180232451A1 · Lev-Tov et al. · 2018 [cited by applicant]
US 20180322855A1 · Lyske et al. · 2018 [cited by applicant]
US 20180358052A1 · Miller et al. · 2018 [cited by applicant]
US 20190005128A1 · Barari et al. · 2019 [cited by applicant]
US 20190035431A1 · Attorre et al. · 2019 [cited by applicant]
US 20190104259A1 · Angquist et al. · 2019 [cited by applicant]
US 20190289372A1 · Merler et al. · 2019 [cited by applicant]
US 20200104319A1 · Jati et al. · 2020 [cited by applicant]
US 20200134456A1 · Li et al. · 2020 [cited by applicant]
US 20200160889A1 · Puri et al. · 2020 [cited by applicant]
US 20200167984A1 · Cappello et al. · 2020 [cited by applicant]
US 20200213662A1 · Wolcott et al. · 2020 [cited by applicant]
US 20200242507A1 · Gan et al. · 2020 [cited by applicant]
US 20200322377A1 · Lakhdhar · 2020 [cited by examiner]
US 20200349387A1 · Krishnamurthy et al. · 2020 [cited by applicant]
US 20200349921A1 · Jansen et al. · 2020 [cited by applicant]
US 20200349975A1 · Krishnamurthy et al. · 2020 [cited by applicant]
US 20210035599A1 · Zhang et al. · 2021 [cited by applicant]
US 20210035610A1 · Krishnamurthy et al. · 2021 [cited by applicant]
CN 1622558A · 2005 [cited by applicant]
CN 102480671A · 2012 [cited by applicant]
CN 104995681A · 2015 [cited by applicant]
CN 105068798A · 2015 [cited by applicant]
CN 107223332A · 2017 [cited by applicant]
CN 109587554A · 2019 [cited by applicant]
JP 2010020133A · 2010 [cited by applicant]
JP 6442102B1 · 2018 [cited by applicant]
TW 201426729A · 2014 [cited by applicant]
Donghuo Zeng et al. “Deep Triplet Neural Networks with Cluster-CCA for Audio-Visual Cross-modal Retrieval” Arxiv.org, Cornell University Library, Dated Aug. 10, 2019, XP081459707. [cited by applicant]
Extend European Search for European Application No. 21787592.1, dated Mar. 27, 2024. [cited by applicant]
Hao, Zhu et al. “Deep Audio-Visual Learning: A Survey,” Arxiv.org, Cornell University Library, dated: Jan. 14, 2020, XP081578387. [cited by applicant]
Hung Sungeun et al. “CBVMR: Content-Based Video-Music Retrieval Using Soft Intramodal Structure Constraint,” Proceedings of the 2018 ACM on International Conference on Multimedia, Retrieved Jun. 5, 2018, pp. 353-361, XP… [cited by applicant]
Japanese Office Action for Japanese Application No. 2022-562558, dated Mar. 14, 2024. [cited by applicant]
B. Kang, Y. Kim and D. Kim, “Deep Convolutional Neural Network Using Triplets of Faces, Deep Ensemble, and Score-Level Fusion for Face Recognition,” 2017 IEEE Conference on Computer Vision and Pattern Recognition Worksh… [cited by applicant]
Di Hu et al., “Deep Multimodal Clustering for Unsupervised Audiovisual Learning”, arXiv:1807.03094v3 [cs.CV] Apr. 19, 2019. [cited by applicant]
Elad Hoffer and Nir Ailo, “Deep Metric Learning Using Triplet Network”, In: Feragen A., Pelillo M., Loog M. (eds) Similarity-Based Pattern Recognition. SIMBAD 2015. Lecture Notes in Computer Science, vol. 9370. Springer… [cited by applicant]
Final Office Action for U.S. Appl. No. 16/848,484, dated Nov. 18, 2022. [cited by applicant]
Final Office Action for U.S. Appl. No. 16/848,512, dated Jul. 27, 2021. [cited by applicant]
Hochreiter & Schmidhuber. “Long Short-term memory.” Neural Computation 9(8):1735-1780 (1997). [cited by applicant]
International Search Report and Written Opinion dated Jul. 15, 2021 for International Patent Application No. PCT/US2021/026550. [cited by applicant]
International Search Report and Written Opinion dated Jul. 15, 2021 for International Patent Application No. PCT/US2021/026553. [cited by applicant]
International Search Report and Written Opinion of the International Searching Authority for International Patent Application No. PCT/US2021/026554. [cited by applicant]
Non-Final Office Action for U.S. Appl. No. 16/848,484, dated Jun. 24, 2022. [cited by applicant]
Non-Final Office Action for U.S. Appl. No. 16/848,499, dated Jun. 27, 2022. [cited by applicant]
Non-Final Office Action for U.S. Appl. No. 16/848,512, dated Mar. 15, 2021. [cited by applicant]
Notice of Allowance for U.S. Appl. No. 16/848,484, dated Feb. 17, 2023. [cited by applicant]
Notice of Allowance for U.S. Appl. No. 16/848,499, dated Nov. 21, 2022. [cited by applicant]
Notice of Allowance for U.S. Appl. No. 16/848,512, dated Mar. 2, 2022. [cited by applicant]
Relja Arandjelovic and Andrew Zisserman, “Objects that Sound”, arXiv:1712.06651v2 [cs.CV] Jul. 25, 2018. [cited by applicant]
Xiang et al., “Person Re-identification Based on Feature Fusion and Triplet Loss Function” 2018 24th International Conference on Pattern Recognition, Aug. 20-24, 2018, pp. 2477-2482 (Year:2018). [cited by applicant]
Japanese Office Action for Japanese Application No. 2022-562558, dated Oct. 18, 2023. [cited by applicant]
Kyosuke Rinsaka, Yoshinobu Kajikawa, Yasuo Nomura, Study on Mutual Search System for Heterogeneous Media: Music Matching Images, Technical Report of the Institute of Electronics, Information and Communication Engineers,… [cited by applicant]
International Preliminary Report on Patentability in International Appln. No. PCT/US2021/026550, mailed on Oct. 13, 2022, 5 pages. [cited by applicant]
merriam-webster.com [online], “still photograph,” available on or before Nov. 18, 2022, retrieved on Jan. 9, 2025, retrieved from URL<https://www.merriam-webster.com/dictionary/still%20photograph>, 5 pages. [cited by applicant]
Owens et al., “Audio-Visual Scene Analysis with Self-Supervised Multisensory Features,” CoRR, Submitted on Oct. 9, 2018, arXiv:1804.03641v2, 19 pages. [cited by applicant]
Takeshi et al., “Music Database Retrieval System with Sensitivity Words Using Music Sensitivity Space,” Journal of the Information Processing Society of Japan, Dec. 2001, 42(12):3201-3212 (English abstract only). [cited by applicant]
Owens et al., “Visually Indicated Sounds,” 2016 IEEE Conference on Computer Vision and Pattern Recognition (CVPR), Jun. 2016, pp. 2405-2413. [cited by applicant]