IP Library › Granted Patent US 12,364,924
Granted Patent B2
US 12,364,924 · App. 17/865,767 · Granted Jul 22, 2025

Audio generation methods and system

Inventor: Adrian Barahona Rios (London, GB)
Assignee: SONY INTERACTIVE ENTERTAINMENT EUROPE LIMITED
A63F13/54A63F13/60G06T11/00G10L21/14G10L21/18G10L25/18G10L25/30H04S3/008A63F2300/6009A63F2300/6072A63F2300/6081H04S2400/01
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,364,924
App. No.
17/865,767
Filed
Jul 15, 2022
Granted
Jul 22, 2025
Kind
B2
Art Unit
3715
USPC
463/35
Abstract

A method of generating audio assets, comprising the steps of: receiving an input multi-layered audio asset comprising a plurality of audio layers, generating an input multi-channel image, wherein each channel of the input multi-channel image comprises an input image representative of one of the audio layers, training a generative model on the input multi-channel image and implementing the trained generative model to generate an output multi-channel image, wherein each channel of the output multi-channel image comprises an output image representative of an output audio layer, and generating an output multi-layered audio asset based on a combination of output audio layers derived from the output images.

Claims (34)

1. A method of generating audio assets, comprising the steps of:

receiving an input multi-layered audio asset comprising a plurality of audio layers,

converting each audio layer into an input graphical representation,

generating an input multi-channel image by stacking each input graphical representation in a separate channel of the multi-channel image,

training a generative model on the input multi-channel image and implementing the trained generative model to generate an output multi-channel image, wherein each channel of the output multi-channel image comprises an output graphical representation of an output audio layer,

separating the output graphical representations from each output multi-channel image, and

generating an output multi-layered audio asset based on a combination of output audio layers derived from the output graphical representations.

2. The method according to claim 1 , wherein the step of generating an output multi-layered audio asset comprises arranging the plurality of output audio layers with a time delay between each of the output audio layers.

3. The method according to claim 2 , wherein the step of receiving an input multi-layered audio asset comprises determining the temporal arrangement of the audio layers, and

the time delay between each of the plurality of output audio layers is configured according to the determined temporal arrangement of the audio layers.

4. The method according to claim 1 , wherein the step of receiving an input multi-layered audio asset comprises labelling the sequence of audio layers in the input multi-layered audio asset, and

the step of generating an output multi-layered audio asset comprises arranging the output audio layers according to the labelled sequence of the input multi-layered audio asset.

5. The method according to claim 1 , wherein the generative model is a single-image generative model comprising a generative adversarial network, GAN, having a generator and a patch discriminator.

6. The method according to claim 1 , wherein each input image is a spectrogram of a respective audio layer, the input multi-channel image is an input multi-channel spectrogram and the output multi-channel image is an output multi-channel spectrogram.

7. The method according to claim 6 , wherein the step of generating an input multi-channel image comprises separating the audio layers in the input multi-channel image and performing a Fourier transform on each audio layer to generate a spectrogram of the respective audio layer.

8. A method according to claim 1 , wherein the step of receiving an input multi-layered audio asset comprises receiving, from a video game environment, video game information, and the step of generating the output multi-channel images comprises feeding the video game information into the single-image generative model such that the output multi-channel image is influenced by the video game information.

9. A method according to claim 1 , further comprising the step of storing the trained generative model on a memory, configured to be accessed to generate further audio assets.

10. A method according to claim 1 , wherein the step of receiving an input multi-layered audio asset comprises receiving a second input multi-layered audio asset having second input layers, and the input multi-channel image further comprises images representative of the second input layers.

11. A computer program comprising computer-implemented instructions that, when run on a computer, cause the computer to implement a method of generating audio assets, comprising the steps of:

receiving an input multi-layered audio asset comprising a plurality of audio layers,

converting each audio layer into an input graphical representation,

generating an input multi-channel image by stacking each input graphical representation in a separate channel of the multi-channel image,

training a generative model on the input multi-channel image and implementing the trained generative model to generate an output multi-channel image, wherein each channel of the output multi-channel image comprises an output graphical representation of an output audio layer,

separating the output graphical representations from each output multi-channel image, and

generating an output multi-layered audio asset based on a combination of output audio layers derived from the output graphical representation.

12. A system for generating audio assets, the system comprising:

an asset input unit configured to receive an input multi-layered audio asset comprising a plurality of audio layers, convert each audio layer into an input graphical representation, and generate an input multi-channel image by stacking each input graphical representation in a separate channel of the multi-channel image, and

an image generation unit configured to implement a generative model to generate one or more output multi-channel images based on the input multi-channel image, each channel of the output multi-channel image comprising an output graphical representation of an output audio layer, and

an asset output unit configured to separate the output graphical representations from each output multi-channel image and generate an output multi-layered audio asset based on a combination of output audio layers derived from the output graphical representations.

13. A system according to claim 12 , further comprising a transform unit configured to perform Fourier transform operations and inverse Fourier transform operations to convert between audio and graphical files, and wherein

the asset input unit is configured to access the transform unit to convert each input audio layer into an input graphical representation, and

the asset output unit is configured to access the transform unit to convert each output graphical representation into an output audio layer.

14. A system according to any of claim 12 , further comprising a video game data processing unit, configured to process video game information derived from or relating to a video game environment and feed through to one or more of the asset input unit, the image generation unit and the asset output unit, and the image generation unit is configured to implement the generative model based at least in part on the video game information.

15. A system according to claim 12 , configured to store the generative model on the memory.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jul 28, 2022
From: RIOS, ADRIAN BARAHONA
To: SONY INTERACTIVE ENTERTAINMENT EUROPE LIMITED
Reel/Frame 060660/0106 →
Priority Claims (1)
GB 2110275 · Jul 16, 2021 · national
Continuity (1)
Related Publication 20230020621A1 · Jan 19, 2023
References Cited (22)
US 10511908B1 · Fisher · 2019 [cited by applicant]
US 20070185705A1 · Hiroe · 2007 [cited by applicant]
US 20080041220A1 · Foust · 2008 [cited by examiner]
US 20100131086A1 · Itoyama et al. · 2010 [cited by applicant]
US 20190392802A1 · Higurashi · 2019 [cited by examiner]
US 20210019572A1 · Munoz Delgado · 2021 [cited by examiner]
US 20220193549A1 · Wakeland · 2022 [cited by examiner]
JP 2019139102A · 2019 [cited by applicant]
KR 20200132352A · 2020 [cited by applicant]
WO 2020084787A1 · 2020 [cited by applicant]
WO 2021013345A1 · 2021 [cited by applicant]
Examination Report for Application No. GB2110275.1 dated Sep. 11, 2023, 4 pages. [cited by applicant]
Combined Search and Examination Report for GB Application No. GB2110275.1 mailed Jan. 17, 2022. 8 pgs. [cited by applicant]
Nugraha. A. et al., “Multichannel Audio Source Separation With Deep Neural Networks” IEEE/ ACM Transactions on Audio, Speech, and Language Processing, Sep. 2016, pp. 1652-1664, vol. 24, Issue No. 9. [cited by applicant]
Ristea. N. et al., “Are you wearing a mask? Improving mask detection from speech using augmentation by cycle-consistent GANs” arXiv. Org, Cornell University Library, Jul. 2020, pp. 1-5. [cited by applicant]
Itakura. K. et al., “Bayesian Multichannel Audio Source Separation Based on Integrated Source and Spatial Models” IEEE/ACM Transactions on Audio, Speech, and Language Processing, Apr. 2018, pp. 831-846, vol. 26. Issue 4. [cited by applicant]
Choi. W. et al., Investigating U-Nets With Various Intermediate Blocks for Spectrogram-Based Singing Voice Separation arXiv. Org, Cornell University Library, Oct. 2020, pp. 1-8. [cited by applicant]
Ngo. D. et al., “Sound Context Classification Basing on Join Learning Model and Multi-Spectrogram Features” arXiv, Cornell University Library, May 2020, pp. 1-12. [cited by applicant]
Pham. L. et al., Robust Acoustic Scene Classification using a Multi-Spectogram Encoder-Decoder Framework IEEE Transactions on Audio, Speech and Language Processing, Feb. 2020, pp. 1-10. [cited by applicant]
Sriram. G. et al., “3-D CNN Models for Far-Field Multi-Channel Speech Recognition” IEEE International Conference on Acoustics, Speech and Signal Processing, Apr. 2018, pp. 5499-5503. [cited by applicant]
Jordan. S. et al., “Nonnegative Tensor Factorization for Source Separation of Loops in Audio” IEEE International Conference on Acoustics, Speech and Signal Processing, Apr. 2018, pp. 171-175. [cited by applicant]
Extended European Search Report including Written Opinion for Application No. 22183549.9 dated Dec. 9, 2022, pp. 1-14. [cited by applicant]