IP Library › Granted Patent US 12,190,851
Granted Patent B2
US 12,190,851 · App. 17/865,869 · Granted Jan 7, 2025

Audio generation methods and systems

Inventor: Adrian Barahona Rios (London, GB)
Assignee: Sony Interactive Entertainment Europe Limited
G10H1/0008A63F13/54G10H2220/135G10H2250/235G10H2250/311
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,190,851
App. No.
17/865,869
Granted
Jan 7, 2025
Kind
B2
Abstract

A method of generating audio assets, comprising the steps of: receiving an input audio asset having a first duration, generating an input image representative of the input audio asset, training a generative model on the input image and implementing the trained generative model to generate an output image representative of an output audio asset having a second duration different to the first duration, and generating the output audio asset based on the output image.

Claims (32)

1. A method of generating audio assets, comprising the steps of:

receiving an input audio asset having a first duration,

generating an input image representative of the input audio asset,

training a generative model on the input image and implementing the trained generative model to generate an output image representative of an output audio asset having a second duration different to the first duration, and

generating the output audio asset based on the output image,

wherein the input image and output image each comprise an axis representative of time duration, and the step of generating an output image comprises retargeting the input image along the axis representative of time duration, and

wherein the output image has a larger dimension along the axis representative of time duration than the input image.

2. A method according to claim 1 , wherein the output image has a dimension along the axis representative of time duration which is multiplied by a scale factor.

3. A method according to claim 1 , wherein the input image is an input spectrogram, and the output image is an output spectrogram.

4. The method according to claim 3 , wherein the step of generating an input image comprises performing a Fourier transform on the input audio asset, and the step of generating the output audio asset comprises performing an inverse Fourier transform on the output image.

5. The method according to claim 1 , wherein the generative model is a single-image generative model comprising a generative adversarial network, GAN, having a generator and a patch discriminator.

6. A method according to claim 1 , wherein the step of receiving an input audio asset comprises receiving, from a video game environment, video game information, and the step of generating the output image comprises feeding the video game information into the generative model such that the output image is influenced by the video game information.

7. A method according to claim 1 , further comprising the step of storing the trained generative model on a memory, configured to be accessed to generate further audio assets.

8. A method according to claim 1 , wherein the step of receiving an input audio asset comprises receiving a second input audio asset, and the step of generating an input image comprises generating a multi-channel image, comprising a first channel having an image representative of the input audio asset, and a second channel having an image representative of the second input audio asset.

9. A non-transitory computer-readable medium having stored thereon a computer program comprising computer-implemented instructions that, when run on a computer, cause the computer to implement a method of generating audio assets, comprising the steps of:

receiving an input audio asset having a first duration,

generating an input image representative of the input audio asset,

training a generative model on the input image and implementing the trained generative model to generate an output image representative of an output audio asset having a second duration different to the first duration, and

generating the output audio asset based on the output image,

wherein the input image and output image each comprise an axis representative of time duration, and the step of generating an output image comprises retargeting the input image along the axis representative of time duration, and

wherein the output image has a larger dimension along the axis representative of time duration than the input image.

10. A system for generating audio assets, the system comprising:

an asset input unit configured to receive an input audio asset having a first duration, and to convert the input audio asset into an input image,

an image generation unit configured to implement a generative model to generate one or more output images based on the input image, the output image representing an output audio asset having a second duration different to the first duration, and

an asset output unit configured to generate an output audio asset based on the output image,

wherein the input image and output image each comprise an axis representative of time duration, and the step of generating an output image comprises retargeting the input image along the axis representative of time duration, and

wherein the output image has a larger dimension along the axis representative of time duration than the input image.

11. A system according to claim 10 , further comprising a transform unit configured to perform Fourier transform operations and inverse Fourier transform operations to convert between audio and graphical files, and wherein

the asset input unit is configured to access the transform unit to convert the input audio asset into an input image, and

the asset output unit is configured to access the transform unit to convert each output image into an output audio asset.

12. A system according to claim 10 , further comprising a video game data processing unit, configured to process video game information derived from or relating to a video game environment and feed through to one or more of the asset input unit, the image generation unit and the asset output unit, and the image generation unit is configured to implement the generative model based at least in part on the video game information.

13. A system according to claim 10 , configured to store the generative model on the memory.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jul 28, 2022
From: RIOS, ADRIAN BARAHONA
To: SONY INTERACTIVE ENTERTAINMENT EUROPE LIMITED
Reel/Frame 060660/0183 →
Priority Claims (1)
GB 2110280 · Jul 16, 2021 · national
Continuity (1)
Related Publication 20230018661A1 · Jan 19, 2023
References Cited (12)
US 10511908B1 · Fisher · 2019 [cited by applicant]
US 20190392802A1 · Higurashi · 2019 [cited by examiner]
US 20220319534A1 · Krishnan Gorumkonda · 2022 [cited by examiner]
JP 2019139102A · 2019 [cited by applicant]
KR 20200132352A · 2020 [cited by applicant]
Extended European Search Report including Written Opinion for Application No. 22183615.8 dated Dec. 9, 2022, pp. 1-12. [cited by applicant]
Huzaifah, M. et al., “Applying Visual Domain Style Transfer and Texture Synthesis Techniques to Audio—Insights and Challenges”, arxiv.org, Cornell University Library, Jan. 29, 2019 (Jan. 29, 2019), pp. 1-15. XP081009556. [cited by applicant]
Hiwang, Y. et al., “Mel-spectrogram augmentation for sequence to sequence voice conversion”, arxiv.org, Cornell University Library, Jan. 6, 2020 (Jan. 6, 2020), pp. 1-5. XP081572541. [cited by applicant]
Vainer, J. et al., “SpeedySpeech: Efficient Neural Speech Synthesis”, arxiv.org, Cornell University Library, Aug. 9, 2020 (Aug. 9, 2020), pp. 1-5. XP081737392. [cited by applicant]
Zhang, T. et al., “Learning long-term filter banks for audio source separation and audio scene classification”, EURASIP Journal on Audio, Speech, and Music Processing, May 30, 2018 (May 30, 2018), pp. 1-13, vol. 2018. X… [cited by applicant]
Examination Report for Appln. No. GB2110280.1 dated Sep. 11, 2023, pp. 1-4. [cited by applicant]
Combined Search and Examination Report for GB Application No. GB2110280.1 mailed Jan. 17, 2022, pp. 1-7. [cited by applicant]