IP Library › Granted Patent US 12,749,481
Granted Patent B2
US 12,749,481 · App. 18/663,899 · Granted Sep 29, 2026

Generating audio using auto-regressive generative neural networks

Inventors: Neil Zeghidour (Paris, FR); David Grangier (Mountain View, CA); Marco Tagliasacchi (Ruvigliana, CH); Raphaël Marinier (Paris, FR); Olivier Teboul (Paris, FR); Zalán Borsos (Zurich, CH)
Assignee: Google LLC
G10L15/16G06N3/0455G06N3/0475G10H1/0008G10L15/063G10L15/1815G10H2210/056G10H2250/311
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,749,481
App. No.
18/663,899
Granted
Sep 29, 2026
Kind
B2
Abstract

Methods, systems, and apparatus, including computer programs encoded on computer storage media, for generating a prediction of an audio signal. One of the methods includes receiving a request to generate an audio signal; obtaining a semantic representation of the audio signal; generating, using one or more generative neural networks and conditioned on at least the semantic representation, an acoustic representation of the audio signal; and processing at least the acoustic representation using a decoder neural network to generate the prediction of the audio signal.

Claims (59)

1 . A computer-implemented method for generating an acoustic representation of an audio signal, the method comprising:

receiving a request to generate an audio signal having a respective audio sample at each of a plurality of output time steps spanning a time window;

obtaining a semantic representation of the audio signal that specifies a respective semantic token at each of a plurality of first time steps spanning the time window; and

generating, using one or more generative neural networks and conditioned on at least the semantic representation, the acoustic representation of the audio signal, the acoustic representation specifying a set of one or more respective acoustic tokens at each of a plurality of second time steps spanning the time window, wherein the set of one or more respective acoustic tokens at each of the plurality of second time steps comprises a plurality of acoustic tokens that collectively represent a prediction of an output of a residual vector quantization applied to an embedding that represents acoustic properties of the audio signal at the second time step,

the residual vector quantization encodes the embedding using a hierarchy of a plurality of vector quantizers that each generate a respective acoustic token from a corresponding vocabulary of acoustic tokens for the vector quantizer, wherein the hierarchy comprises one or more coarse vector quantizers and one or more fine vector quantizers, and

the set of acoustic tokens at each second time step comprises, for each vector quantizer, a respective acoustic token selected from the vocabulary for the vector quantizer.

2 . The method of claim 1 , wherein each semantic token is selected from a vocabulary of semantic tokens and represents semantic content of the audio signal at the corresponding first time step.

3 . The method of claim 1 , wherein the one or more respective acoustic tokens at each second time step represent acoustic properties of the audio signal at the corresponding second time step.

4 . The method of claim 1 , further comprising:

processing at least the acoustic representation using a decoder neural network to generate a prediction of the audio signal.

5 . The method of claim 4 , wherein the decoder neural network is a decoder neural network of a neural audio codec that has been trained jointly with an encoder neural network on an objective that measures reconstruction quality of predicted audio signals generated by the decoder neural network from acoustic representations generated using outputs generated by the encoder neural network.

6 . The method of claim 1 , wherein the acoustic representation is a prediction of a ground truth acoustic representation that would be generated from outputs of an encoder neural network by processing the audio signal.

7 . The method of claim 6 , wherein the encoder neural network outputs a respective embedding at each of the plurality of second time steps, and wherein the ground truth acoustic representation is generated by applying quantization to each of the respective embeddings.

8 . The method of claim 7 , wherein:

the quantization is residual vector quantization that encodes each respective embedding using the hierarchy of the plurality of vector quantizers, and

the set of one or more respective acoustic tokens at each second time step comprise, for each vector quantizer, a respective acoustic token that is a prediction of a ground truth acoustic token that would be generated by the vector quantizer from a ground truth embedding generated by the encoder neural network at the second time step.

9 . The method of claim 1 , wherein the hierarchy comprises the one or more coarse vector quantizers at one or more first positions in the hierarchy and the one or more fine vector quantizers at one or more last positions in the hierarchy.

10 . The method of claim 1 , wherein generating, using one or more generative neural networks and conditioned on at least the semantic representation, an acoustic representation of the audio signal comprises:

generating, using a first generative neural network and for each of the one or more coarse vector quantizers in the hierarchy, the respective acoustic tokens for the second time steps for the coarse vector quantizer conditioned on at least the semantic representation.

11 . The method of claim 10 , wherein the first generative neural network is an auto-regressive neural network that is configured to generate the acoustic tokens auto-regressively according to a first generation order, and wherein each particular acoustic token for each particular coarse vector quantizer and at each particular second time step is conditioned on at least the semantic representation and any acoustic tokens that precede the particular acoustic token in the first generation order.

12 . The method of claim 11 , wherein each particular acoustic token for each particular coarse vector quantizer and at each particular second time step is preceded in the first generation order by (i) any acoustic token for any of the coarse vector quantizers at any second time step that precedes the particular second time step and (ii) any acoustic tokens at the particular second time step for any coarse vector quantizers that precede the particular coarse vector quantizer in the hierarchy.

13 . The method of claim 10 , wherein the first generative neural network has a decoder-only Transformer architecture or an encoder-decoder Transformer architecture.

14 . The method of claim 10 , wherein generating, using one or more generative neural networks and conditioned on at least the semantic representation, an acoustic representation of the audio signal comprises:

generating, using a second generative neural network and for each of the one or more fine vector quantizers in the hierarchy, the respective acoustic tokens for the second time steps for the fine vector quantizer conditioned on the respective acoustic tokens for the second time steps for the one or more coarse vector quantizers in the hierarchy.

15 . The method of claim 14 , wherein the second generative neural network is not conditioned on the semantic representation.

16 . The method of claim 14 , wherein the second generative neural network is an auto-regressive neural network that is configured to generate the acoustic tokens auto-regressively according to a second generation order, and wherein each particular acoustic token for each particular fine vector quantizer and at each particular second time step is conditioned on (i) the respective acoustic tokens for at least a subset of the second time steps for the one or more coarse vector quantizers and (ii) at least a subset of the acoustic tokens that precede the particular acoustic token in the second generation order.

17 . The method of claim 16 , wherein each particular acoustic token for each particular fine vector quantizer and at each particular second time step is preceded in the second generation order by (i) any acoustic token for any of the fine vector quantizers at any second time step that precedes the particular second time step and (ii) any acoustic tokens at the particular second time step for any fine vector quantizers that precede the particular fine vector quantizer in the hierarchy.

18 . The method of claim 16 , wherein each particular acoustic token for each particular fine vector quantizer and at each particular second time step is conditioned on (i) the respective acoustic tokens for the one or more coarse vector quantizers that are at most a threshold number of second time steps before the second time step and (ii) any acoustic tokens that precede the particular second time step in the second generation order and that are at second time steps that are at most a threshold number of second time steps before the second time step.

19 . The method of claim 14 , wherein the second generative neural network has a decoder-only Transformer architecture or an encoder-decoder Transformer architecture.

20 . The method of claim 1 , wherein obtaining a semantic representation of the audio signal comprises:

generating the semantic representation auto-regressively using a third generative neural network.

21 . The method of claim 1 , wherein the request specifies a context for the audio signal and the audio signal is conditioned on the context.

22 . The method of claim 21 , wherein the context specifies semantic properties of the audio signal and wherein obtaining a semantic representation of the audio signal comprises:

generating the semantic representation conditioned on the context.

23 . The method of claim 21 , wherein the context specifies acoustic properties of the audio signal, and wherein generating, using one or more generative neural networks and conditioned on at least the semantic representation, the acoustic representation of the audio signal comprises:

generating, using one or more generative neural networks and conditioned on the semantic representation and the context, the acoustic representation of the audio signal.

24 . The method of claim 23 , wherein the set of one or more respective acoustic tokens at each of the plurality of second time steps comprises a plurality of acoustic tokens that collectively represent a prediction of an output of a residual vector quantization applied to an embedding that represents acoustic properties of the audio signal at the second time step,

the residual vector quantization encodes the embedding using a hierarchy of a plurality of vector quantizers that each generate a respective acoustic token from a corresponding vocabulary of acoustic tokens for the vector quantizer, wherein the hierarchy comprises one or more coarse vector quantizers at one or more first positions in the hierarchy and one or more fine vector quantizers at one or more last positions in the hierarchy, and

the set of acoustic tokens at each second time step comprises, for each vector quantizer, a respective acoustic token selected from the vocabulary for the vector quantizer, and wherein generating, using one or more generative neural networks and conditioned on the semantic representation and the context, an acoustic representation of the audio signal comprises:

generating, using a first generative neural network and for each of the one or more coarse vector quantizers in the hierarchy, the respective acoustic tokens for the second time steps for the coarse vector quantizer conditioned on the semantic representation and the context.

25 . The method of claim 23 , wherein processing at least the acoustic representation using a decoder neural network to generate the prediction of the audio signal comprises:

processing the acoustic representation and an acoustic representation of the context using the decoder neural network to generate the prediction of the audio signal.

26 . The method of claim 21 , wherein the context comprises an audio input.

27 . The method of claim 21 , wherein the context comprises visual data.

28 . The method of claim 21 , wherein the context comprises text data.

29 . The method of claim 1 , wherein a number of first time steps and a number of second time steps that span the time window is less than the number of output time steps that span the time window.

30 . The method of claim 29 , wherein the number of first time steps that span the time window is less than the number of second time steps that span the time window.

31 . A system comprising one or more computers and one or more storage devices storing instructions that, when executed by the one or more computers, cause the one or more computers to perform operations for generating an acoustic representation of an audio signal, wherein the operations comprise:

receiving a request to generate an audio signal having a respective audio sample at each of a plurality of output time steps spanning a time window;

obtaining a semantic representation of the audio signal that specifies a respective semantic token at each of a plurality of first time steps spanning the time window; and

generating, using one or more generative neural networks and conditioned on at least the semantic representation, the acoustic representation of the audio signal, the acoustic representation specifying a set of one or more respective acoustic tokens at each of a plurality of second time steps spanning the time window, wherein the set of one or more respective acoustic tokens at each of the plurality of second time steps comprises a plurality of acoustic tokens that collectively represent a prediction of an output of a residual vector quantization applied to an embedding that represents acoustic properties of the audio signal at the second time step,

the residual vector quantization encodes the embedding using a hierarchy of a plurality of vector quantizers that each generate a respective acoustic token from a corresponding vocabulary of acoustic tokens for the vector quantizer, wherein the hierarchy comprises one or more coarse vector quantizers and one or more fine vector quantizers, and

the set of acoustic tokens at each second time step comprises, for each vector quantizer, a respective acoustic token selected from the vocabulary for the vector quantizer.

32 . One or more non-transitory computer-readable storage media storing instructions that when executed by one or more computers cause the one or more computers to perform operations for generating an acoustic representation of an audio signal, wherein the operations comprise:

receiving a request to generate an audio signal having a respective audio sample at each of a plurality of output time steps spanning a time window;

obtaining a semantic representation of the audio signal that specifies a respective semantic token at each of a plurality of first time steps spanning the time window; and

generating, using one or more generative neural networks and conditioned on at least the semantic representation, the acoustic representation of the audio signal, the acoustic representation specifying a set of one or more respective acoustic tokens at each of a plurality of second time steps spanning the time window, wherein the set of one or more respective acoustic tokens at each of the plurality of second time steps comprises a plurality of acoustic tokens that collectively represent a prediction of an output of a residual vector quantization applied to an embedding that represents acoustic properties of the audio signal at the second time step,

the residual vector quantization encodes the embedding using a hierarchy of a plurality of vector quantizers that each generate a respective acoustic token from a corresponding vocabulary of acoustic tokens for the vector quantizer, wherein the hierarchy comprises one or more coarse vector quantizers and one or more fine vector quantizers, and

the set of acoustic tokens at each second time step comprises, for each vector quantizer, a respective acoustic token selected from the vocabulary for the vector quantizer.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jul 11, 2024
From: ZEGHIDOUR, NEIL; GRANGIER, DAVID; TAGLIASACCHI, MARCO; MARINIER, RAPHAËL; TEBOUL, OLIVIER; BORSOS, ZALÁN
To: GOOGLE LLC
Reel/Frame 067970/0208 →
Continuity (4)
Continuation 18463092 · Sep 7, 2023
Provisional Application 63441412 · Jan 26, 2023
Provisional Application 63404528 · Sep 7, 2022
Related Publication 20240371366A1 · Nov 7, 2024
References Cited (199)
US 9190054B1 · Riley et al. · 2015 [cited by applicant]
US 10971170B2 · Wu · 2021 [cited by examiner]
US 11024288B2 · McCallum · 2021 [cited by examiner]
US 11410684B1 · Klimkov et al. · 2022 [cited by applicant]
US 11551663B1 · Bissell et al. · 2023 [cited by applicant]
US 11580145B1 · Kumar et al. · 2023 [cited by applicant]
US 11645947B1 · Blair et al. · 2023 [cited by applicant]
US 11691650B2 · Li · 2023 [cited by examiner]
US 11934935B2 · van den Oord · 2024 [cited by examiner]
US 12020138B2 · Zeghidour · 2024 [cited by examiner]
US 12062380B2 · Kleijn · 2024 [cited by examiner]
US 20060173686A1 · Hwang · 2006 [cited by applicant]
US 20130339035A1 · Chordia et al. · 2013 [cited by applicant]
US 20150120308A1 · Leistikow et al. · 2015 [cited by applicant]
US 20150170640A1 · Sak · 2015 [cited by examiner]
US 20160351188A1 · Rao · 2016 [cited by examiner]
US 20160372118A1 · Senior et al. · 2016 [cited by applicant]
US 20170103752A1 · Senior et al. · 2017 [cited by applicant]
US 20180025721A1 · Li et al. · 2018 [cited by applicant]
US 20180174576A1 · Soltau · 2018 [cited by examiner]
US 20180358005A1 · Tomar · 2018 [cited by examiner]
US 20180365554A1 · van den Oord · 2018 [cited by examiner]
US 20190057683A1 · Sak et al. · 2019 [cited by applicant]
US 20190103093A1 · Huang et al. · 2019 [cited by applicant]
US 20190115013A1 · Bengio · 2019 [cited by examiner]
US 20190378498A1 · Sainath et al. · 2019 [cited by applicant]
US 20200051583A1 · Wu · 2020 [cited by examiner]
US 20200075019A1 · Steelberg et al. · 2020 [cited by applicant]
US 20200090641A1 · Kim · 2020 [cited by examiner]
US 20200110803A1 · Djalali et al. · 2020 [cited by applicant]
US 20200126539A1 · van den Oord · 2020 [cited by examiner]
US 20200176004A1 · Kleijn · 2020 [cited by examiner]
US 20200349965A1 · Nesta · 2020 [cited by examiner]
US 20200410997A1 · Bar-On et al. · 2020 [cited by applicant]
US 20210065712A1 · Holm · 2021 [cited by applicant]
US 20210074266A1 · Lu et al. · 2021 [cited by applicant]
US 20210089909A1 · Binkowski · 2021 [cited by examiner]
US 20210098098A1 · Pinto · 2021 [cited by applicant]
US 20210117624A1 · Aghajanyan et al. · 2021 [cited by applicant]
US 20210217404A1 · Jia · 2021 [cited by examiner]
US 20210280165A1 · Yu et al. · 2021 [cited by applicant]
US 20210312923A1 · Gaur · 2021 [cited by examiner]
US 20210342670A1 · van den Oord · 2021 [cited by examiner]
US 20210383789A1 · Donahue et al. · 2021 [cited by applicant]
US 20220028367A1 · Shekhar · 2022 [cited by examiner]
US 20220157316A1 · Rebryk · 2022 [cited by examiner]
US 20220310073A1 · Audhkhasi · 2022 [cited by examiner]
US 20220343903A1 · Mostafazadeh · 2022 [cited by examiner]
US 20230013370A1 · Li et al. · 2023 [cited by applicant]
US 20230014624A1 · Rios · 2023 [cited by examiner]
US 20230017503A1 · Moritz et al. · 2023 [cited by applicant]
US 20230020621A1 · Rios · 2023 [cited by applicant]
US 20230058447A1 · Rosenberg et al. · 2023 [cited by applicant]
US 20230075891A1 · Zheng et al. · 2023 [cited by applicant]
US 20230096805A1 · Kim · 2023 [cited by examiner]
US 20230124296A1 · Ngo · 2023 [cited by examiner]
US 20230134235A1 · Setlur et al. · 2023 [cited by applicant]
US 20230215420A1 · Yu · 2023 [cited by examiner]
US 20230269291A1 · Ramadas et al. · 2023 [cited by applicant]
US 20230281427A1 · Sakhinana · 2023 [cited by examiner]
US 20230343319A1 · Hu et al. · 2023 [cited by applicant]
US 20230368804A1 · Kleijn · 2023 [cited by examiner]
US 20230376833A1 · Tanski et al. · 2023 [cited by applicant]
US 20240371366A1 · Zeghidour · 2024 [cited by examiner]
CN 108885870A · 2018 [cited by applicant]
CN 109937446A · 2019 [cited by applicant]
CN 110476206A · 2019 [cited by applicant]
CN 113724683A · 2021 [cited by applicant]
CN 114503191A · 2022 [cited by applicant]
CN 114664282A · 2022 [cited by applicant]
CN 114822492A · 2022 [cited by applicant]
GB 2508411A · 2014 [cited by applicant]
JP 2007148039A · 2007 [cited by applicant]
JP 2018141915A · 2018 [cited by applicant]
WO WO2022035586A1 · 2022 [cited by applicant]
International Preliminary Report on Patentability in International Appln. No. PCT/US2023/032168, mailed on Mar. 20, 2025, 17 pages. [cited by applicant]
Jian-Tao et al., “FreeVoiceCAD—A Multimodal User Interface Prototype System,” Journal of Computer Research and Development, Sep. 2003, 40(9):1382-1388 (English Abstract). [cited by applicant]
Notice of Allowance in Korean Appln. No. 10-2024-7017365, mailed on Apr. 3, 2025, 5 pages (with English translation). [cited by applicant]
Office Action in Chinese Appln. 202410903951.3, mailed on Feb. 22, 2025, 10 pages (with machine translation). [cited by applicant]
Extended European Search Report in European Appln. No. 25165951.2, mailed on May 30, 2025, 6 pages. [cited by applicant]
Office Action in Australian Appln. No. 2024219995, mailed on Sep. 18, 2025, 12 pages. [cited by applicant]
Office Action in Chinese Appln. No. 202380015080.5, mailed on Aug. 21, 2025, 9 pages. [cited by applicant]
Qui et al., “End-to-end Speech Synthesis Based on WaveNet,” Journal of Computer Applications, May 2019, 39(5):1325-1329. [cited by applicant]
Office Action in Japanese Appln. No. 2024-531121, mailed on Dec. 23, 2024, 11 pages (with machine translation). [cited by applicant]
Abu-El-Haija et al., “YouTube-8m: A large-scale video classification benchmark,” CoRR, Submitted on Sep. 27, 2016, arXiv:1609.08675v1, 10 pages. [cited by applicant]
Agustsson et al., “Soft-To-Hard Vector Quantization for End-to-End Learning Compressible Representations,” Paper, Presented at NIPS'17: Proceedings of the 31st International Conference on Neural Information Processing S… [cited by applicant]
Baevski et al., “vq-wav2vec: Self-supervised learning of discrete speech representations,” CoRR, Submitted on Oct. 12, 2019, arXiv:1910.05453v1, 11 pages. [cited by applicant]
Baevski et al., “wav2vec 2.0: A Framework for Self-Supervised Learning of Speech Representations,” CoRR, Submitted on Jun. 20, 2020, arXiv:2006.11477v1, 12 pages. [cited by applicant]
Bińkowski et al., “High Fidelity Speech Synthesis with Adversarial Networks,” CoRR, Submitted on Sep. 25, 2019, arXiv:1909.11646v1, 15 pages. [cited by applicant]
Borsos et al., “A Language Modeling Approach to Audio Generation,” CoRR, Submitted on Jun. 21, 2023, arXiv:2209.03143v2, 11 pages. [cited by applicant]
Brown et al., “Language Models Are Few-Shot Learners,” CoRR, Submitted on May 28, 2020, arXiv:2005.14165v1, 25 pages. [cited by applicant]
Carlini et al., “Extracting Training Data From Large Language Models,” Paper, Presented at the 30th USENIX Security Symposium, Virtual Conference, Aug. 11-13, 2021, 19 pages. [cited by applicant]
Carlini et al., “Quantifying Memorization Across Neural Language Models,” CoRR, Submitted on Feb. 15, 2022, arXiv:2202.07646v1, 19 pages. [cited by applicant]
Casanova et al., “YourTTS: Towards Zero-Shot Multi-Speaker TTS and Zero-Shot Voice Conversion for Everyone,” Paper, Presented at the 39th International Conference on Machine Learning, Baltimore, Maryland, Jul. 17-23, 20… [cited by applicant]
Chen et al., “Evaluating Large Language Models Trained on Code,” CoRR, Submitted on Jul. 7, 2021, arXiv:2107.03374v1, 35 pages. [cited by applicant]
Chen et al., “Fine-grained style control in transformer-based text-to-speech synthesis,” Paper, Presented at IEEE International Conference on acoustics, speech and signal processing, Singapore, Singapore, May 22-27, 202… [cited by applicant]
Chen et al., “Wavegrad: Estimating Gradients for Waveform Generation,” CoRR, Submitted on Sep. 2, 2020, arXiv:2009.00713v1, 15 pages. [cited by applicant]
Chinen et al., “ViSQOL v3: An Open Source Production Ready Objective Speech and Audio Metric,” Paper, Presented at the 12th International Conference on Quality of Multimedia Experience, Athlone, Ireland, May 26-28, 2020… [cited by applicant]
Choromanski et al., “Rethinking Attention with Performers,” CoRR, Submitted on Sep. 30, 2023, arXiv:2009.14794v1, 38 pages. [cited by applicant]
Chowdhery et al., “PaLM: Scaling Language Modeling with Pathways,” CoRR, Submitted on Apr. 5, 2022, arXiv:2204.02311v1, 87 pages. [cited by applicant]
Chung et al., “W2v-bert: Combining Contrastive Learning and Masked Language Modeling for Self-Supervised Speech Pre-Training,” CoRR, Submitted on Sep. 13, 2021, arXiv:2108.06209v2, 7 pages. [cited by applicant]
Conneau et al., “Xtreme-s: Evaluating Cross-Lingual Speech Representations,” CoRR, Submitted on Mar. 21, 2022, arXiv:2203.10752v1, 13 pages. [cited by applicant]
Cuturi, “Sinkhorn Distances: Lightspeed Computation of Optimal Transport,” CoRR, Submitted on Jun. 3, 2013, arXiv:1306.0895v1, 9 pages. [cited by applicant]
Defferrard et al., “FMA: A Dataset for Music Analysis,” CoRR, Submitted on Dec. 6, 2016, arXiv:1612.01840v1, 8 pages. [cited by applicant]
Défossez et al., “High Fidelity Neural Audio Compression,” CoRR, Submitted on Oct. 24, 2022, arXiv:2210.13438v1, 19 pages. [cited by applicant]
Delgado et al., “ASVspoof 2021: Automatic Speaker Verification Spoofing and Countermeasures Challenge Evaluation Plan,” CoRR, Submitted on Sep. 1, 2021, arXiv:2109.00535v1, 13 pages. [cited by applicant]
Devlin et al., “Bert: Pre-Training of Deep Bidirectional Transformers for Language Understanding,” CoRR, Submitted on Oct. 11, 2018, arXiv:1810.04805v1, 16 pages. [cited by applicant]
Dhariwal et al., “Jukebox: A Generative Model for Music,” CoRR, Submitted on Apr. 30, 2020, arXiv:2005.00341v1, 20 pages. [cited by applicant]
Dosovitskiy et al., “An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale,” CoRR, Submitted on Oct. 22, 2020, arXiv:2010.11929v1, 22 pages. [cited by applicant]
Duan et al., “Interpretable Melody Generation from Lyrics with Discrete-Valued Adversarial Training,” CoRR, Submitted on Jun. 30, 2022, arXiv:2006.15027v1, 3 pages. [cited by applicant]
Dunbar et al., “The Zero Resource Speech Challenge 2021: Spoken Language Modelling,” CoRR, Submitted on Apr. 29, 2021, arXiv:2104.14700v1, 5 pages. [cited by applicant]
Engel et al., “DDSP: Differentiable Digital Signal Processing,” CoRR, Submitted on Jan. 14, 2020, arXiv:2001.04643v1, 19 pages. [cited by applicant]
Esser et al., “Taming Transformers for High-Resolution Image Synthesis,” Paper, Presented at the Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, Nashville, Tennessee, Jun. 20-25, 2021,… [cited by applicant]
Gao et al., “Interactive text-to-speech system via joint style analysis,” Paper, Proceedings of the annual conference of the international speech communication association, Shanghai, China, Oct. 25-29, 2020, pp. 4447-44… [cited by applicant]
Ge et al., “Long Video Generation with Time-Agnostic VQGAN and Time-Sensitive Transformer,” CoRR, Submitted on Sep. 24, 2022, arXiv:2204.03638v4, 30 pages. [cited by applicant]
Gemmeke et al., “Audio set: An Ontology and Human-Labeled Dataset for Audio Events,” Paper, Presented at the 2017 IEEE International Conference on Acoustics, Speech and Signal Processing, New Orleans, Louisiana, Mar. 5-… [cited by applicant]
Github.com [online], “MubertAI/Mubert-Text-To-Music,” May 25, 2023, retrieved on Oct. 3, 2023, retrieved from URL<https://github.com/MubertAI/Mubert-Text-to-Music>, 1 page. [cited by applicant]
Gong et al., “Ast: Audio spectrogram transformer,” CoRR, Submitted on Jul. 8, 2021, arXiv:2104.01778v3, 5 pages. [cited by applicant]
Goodfellow et al., “Generative Adversarial Networks,” CoRR, Submitted on Jun. 10, 2014, arXiv:1406.2661v1, 9 pages. [cited by applicant]
Gulati et al., “Conformer: Convolution-Augmented Transformer for Speech Recognition,” CoRR, Submitted on May 16, 2020, arXiv:2005.08100v1, 5 pages. [cited by applicant]
Hawthorne et al., “Enabling Factorized Piano Music Modeling and Generation with the MAESTRO Dataset,” CoRR, Submitted on Oct. 29, 2018, arXiv:1810.12247v1, 12 pages. [cited by applicant]
Hawthorne et al., “General-Purpose, Long-Context Autoregressive Modeling with Perceiver AR,” Paper, Presented at the 39th International Conference on Machine Learning, Baltimore, Maryland, Jul. 17-23, 2022, 24 pages. [cited by applicant]
Hawthorne et al., “Multi-Instrument Music Synthesis with Spectrogram Diffusion,” CoRR, Submitted on Jun. 11, 2022, arXiv:2206.05408v1, 12 pages. [cited by applicant]
Hershey et al., “CNN Architectures for Large-Scale Audio Classification,” Paper, Presented at the 2017 IEEE International Conference on Acoustics, Speech and Signal Processing, New Orleans, Louisiana, Mar. 5-9, 2017, 5 … [cited by applicant]
Hines et al., “ViSQOL: An Objective Speech Quality Model,” EURASIP Journal on Audio, Speech, and Music Processing, Dec. 2015, 13:1-18. [cited by applicant]
Ho et al., “Denoising Diffusion Probabilistic Models,” CoRR, Submitted on Jun. 19, 2020, arXiv:2006.11239v1, 12 pages. [cited by applicant]
Ho et al., “Video Diffusion Models,” CoRR, Submitted on Apr. 7, 2022, arXiv:2204.03458v1, 14 pages. [cited by applicant]
Hong et al., “Cogvideo: Large-Scale Pretraining for Text-to-Video Generation via Transformers,” CoRR, Submitted on May 29, 2022, arXiv:2205.15868v1, 15 pages. [cited by applicant]
Hsu et al., “HuBERT: How Much Can a Bad Teacher Benefit ASR Pre-Training?,” Paper, Presented at 2021 IEEE International Conference on Acoustics, Speech and Signal Processing , Toronto, Canada, Jun. 6-11, 2021, 10 pages. [cited by applicant]
Hsu et al., “HuBERT: Self-Supervised Speech Representation Learning by Masked Prediction of Hidden Units,” IEEE/ACM Transactions on Audio, Speech, and Language Processing, Oct. 26, 2021, 29:3451-3460. [cited by applicant]
Huang et al., “Mulan: A Joint Embedding of Music Audio and Natural Language,” CoRR, Submitted on Aug. 26, 2022, arXiv:2208.12415v1, 8 pages. [cited by applicant]
Huang et al., “Music transformer,” CoRR, Submitted on Sep. 12, 2018, arXiv:1809.04281v1, 14 pages. [cited by applicant]
International Search Report and Written Opinion in International Appln. No. PCT/US2023/032168, mailed on Apr. 18, 2024, 27 pages. [cited by applicant]
Kahn et al., “Libri-Light: A Benchmark for ASR with Limited or No Supervision,” Paper, Presented at 2020 IEEE International Conference on Acoustics, Speech and Signal Processing, Virtual Conference, May 4-8, 2020, 7 pag… [cited by applicant]
Kalchbrenner et al., “Efficient Neural Audio Synthesis,” CoRR, Submitted on Jun. 25, 2018, arXiv:1802.08435v2, 10 pages. [cited by applicant]
Kankanahalli, “End-to-End Optimized Speech Coding with Deep Neural Networks,” Paper, Presented at the 2018 IEEE International Conference on Acoustics, Speech and Signal Processing, Calgary, Canada, Apr. 15-20, 2018, 5 p… [cited by applicant]
Kharitonov et al., “Data Augmenting Contrastive Learning of Speech Representations in the Time Domain,” Paper, Presented at the 2021 IEEE Spoken Language Technology Workshop, Virtual Conference, Jan. 19-22, 2021, 6 page… [cited by applicant]
Kharitonov et al., “Text-Free Prosody-Aware Generative Spoken Language Modeling,” CoRR, Submitted on Sep. 7, 2021, arXiv:2109.03264v1, 16 pages. [cited by applicant]
Kharitonov et al., “textless-lib: A Library for Textless Spoken Language Processing,” CoRR, Submitted on Feb. 15, 2022, arXiv:2202.07359v1, 9 pages. [cited by applicant]
Kilgour et al., “Fréchet Audio Distance: A Reference-Free Metric for Evaluating Music Enhancement Algorithms,” Paper, Presented at 20th Annual Conference of the International Speech Communication Association, Graz, Aust… [cited by applicant]
Kim et al., “Audiocaps: Generating Captions for Audios in the Wild,” Paper, Presented at the Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Langu… [cited by applicant]
Kong et al., “Diffwave: A Versatile Diffusion Model for Audio Synthesis, ” CoRR, Submitted on Sep. 21, 2020, arXiv:2009.09761v1, 17 pages. [cited by applicant]
Kong et al., “HiFi-GAN: Generative Adversarial Networks for Efficient and High Fidelity Speech Synthesis,” Advances in Neural Information Processing Systems, Dec. 2020, 33:17022-17033. [cited by applicant]
Kreuk et al., “Audiogen: Textually Guided Audio Generation,” CoRR, Submitted on Sep. 30, 2022, arXiv:2209.15352v1, 16 pages. [cited by applicant]
Kumar et al., “MelGAN: Generative Adversarial Networks for Conditional Waveform Synthesis,” Advances in neural information processing systems, Dec. 2019, 32:12 pages. [cited by applicant]
Lakhotia et al., “On Generative Spoken Language Modeling from Raw Audio,” Transactions of the Association for Computational Linguistics, Dec. 6, 2021, 9:1336-1354. [cited by applicant]
Lample et al., “Deep learning for symbolic mathematics,” CoRR, Submitted on Dec. 2, 2019, arXiv:1912.01412v1, 24 pages. [cited by applicant]
Latif et al., “Sparks of large audio models: a survey and outlook,” CoRR, Submitted on Sep. 3, 2023, arXiv:0916.03791v2, 34 pages. [cited by applicant]
Liu et al., “DelightfulTTS 2: End-to-End Speech Synthesis with Adversarial Vector-Quantized Auto-Encoders,” CoRR, Submitted on Jul. 11, 2022, arXiv:2007.04646v1, 5 pages. [cited by applicant]
Liu et al., “RoBERTa: A Robustly Optimized Bert Pretraining Approach,” CoRR, Submitted on Jul. 26, 2019, arXiv:1907.11692v1, 13 pages. [cited by applicant]
Nguyen et al., “Are Discrete Units Necessary for Spoken Language Modeling?,” IEEE Journal of Selected Topics in Signal Processing, Aug. 23, 2022, 16(6):1415-1423. [cited by applicant]
Nichol et al., “GLIDE: Towards Photorealistic Image Generation and Editing with Text-Guided Diffusion Models,” CoRR, Submitted on Dec. 20, 2021, arXiv:2112.10741v1, 20 pages. [cited by applicant]
Notice of Allowance in Australian Appln. No. 2023337867, mailed on Jun. 20, 2024, 3 pages. [cited by applicant]
Office Action in Korean Appln. No. 10-2024-7017365, mailed on Oct. 22, 2024, 16 pages (with machine translation). [cited by applicant]
Oord et al., “Neural Discrete Representation Learning,” CoRR, Submitted on Nov. 2, 2017, arXiv:1711.00937v1, 11 pages. [cited by applicant]
Oord et al., “Parallel WaveNet: Fast High-Fidelity Speech Synthesis,” Paper, Presented at the 35th International Conference on Machine Learning, Stockholm, Sweden, May 26-28, 2018, 9 pages. [cited by applicant]
Oord et al., “Representation Learning with Contrastive Predictive Coding, ” CoRR, Submitted on Jul. 10, 2018, arXiv:1807.03748v1, 13 pages. [cited by applicant]
Panayotov et al., “Librispeech: An ASR Corpus Based on Public Domain Audio Books,” Paper, Presented at 2015 IEEE International Conference on Acoustics, Speech and Signal Processing, Brisbane, Queensland, Apr. 19-24, 201… [cited by applicant]
Pasad et al., “Layer-Wise Analysis of a Self-Supervised Speech Representation Model,” Paper, Presented at the 2021 IEEE Automatic Speech Recognition and Understanding Workshop, Virtual Conference, Dec. 13-17, 2021, 9 pa… [cited by applicant]
Peng et al., “Self-Supervised Representation Learning for Speech Using Visual Grounding and Masked Language Modeling,” CoRR, Submitted on Feb. 7, 2022, arXiv:2202.03543v1, 9 pages. [cited by applicant]
Petermann et al., “Hyper-Autoencoded Reconstruction Propagation for Scalable Neural Audio Coding,” Paper, Presented at the 2021 IEEE Workshop on Applications of Signal Processing to Audio and Acoustics, Virtual Conferen… [cited by applicant]
Radford et al., “Learning Transferable Visual Models from Natural Language Supervision,” Paper, Presented at the International Conference on Machine Learning, Virtual Conference, Jul. 18-24, 2021, 16 pages. [cited by applicant]
Rae et al., “Scaling Language Models: Methods, Analysis & Insights from Training Gopher,” CoRR, Submitted on Dec. 8, 2021, arXiv:2112.11446v1, 120 pages. [cited by applicant]
Raffel et al., “Exploring the Limits of Transfer Learning with a Unified Text-to-Text Transformer,” The Journal of Machine Learning Research, Jun. 1, 2020, 21(1):67 pages. [cited by applicant]
Ramesh et al., “Hierarchical Text-Conditional Image Generation with Clip Latents,” CoRR, Submitted on Apr. 13, 2022, arXiv:2204.06125v1, 27 pages. [cited by applicant]
Ramesh et al., “Zero-Shot Text-to-Image Generation,” Paper, Presented at the International Conference on Machine Learning, Virtual Conference, Jul. 18-24, 2021, 27 pages. [cited by applicant]
Razavi et al., “Generating Diverse High-Fidelity Images with VQ-VAE-2,” Advances in neural information processing systems, Dec. 2019, 32:11 pages. [cited by applicant]
Riffusion.com [online], “[Riffusion] (noun): riff + diffusion,” available on or before Dec. 15, 2022, via Internet Archive: Wayback Machine URL<http://web.archive.org/web/20221215132646/https://www.riffusion.com/about>,… [cited by applicant]
Rivière et al., “Towards Unsupervised Learning of Speech Features in the Wild,” Paper, Presented at the 2021 IEEE Spoken Language Technology Workshop (SLT), Virtual Conference, Jan. 19-22, 2021, 9 pages. [cited by applicant]
Roberts et al., “Scaling Up Models and Data with T5X and Seqio,” CoRR, Submitted on Mar. 31, 2022, arXiv:2203.17189v1, 12 pages. [cited by applicant]
Rombach et al., “High-Resolution Image Synthesis with Latent Diffusion Models,” Paper, Presented at the Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, New Orleans, Louisiana, Jun. 18-… [cited by applicant]
Roy et al., “Efficient Content-Based Sparse Attention with Routing Transformers,” Transactions of the Association for Computational Linguistics, Feb. 1, 2021, 9:53-68. [cited by applicant]
Saeed et al., “Contrastive Learning of General-Purpose Audio Representations,” Paper, Presented at the 2021 IEEE International Conference on Acoustics, Speech and Signal Processing, Toronto, Ontario, Jun. 6-11, 2021, 5 … [cited by applicant]
Saharia et al., “Photorealistic Text-to-Image Diffusion Models with Deep Language Understanding,” Advances in Neural Information Processing Systems, Dec. 6, 2022, 35:36479-36494. [cited by applicant]
Schatz et al., “Evaluating Speech Features with the Minimal-Pair ABX task: Analysis of the classical MFC/PLP pipeline,” Paper, Presented at Interspeech 2013: 14th Annual Conference of the International Speech Communicat… [cited by applicant]
Schatz, “ABX-Discriminability Measures and Applications,” Cognitive Science, Universite Paris 6 (UMPC), 2016, 6 pages. [cited by applicant]
Schroff et al., “Facenet: A Unified Embedding for Face Recognition and Clustering,” Paper, Presented at the Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Boston, Massachusetts, Jun. 7-12… [cited by applicant]
Shor et al., “Towards Learning a Universal Non-Semantic Representation of Speech,” CoRR, Submitted on Feb. 25, 2020, arXiv:2002.12764v1, 5 pages. [cited by applicant]
Tagliasacchi et al., “SEANet: A Multi-Modal Speech Enhancement Network,” CoRR, Submitted on Sep. 4, 2020, arXiv:2009.02095v1, 5 pages. [cited by applicant]
Tagliasacchi et al., “Self-Supervised Audio Representation Learning for Mobile Devices,” CoRR, Submitted on May 24, 2019, arXiv:1905.11796v1, 11 pages. [cited by applicant]
Thoppilan et al., “LaMDA: Language Models for Dialog Applications,” CoRR, Submitted on Jan. 20, 2022, arXiv:2201.08239v1, 47 pages. [cited by applicant]
Valle et al., “Flowtron: an Autoregressive Flow-based Generative Network for Text-to-Speech Synthesis,” CoRR, Submitted on Jul. 16, 2020, arXiv:2005.05957v3, 10 pages. [cited by applicant]
Van den Oord et al., “WaveNet: A Generative Model for Raw Audio,” CoRR, Submitted on Sep. 12, 2016, arXiv:1609.03499v1, 15 pages. [cited by applicant]
Van Niekerk et al., “Analyzing Speaker Information in Self-Supervised Models to Improve Zero-Resource Speech Processing,” CoRR, Submitted on Aug. 2, 2021, arXiv:2108.00917v1, 5 pages. [cited by applicant]
Vaswani et al., “Attention is All You Need,” Advances in Neural Information Processing Systems, Dec. 2017, 30:11 pages. [cited by applicant]
Villegas et al., “Phenaki: Variable Length Video Generation from Open Domain Textual Description,” CoRR, Submitted on Oct. 5, 2022, arXiv:2210.02399v1, 17 pages. [cited by applicant]
Wu et al., “Godiva: Generating open-domain videos from natural descriptions,” CoRR, Submitted on Apr. 30, 2021, arXiv:2104.14806v1, 12 pages. [cited by applicant]
Wu et al., “NUWA: Visual Synthesis Pre-Training for Neural Visual World Creation,” Paper, Presented at the European Conference on Computer Vision, Tel Aviv, Israel, Oct. 23-27, 2022, 28 pages. [cited by applicant]
Wu et al., “Nuwa-infinity: Autoregressive over autoregressive generation for infinite visual synthesis,” CoRR, Submitted on Jul. 20, 2022, arXiv:2207.09814v1, 24 pages. [cited by applicant]
Wu et al., “Wav2clip: Learning Robust Audio Representations From Clip,” Paper, Presented at the 2022 IEEE International Conference on Acoustics, Speech and Signal Processing, Singapore, Singapore, May 23-27, 2022, 5 pag… [cited by applicant]
Yang et al., “Diffsound: Discrete Diffusion Model for Text-to-Sound Generation,” Journal of Latex Class Files, Aug. 2021, 31(8):1-13. [cited by applicant]
Yu et al., “Scaling Autoregressive Models for Content-Rich Text-to-Image Generation,” CoRR, Submitted on Jun. 22, 2022, arXiv:2206.10789v1, 49 pages. [cited by applicant]
Yu et al., “Vector-Quantized Image Modeling with Improved VQGAN,” CoRR, Submitted on Oct. 9, 2021, arXiv:2110.04627v1, 17 pages. [cited by applicant]
Zeghidour et al., “LEAF: A Learnable Frontend for Audio Classification,” CoRR, Submitted on Jan. 21, 2021, arXiv:2101.08596v1, 16 pages. [cited by applicant]
Zeghidour et al., “Soundstream: An End-to-End Neural Audio Codec, ” IEEE/ACM Transactions on Audio, Speech, and Language Processing, Nov. 23, 2021, 30:495-507. [cited by applicant]
Zen et al., “Statistical Parametric Speech Synthesis Using Deep Neural Networks,” Paper, Presented at 2013 IEEE International Conference on Acoustics, Speech and Signal Processing, Vancouver, British Columbia, May 26-31… [cited by applicant]
Zhen et al., “Cascaded Cross-Module Residual Learning Towards Lightweight End-to-End Speech Coding,” CoRR, Submitted on Jun. 18, 2019, arXiv:1906.07769v1, 5 pages. [cited by applicant]
Zhu et al., “Quantized GAN for Complex Music Generation from Dance Videos,” CoRR, Submitted on Jul. 19, 2022, arXiv:2004.00604v2, 18 pages. [cited by applicant]
Extended European Search Report in European Appln. No. 26174694.5, mailed on Jul. 16, 2026, 10 pages. [cited by applicant]