IP Library Granted Patent US 12706108
Granted Patent B2
US 12706108 · App. 18/104,071 · Granted Aug 11, 2026

Signal transformation based on unique key-based network guidance and conditioning

Inventors: Atti Venkatraman (Calabasas, CA); Zoran Fejzo (Calabasas, CA); Antonius Kalker (Calabasas, CA)
Assignee: DTS Inc.
G10L19/08G10L25/30
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12706108
App. No.
18/104,071
Granted
Aug 11, 2026
Kind
B2
Abstract

A method comprises receiving input audio and target audio having a target audio characteristic. The method includes estimating key parameters that represent the target audio characteristic based on one or more of the target audio and the input audio. The method further comprises configuring a neural network, trained to be configured by the key parameters, with the key parameters to cause the neural network to perform a signal transformation of the input audio, to produce output audio having an output audio characteristic corresponding to and that matches the target audio characteristic.

Claims (38)

1 . A method comprising:

receiving input audio and target audio having a target audio characteristic, wherein the input audio and the target audio are received as separate signals;

estimating frame-specific key parameters that include quantized representations of the target audio characteristic derived from a joint analysis of corresponding frames from the target audio and the input audio, wherein the frame-specific key parameters vary from frame to frame, wherein the quantized representations are based on an analysis including a temporal analysis and at least one of a spectral analysis or a frequency harmonic analysis;

encoding the input audio and the frame-specific key parameters;

multiplexing the encoded frame-specific key parameters in alignment with the encoded input audio at a frame level into a combined bit-stream for transmission to a neural network;

demultiplexing the combined bit-stream to generate decoded input audio and decoded frame-specific key parameters;

configuring the neural network, trained to be configured by the frame-specific key parameters, with the decoded frame-specific key parameters to cause the neural network to perform a signal transformation of the decoded input audio; and

producing output audio having an output audio characteristic corresponding to and that matches the target audio characteristic, wherein the output audio has perceptually-improved signal quality with frequency bandwidth extension as compared to the decoded input audio.

2 . The method of claim 1 , wherein the output audio includes a target temporal characteristic and a temporal characteristic, which are each a respective temporal amplitude characteristic.

3 . The method of claim 1 , wherein the estimating the frame-specific key parameters includes: estimating as temporal key parameters temporal key parameters that represent a temporal amplitude characteristic of the target audio, wherein the estimating further includes at least one of: spectral envelope key parameters including LP coefficients (LPCs) or line spectral frequencies (LSFs) representative of a target spectral envelope of the target audio; and harmonic key parameters that represent harmonics present in the target audio.

4 . The method of claim 1 , wherein:

the input audio and the target audio include respective sequences of audio frames;

the estimating the frame-specific key parameters includes estimating the key parameters on a frame-by-frame basis; and

the configuring the neural network includes configuring the neural network with the decoded frame-specific key parameters estimated on a frame-by-frame basis to cause the neural network to perform the signal transformation on the frame-by-frame basis, to produce the output audio as a sequence of audio frames.

5 . An apparatus comprising:

a decoder to decode encoded input audio and encoded frame-specific key parameters in a received combined bit-stream from a transmission channel to produce decoded input audio and decoded frame-specific key parameters, respectively, wherein the received combined bit-stream includes the encoded frame-specific key parameters in alignment with the encoded input audio at a frame level, wherein the decoded frame-specific key parameters include quantized representations of a target audio characteristic derived from a joint analysis of corresponding frames from target audio and the decoded input audio, wherein the decoded frame-specific key parameters vary from frame to frame and the quantized representations are based on an analysis including a temporal analysis and at least one of a spectral analysis or a frequency harmonic analysis; and

a neural network trained to be configured by the decoded frame-specific key parameters as produced by the decoder to perform a signal transformation of audio representative of the decoded input audio and

to produce output audio having an output audio characteristic corresponding to and that matches the target audio characteristic, wherein the output audio has perceptually-improved signal quality with frequency bandwidth extension as compared to the decoded input audio.

6 . The apparatus of claim 5 , wherein:

the audio representative of the decoded input audio includes a sequence of audio frames;

the decoded frame-specific key parameters include a sequence of frame-by-frame key parameters that represent the target audio characteristic on a frame-by-frame basis; and

the neural network is configured by the sequence of frame-by-frame key parameters to perform the signal transformation of the audio representative of the decoded input audio on a frame-by frame basis, to produce the output audio as a sequence of output audio frames.

7 . The apparatus of claim 5 , further comprising a pre-processor to pre-process the input audio to produce pre-processed input audio as the audio representative of the input audio.

8 . The apparatus of claim 5 , wherein the audio representative of the input audio includes the input audio.

9 . The apparatus of claim 5 , wherein the decoder is further configured to demultiplex the encoded input audio and the encoded frame-specific key parameters from a multiplexed signal, and then decode of the encoded input audio and the encoded key parameters.

10 . The apparatus of claim 5 , further comprising:

a blending unit providing a blending operation to blend the decoded input audio with the output audio produced by the neural network.

11 . A method comprising:

receiving input audio and frame-specific key parameters that are representative of a target audio characteristic in a multiplexed and combined bit-stream in which both the input audio and the frame-specific key parameters are encoded, wherein the received combined bit-stream includes the encoded frame-specific key parameters in alignment with the encoded input audio at a frame level, wherein the frame-specific key parameters include quantized representations of the target audio characteristic derived from a joint analysis of corresponding frames from target audio and the input audio, wherein the frame-specific key parameters vary from frame to frame and the quantized representations are based on an analysis including a temporal analysis and at least one of a spectral analysis or a frequency harmonic analysis;

demultiplexing and decoding the encoded input audio and the encoded frame-specific key parameters to recover the input audio and the frame-specific key parameters;

configuring a neural network, that was previously trained to be configured by the frame-specific key parameters, with the frame-specific key parameters as decoded to cause the neural network to perform a signal transformation of audio that is representative of the input audio; and

producing output audio with an output audio characteristic that matches the target audio characteristic, wherein the output audio has perceptually-improved signal quality with frequency bandwidth extension as compared to the input audio.

12 . The method of claim 11 , wherein:

the input audio and the audio include respective sequences of audio frames;

the frame-specific key parameters represent the target audio characteristic on a frame-by-frame basis;

and

the neural network is configured by the key parameters to perform the signal transformation on a frame-by-frame basis, to produce the output audio as a sequence of output audio frames.

13 . The method of claim 11 , further comprising pre-processing the input audio to produce pre-processed input audio as the audio.