IP Library Granted Patent US 12664990
Granted Patent B2
US 12664990 · App. 18/207,554 · Granted Jun 23, 2026

Audio processing of missing audio information

Inventors: Wei Xiong (Shenzhen, CN); Fei Huang (Shenzhen, CN)
Assignee: Tencent Technology (Shenzhen) Company Limited
G10L19/005G10L19/06G10L21/0316
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12664990
App. No.
18/207,554
Granted
Jun 23, 2026
Kind
B2
Abstract

Target audio data and frequency spectrum information of the target audio data is acquired. The target audio data includes an audio missing segment and context audio segments of the audio missing segment. The frequency spectrum information includes frequency spectrum features of the context audio segments. Feature compensation is performed on the frequency spectrum information of the target audio data based on the frequency spectrum features of the context audio segments to obtain compensated frequency spectrum information corresponding to the target audio data. The compensated frequency spectrum information indicates upsampled frequency spectrum information of the target audio data. Audio prediction is performed based on the compensated frequency spectrum information to obtain predicted audio data. The audio missing segment in the target audio data is compensated by replacing the audio missing segment with a predicted segment in the predicted audio data to obtain compensated audio data of the target audio data.

Claims (100)

1 . An audio processing method, comprising:

acquiring target audio data and frequency spectrum information of the target audio data, the target audio data corresponding to a time interval in a time domain and including an audio missing segment, context audio segments, and other valid audio segments, the context audio segments including a first valid audio segment preceding the audio missing segment and a second valid audio segment succeeding the audio missing segment, and the frequency spectrum information including frequency spectrum features of the target audio data corresponding to the time interval irrespective of a position of the audio missing segment within the target audio data;

performing feature compensation on the frequency spectrum information of the target audio data based on the frequency spectrum features of the target audio data to obtain compensated frequency spectrum information corresponding to the target audio data, the compensated frequency spectrum information indicating upsampled frequency spectrum information associated with the target audio data and including additional feature information that is not included in the frequency spectrum information;

generating predicted audio data corresponding to the time interval in the time domain based on the compensated frequency spectrum information; and

obtaining compensated audio data based on modification of a portion of the target audio data according to a corresponding portion of the predicted audio data, the portion of the target audio data corresponding to the audio missing segment and at least a subset of the context audio segments.

2 . The method according to claim 1 , wherein the obtaining the compensated audio data comprises:

replacing the audio missing segment in the predicted audio data with a predicted segment in the predicted audio data corresponding to the audio missing segment in the predicted audio data.

3 . The method according to claim 2 , wherein the obtaining the compensated audio data comprises:

acquiring one or more fusion parameters;

smoothing one or more associated predicted segments in the predicted audio data that correspond to the subset of the context audio segments based on the one or more fusion parameters to obtain one or more smoothed associated predicted segments; and

replacing the subset of the context audio segments with the one or more smoothed associated predicted segments.

4 . The method according to claim 1 , wherein:

the target audio data is extracted by a generator from training audio data,

the generator is configured to specify a sampling rate and an audio extraction length, and

the acquiring the target audio data includes:

based on the generator,

extracting intermediate audio data with a length equal to the audio extraction length from the training audio data;

sampling the intermediate audio data according to the sampling rate to obtain a sampling sequence of the intermediate audio data; and

performing a simulative packet loss adjustment on a plurality of sampling points in the sampling sequence according to a preset packet loss length to obtain the target audio data such that audio data of the plurality of sampling points is zero, the plurality of sampling points subjected to the simulative packet loss adjustment being set as the audio missing segment in the target audio data.

5 . The method according to claim 4 , the method further comprises:

extracting feature maps at a plurality of resolutions from the compensated audio data based on a discriminator;

determining a feature difference between the compensated audio data and the intermediate audio data according to the feature maps at the plurality of resolutions; and

training the generator and the discriminator based on the feature difference to obtain a trained generator and a trained discriminator.

6 . The method according to claim 5 , wherein:

the generator and the discriminator are trained based on a loss function,

the loss function includes a multi-resolution loss function, and

the method further includes:

determining a frequency spectrum feature of each of the feature maps of the compensated audio data at a respective resolution;

acquiring a frequency spectrum feature of the intermediate audio data;

obtaining a spectrum convergence function associated with each of the feature maps at a respective resolution based on the frequency spectrum feature of the corresponding one of the feature maps at the respective resolution and the frequency spectrum feature of the intermediate audio data, the spectrum convergence function indicating a frequency spectrum difference between the frequency spectrum feature of the intermediate audio data and the corresponding one of the feature maps at the respective resolution;

solving the multi-resolution loss function based on the spectrum convergence functions associated with the feature maps at the plurality of resolutions; and

training the generator and the discriminator based on the multi-resolution loss function.

7 . The method according to claim 6 , wherein:

the obtaining the spectrum convergence function includes:

determining, based on the frequency spectrum feature of the intermediate audio data being a reference feature, the spectrum convergence function of each of the feature maps at the respective resolution according to the corresponding one of the frequency spectrum features of the feature maps at the respective resolution and the frequency spectrum feature of the intermediate audio data, and

the solving the multi-resolution loss function includes:

acquiring a frequency spectrum magnitude difference between each of the frequency spectrum features of the feature maps at the respective resolution and the frequency spectrum feature of the intermediate audio data to obtain a magnitude difference function; and

determining the multi-resolution loss function by weighting the spectrum convergence functions associated with the feature maps at the plurality of resolutions and the magnitude difference functions corresponding to the feature maps.

8 . The method according to claim 5 , wherein:

the generator and the discriminator are trained based on a loss function,

the loss function further includes a discriminator loss function and a generator loss function, and

the method further comprises:

determining a time domain feature of the compensated audio data according to the feature maps of the compensated audio data at the plurality of resolutions;

acquiring a time domain feature difference between the compensated audio data and the intermediate audio data from which the target audio data is obtained;

acquiring a consistency discrimination result that indicates whether the compensated audio data and the intermediate audio data are consistent by the discriminator based on the time domain feature difference;

solving the discriminator loss function according to the consistency discrimination result;

solving the generator loss function based on the time domain feature difference; and

training the generator and the discriminator based on the generator loss function and the discriminator loss function.

9 . The method according to claim 4 , wherein:

the frequency spectrum information of the target audio data further includes a frequency spectrum feature of the audio missing segment, and

the performing the feature compensation includes:

smoothing the frequency spectrum feature of the audio missing segment based on the frequency spectrum features of the context audio segments to obtain a smoothed frequency spectrum feature; and

obtaining the compensated frequency spectrum information corresponding to the target audio data based on the frequency spectrum features of the context audio segments and the smoothed frequency spectrum feature of the audio missing segment.

10 . The method according to claim 9 , wherein the performing the feature compensation further comprises:

acquiring a frequency spectrum length of the frequency spectrum information that includes the frequency spectrum features of the context audio segments and the smoothed frequency spectrum feature of the audio missing segment, the frequency spectrum length indicating a number of feature points in the frequency spectrum information;

determining a number of sampling points in the sampling sequence corresponding to the target audio data based on the sampling rate and the audio extraction length;

upsampling the frequency spectrum information according to the number of the sampling points in the sampling sequence such that the number of feature points in the frequency spectrum information is equal to the number of sampling points in the sampling sequence; and

setting the upsampled frequency spectrum information as the compensated frequency spectrum information corresponding to the target audio data.

11 . The method according to claim 10 , wherein:

the upsampling is performed one or more times on the frequency spectrum information according to the number of sampling points, and

the method further includes:

performing multi-scale convolution operation on the upsampled frequency spectrum information to obtain the frequency spectrum features of the context audio segments.

12 . The method according to claim 1 , wherein the generating the predicted audio data comprises:

adjusting a number of frequency spectrum channels of the compensated frequency spectrum information to 1; and

using the adjusted compensated frequency spectrum information to generate the predicted audio data.

13 . An apparatus for audio processing, the apparatus comprising:

processing circuitry configured to:

acquire target audio data and frequency spectrum information of the target audio data, the target audio data corresponding to a time interval in a time domain and including an audio missing segment, context audio segments, and other valid audio segments, the context audio segments including a first valid audio segment preceding the audio missing segment and a second valid audio segment succeeding the audio missing segment, and the frequency spectrum information including frequency spectrum features of the target audio data corresponding to the time interval irrespective of a position of the audio missing segment within the target audio data;

perform feature compensation on the frequency spectrum information of the target audio data based on the frequency spectrum features of the target audio data to obtain compensated frequency spectrum information corresponding to the target audio data, the compensated frequency spectrum information indicating upsampled frequency spectrum information associated with the target audio data and including additional feature information that is not included in the frequency spectrum information;

generate predicted audio data corresponding to the time interval in the time domain based on the compensated frequency spectrum information; and

obtain compensated audio data based on modification of a portion of the target audio data according to a corresponding portion of the predicted audio data, the portion of the target audio data corresponding to the audio missing segment and at least a subset of the context audio segments.

14 . The apparatus according to claim 13 , wherein the processing circuitry is configured to:

replace the audio missing segment in the predicted audio data with a predicted segment in the predicted audio data corresponding to the audio missing segment in the predicted audio data.

15 . The apparatus according to claim 14 , wherein the processing circuitry is configured to:

acquire one or more fusion parameters;

smooth one or more associated predicted segments in the predicted audio data that correspond to the subset of the context audio segments based on the one or more fusion parameters to obtain one or more smoothed associated predicted segments; and

replace the subset of the context audio segments with the one or more smoothed associated predicted segments.

16 . The apparatus according to claim 13 , wherein:

the target audio data is extracted by a generator from training audio data,

the generator is configured to specify a sampling rate and an audio extraction length, and

the processing circuitry is configured to:

based on the generator,

extract intermediate audio data with a length equal to the audio extraction length from the training audio data;

sample the intermediate audio data according to the sampling rate to obtain a sampling sequence of the intermediate audio data; and

perform a simulative packet loss adjustment on a plurality of sampling points in the sampling sequence according to a preset packet loss length to obtain the target audio data such that audio data of the plurality of sampling points is zero, the plurality of sampling points subjected to the simulative packet loss adjustment being set as the audio missing segment in the target audio data.

17 . The apparatus according to claim 16 , wherein the processing circuitry is configured to:

extract feature maps at a plurality of resolutions from the compensated audio data based on a discriminator;

determine a feature difference between the compensated audio data and the intermediate audio data according to the feature maps at the plurality of resolutions; and

train the generator and the discriminator based on the feature difference to obtain a trained generator and a trained discriminator.

18 . A non-transitory computer readable storage medium storing instructions which when executed by at least one processor cause the at least one processor to perform:

acquiring target audio data and frequency spectrum information of the target audio data, the target audio data corresponding to a time interval in a time domain and including an audio missing segment, context audio segments, and other valid audio segments, the context audio segments including a first valid audio segment preceding the audio missing segment and a second valid audio segment succeeding the audio missing segment, and the frequency spectrum information including frequency spectrum features of the target audio data corresponding to the time interval irrespective of a position of the audio missing segment within the target audio data;

performing feature compensation on the frequency spectrum information of the target audio data based on the frequency spectrum features of the target audio data to obtain compensated frequency spectrum information corresponding to the target audio data, the compensated frequency spectrum information indicating upsampled frequency spectrum information associated with the target audio data and including additional feature information that is not included in the frequency spectrum information;

generating predicted audio data corresponding to the time interval in the time domain based on the compensated frequency spectrum information; and

obtaining compensated audio data based on modification of a portion of the target audio data according to a corresponding portion of the predicted audio data, the portion of the target audio data corresponding to the audio missing segment and at least a subset of the context audio segments.

19 . The non-transitory computer readable storage medium according to claim 18 , wherein the obtaining the compensated audio data comprises:

replacing the audio missing segment in the predicted audio data with a predicted segment in the predicted audio data corresponding to the audio missing segment in the predicted audio data.

20 . The non-transitory computer readable storage medium according to claim 19 , wherein the obtaining the compensated audio data comprises:

acquiring one or more fusion parameters;

smoothing one or more associated predicted segments in the predicted audio data that correspond to the subset of the context audio segments based on the one or more fusion parameters to obtain one or more smoothed associated predicted segments; and

replacing the subset of the context audio segments with the one or more smoothed associated predicted segments.