IP Library Granted Patent US 9,761,239
Granted Patent B2
US 9,761,239 · App. 15/386,246 · Granted Sep 12, 2017

Hybrid encoding method and apparatus for encoding speech or non-speech frames using different coding algorithms

Inventor: Zhe Wang (Beijing, CN)
Assignee: HUAWEI TECHNOLOGIES CO., LTD.
G10L19/22G10L19/0204G10L19/035G10L19/06G10L19/20G10L19/02G10L19/04G10L25/03
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 9,761,239
App. No.
15/386,246
Granted
Sep 12, 2017
Kind
B2
Abstract

An audio encoding method and an apparatus are provided. The method includes: determining sparseness of distribution, on spectrums, of energy of N input audio frames ( 101 ), where the N audio frames include a current audio frame, and N is a positive integer; and determining, according to the sparseness of distribution, on the spectrums, of the energy of the N audio frames, whether to use a first encoding method or a second encoding method to encode the current audio frame ( 102 ), where the first encoding method is an encoding method that is based on time-frequency transform and transform coefficient quantization and that is not based on linear prediction, and the second encoding method is a linear-predication-based encoding method. The method can reduce encoding complexity and ensure that encoding is of relatively high accuracy.

Claims (82)

1. An audio encoding method, wherein the method comprises:

determining sparseness of distribution in energy spectrums of N audio frames, wherein the N audio frames comprise a current audio frame, and N is a positive integer; and

determining, according to the sparseness of distribution, whether to use a first encoding method or a second encoding method to encode the current audio frame, wherein the first encoding method is based on time-frequency transform and transform coefficient quantization, the first encoding method is not based on linear prediction, and the second encoding method is a linear-predication-based encoding method,

wherein the determining the sparseness of distribution comprises:

dividing an energy spectrum of each of the N audio frames into P spectral envelopes, wherein P is a positive integer, and

determining a general sparseness parameter according to energy of the P spectral envelopes of each of the N audio frames, wherein the general sparseness parameter indicates the sparseness of distribution.

2. The method according to claim 1 , wherein the general sparseness parameter comprises a first minimum bandwidth, and wherein the determining the general sparseness parameter comprises:

determining an average value of minimum bandwidths, distributed on the energy spectrums, of a first preset proportion of energy of the N audio frames according to the energy of the P spectral envelopes of each of the N audio frames, wherein the average value of the minimum bandwidths of the first preset proportion of the energy of the N audio frames is used as the first minimum bandwidth, and

wherein the first encoding method is determined to be used to encode the current audio frame when the first minimum bandwidth is less than a first preset value, or the second encoding method is determined to be used to encode the current audio frame when the first minimum bandwidth is greater than the first preset value.

3. The method according to claim 2 , wherein the determining the average value of minimum bandwidths of the first preset proportion of the energy of the N audio frames comprises:

sorting the energy of the P spectral envelopes of each audio frame in descending order;

determining, according to the energy, sorted in descending order, of the P spectral envelopes of each of the N audio frames, a minimum bandwidth, distributed on the energy spectrums, of energy that accounts for not less than the first preset proportion of each of the N audio frames; and

determining, according to the minimum bandwidth, distributed on the energy spectrums, of the energy that accounts for not less than the first preset proportion of each of the N audio frames, an average value of minimum bandwidths, distributed on the energy spectrums, of energy that accounts for not less than the first preset proportion of the N audio frames.

4. The method according to claim 1 , wherein the general sparseness parameter comprises a first energy proportion, and wherein

the determining the general sparseness parameter comprises:

selecting P 1 spectral envelopes from the P spectral envelopes of each of the N audio frames; and

determining the first energy proportion according to energy of the P 1 spectral envelopes of each of the N audio frames and total energy of the N audio frames, wherein P 1 is a positive integer less than P,

wherein the first encoding method is determined to be used to encode the current audio frame when the first energy proportion is greater than a second preset value, or the second encoding method is determined to be used to encode the current audio frame when the first energy proportion is less than the second preset value.

5. The method according to claim 4 , wherein energy of any one of the P 1 spectral envelopes is greater than energy of any one of spectral envelopes in the P spectral envelopes other than the P 1 spectral envelopes.

6. The method according to claim 1 , wherein the general sparseness parameter comprises a second minimum bandwidth and a third minimum bandwidth, and wherein

the determining the general sparseness parameter comprises:

determining an average value of minimum bandwidths, distributed on the energy spectrums, of a second preset proportion of the energy of the N audio frames according to the energy of the P spectral envelopes of each of the N audio frames; and

determining an average value of minimum bandwidths, distributed on the energy spectrums, of a third preset proportion of the energy of the N audio frames according to the energy of the P spectral envelopes of each of the N audio frames,

wherein the average value of the minimum bandwidths of the second preset proportion of the energy of the N audio frames is used as the second minimum bandwidth,

wherein the average value of the minimum bandwidths of the third preset proportion of the energy of the N audio frames is used as the third minimum bandwidth,

wherein the second preset proportion is less than the third preset proportion,

wherein the first encoding method is determined to be used to encode the current audio frame when the second minimum bandwidth is less than a third preset value and the third minimum bandwidth is less than a fourth preset value, or

the first encoding method is determined to be used to encode the current audio frame when the third minimum bandwidth is less than a fifth preset value, or

the second encoding method is determined to be used to encode the current audio frame when the third minimum bandwidth is greater than a sixth preset value, and wherein

the fourth preset value is greater than or equal to the third preset value, the fifth preset value is less than the fourth preset value, and the sixth preset value is greater than the fourth preset value.

7. The method according to claim 6 , wherein the determining the average value of minimum bandwidths of the second preset proportion of the energy of the N audio frames and the determining the average value of minimum bandwidths of the third preset proportion of the energy of the N audio frames comprises:

sorting the energy of the P spectral envelopes of each audio frame in descending order;

determining, according to the energy, sorted in descending order, of the P spectral envelopes of each of the N audio frames, a minimum bandwidth, distributed on the energy spectrum, of energy that accounts for not less than the second preset proportion of each of the N audio frames;

determining, according to the minimum bandwidth, distributed on the energy spectrums, of the energy that accounts for not less than the second preset proportion of each of the N audio frames, an average value of minimum bandwidths, distributed on the energy spectrums, of energy that accounts for not less than the second preset proportion of the N audio frames;

determining, according to the energy, sorted in descending order, of the P spectral envelopes of each of the N audio frames, a minimum bandwidth, distributed on the energy spectrums, of energy that accounts for not less than the third preset proportion of each of the N audio frames; and

determining, according to the minimum bandwidth, distributed on the energy spectrums, of the energy that accounts for not less than the third preset proportion of each of the N audio frames, an average value of minimum bandwidths, distributed on the energy spectrums, of energy that accounts for not less than the third preset proportion of the N audio frames.

8. The method according to claim 1 , wherein the general sparseness parameter comprises a second energy proportion and a third energy proportion, and wherein

the determining the general sparseness parameter comprises:

determining the second energy proportion according to energy of P 2 spectral envelopes of each of the N audio frames and total energy of the N audio frames;

determining the third energy proportion according to energy of P 3 spectral envelopes of each of the N audio frames and the total energy of the N audio frames, wherein P 2 and P 3 are positive integers less than P, and P 2 is less than P 3 ,

and wherein the first encoding method is determined to be used to encode the current audio frame when the second energy proportion is greater than a seventh preset value and the third energy proportion is greater than an eighth preset value, or

the first encoding method is determined to be used to encode the current audio frame when the second energy proportion is greater than a ninth preset value, or

the second encoding method is determined to be used to encode the current audio frame when the third energy proportion is less than a tenth preset value.

9. The method according to claim 8 , wherein the P 2 spectral envelopes have maximum energy among possible selections of P 2 spectral envelopes from the P spectral envelopes, and wherein

the P 3 spectral envelopes have maximum energy among possible selections of P 3 spectral envelopes from the P spectral envelopes.

10. An audio encoder, comprising:

a memory comprising instructions; and one or more processors in communication with the memory, wherein the one or more processors execute the instructions to:

obtain N audio frames, wherein the N audio frames comprise a current audio frame, and N is a positive integer;

determine sparseness of distribution in energy spectrums of the N audio frames; and

determine, according to the sparseness of distribution, whether to use a first encoding method or a second encoding method to encode the current audio frame, wherein the first encoding method is based on time-frequency transform and transform coefficient quantization, the first encoding method is not based on linear prediction, and the second encoding method is a linear-predication-based encoding method,

wherein, to determine the sparseness of distribution, the one or more processors execute instructions to:

divide an energy spectrum of each of the N audio frames into P spectral envelopes, and determine a general sparseness parameter according to energy of the P spectral envelopes of each of the N audio frames, wherein P is a positive integer, and the general sparseness parameter indicates the sparseness of distribution.

11. The audio encoder according to claim 10 , wherein the general sparseness parameter comprises a first minimum bandwidth, and wherein

to determine the general sparseness parameter, the one or more processors execute instructions to:

determine an average value of minimum bandwidths, distributed on the energy spectrums, of a first preset proportion energy of the N audio frames according to the energy of the P spectral envelopes of each of the N audio frames,

wherein the average value of the minimum bandwidths of the first preset proportion of the energy of the N audio frames is used as first minimum bandwidth,

and wherein the first encoding method is determined to be used to encode the current audio frame when the first minimum bandwidth is less than a first preset value, or the second encoding method is determined to be used to encode the current audio frame when the first minimum bandwidth is greater than the first preset value.

12. The audio encoder according to claim 11 , wherein, to determine the average value of minimum bandwidths, the one or more processors execute instructions to:

sort the energy of the P spectral envelopes of each audio frame in descending order;

determine, according to the energy, sorted in descending order, of the P spectral envelopes of each of the N audio frames, a minimum bandwidth, distributed on the energy spectrums, of energy that accounts for not less than the first preset proportion of each of the N audio frames; and

determine, according to the minimum bandwidth, distributed on the energy spectrums, of the energy that accounts for not less than the first preset proportion of each of the N audio frames, an average value of minimum bandwidths, distributed on the energy spectrums, of energy that accounts for not less than the first preset proportion of the N audio frames.

13. The audio encoder according to claim 10 , wherein the general sparseness parameter comprises a first energy proportion, and wherein,

to determine the general sparseness parameter, the one or more processors execute instructions to: select P 1 spectral envelopes from the P spectral envelopes of each of the N audio frames, and determine the first energy proportion according to energy of the P 1 spectral envelopes of each of the N audio frames and total energy of the N audio frames, wherein P 1 is a positive integer less than P; and

wherein the first encoding method is determined to be used to encode the current audio frame when the first energy proportion is greater than a second preset value, or the second encoding method is determined to be used to encode the current audio frame when the first energy proportion is less than the second preset value.

14. The audio encoder according to claim 13 , wherein energy of any one of the P 1 spectral envelopes is greater than energy of any one of spectral envelopes in the P spectral envelopes other than the P 1 spectral envelopes.

15. The audio encoder according to claim 10 , wherein the general sparseness parameter comprises a second minimum bandwidth and a third minimum bandwidth, and wherein,

to determine the general sparseness parameter, the one or more processors execute instructions to: determine an average value of minimum bandwidths, distributed on the energy spectrums, of a second preset proportion of the energy of the N audio frames according to the energy of the P spectral envelopes of each of the N audio frames and determine an average value of minimum bandwidths, distributed on the spectrums, of third preset proportion energy of the N audio frames according to the energy of the P spectral envelopes of each of the N audio frames, wherein the average value of the minimum bandwidths of the second preset proportion of the energy of the N audio frames is used as the second minimum bandwidth, the average value of the minimum bandwidths of the third preset proportion of the energy of the N audio frames is used as the third minimum bandwidth, and the second preset proportion is less than the third preset proportion;

wherein the first encoding method is determined to be used to encode the current audio frame when the second minimum bandwidth is less than a third preset value and the third minimum bandwidth is less than a fourth preset value, or the first encoding method is determined to be used to encode the current audio frame when the third minimum bandwidth is less than a fifth preset value, or the second encoding method is determined to be used to encode the current audio frame when the third minimum bandwidth is greater than a sixth preset value; and

wherein the fourth preset value is greater than or equal to the third preset value, the fifth preset value is less than the fourth preset value, and the sixth preset value is greater than the fourth preset value.

16. The audio encoder according to claim 15 , wherein, to determine the average value of minimum bandwidths, the one or more processors execute instructions to:

sort the energy of the P spectral envelopes of each audio frame in descending order;

determine, according to the energy, sorted in descending order, of the P spectral envelopes of each of the N audio frames, a minimum bandwidth, distributed on the energy spectrum, of energy that accounts for not less than the second preset proportion of each of the N audio frames;

determine, according to the minimum bandwidth, distributed on the energy spectrums, of the energy that accounts for not less than the second preset proportion of each of the N audio frames, an average value of minimum bandwidths, distributed on the energy spectrums, of energy that accounts for not less than the second preset proportion of the N audio frames;

determine, according to the energy, sorted in descending order, of the P spectral envelopes of each of the N audio frames, a minimum bandwidth, distributed on the energy spectrums, of energy that accounts for not less than the third preset proportion of each of the N audio frames; and

determine, according to the minimum bandwidth, distributed on the energy spectrums, of the energy that accounts for not less than the third preset proportion of each of the N audio frames, an average value of minimum bandwidths, distributed on the energy spectrums, of energy that accounts for not less than the third preset proportion of the N audio frames.

17. The audio encoder according to claim 10 , wherein the general sparseness parameter comprises a second energy proportion and a third energy proportion, and wherein

to determine the general sparseness parameter, the one or more processors specifically execute instructions to:

determine the second energy proportion according to energy of P 2 spectral envelopes of each of the N audio frames and total energy of the respective N audio frames;

determine the third energy proportion according to energy of P 3 spectral envelopes of each of the N audio frames and the total energy of the N audio frames, wherein P 2 and P 3 are positive integers less than P, and P 2 is less than P 3 ; and

wherein the first encoding method is determined to be used to encode the current audio frame when the second energy proportion is greater than a seventh preset value and the third energy proportion is greater than an eighth preset value, or the first encoding method is determined to be used to encode the current audio frame when the second energy proportion is greater than a ninth preset value, or the second encoding method is determined to be used to encode the current audio frame when the third energy proportion is less than a tenth preset value.

18. The audio encoder according to claim 17 , wherein the P 2 spectral envelopes have maximum energy among possible selections of P 2 spectral envelopes from the P spectral envelopes; and

wherein the P 3 spectral envelopes have maximum energy among possible selections of P 3 spectral envelopes from the P spectral envelopes.

Assignments (2)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Aug 30, 2023
From: HUAWEI TECHNOLOGIES CO., LTD.
To: TOP QUALITY TELEPHONY, LLC
Reel/Frame 064757/0541 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Feb 10, 2017
From: WANG, ZHE
To: HUAWEI TECHNOLOGIES CO., LTD.
Reel/Frame 041227/0300 →
Priority Claims (1)
CN 2014 1 0288983 · Jun 24, 2014 · national
Continuity (2)
Continuation PCTCN2015082076 · Jun 23, 2015
Related Publication 20170103768A1 · Apr 13, 2017