IP Library Granted Patent US 9,230,555
Granted Patent B2
US 9,230,555 · App. 13/260,846 · Granted Jan 5, 2016

Apparatus and method for generating an output audio data signal

Inventors: Holly Francois (Guildford, GB); Jonathan A. Gibbs (Windermere, GB)
Assignee: GOOGLE TECHNOLOGY HOLDINGS LLC
G10L19/24
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 9,230,555
App. No.
13/260,846
Granted
Jan 5, 2016
Kind
B2
Abstract

An apparatus receives an input encoded audio data signal comprising a base layer and at least one enhancement layer. A reference unit ( 103 ) generates reference audio data corresponding to audio data of a reference set of layers. A layer unit ( 105 ) divides the layers of the input signal into a first subset and a second subset. A sample unit ( 107 ) generates sample audio data corresponding to the audio data of the first subset. A comparison unit ( 109 ) generates a difference measure by comparing the sample audio data to the reference audio data based on a perceptual model. An output unit ( 111 ) then determines if the difference measure meets a similarity criterion and generates an output signal without audio data from a layer of the second subset if the similarity criterion is met and including the audio data of the layer otherwise. The invention may provide reduced data rates without an unacceptable degradation of quality.

Claims (74)

1. An apparatus for generating an output audio data signal, the apparatus comprising:

a receiving device for receiving an input encoded audio data signal comprising a plurality of encoding layers including a base layer and a plurality of enhancement layers;

a reference unit for generating reference audio data from a reference set of layers of the plurality of encoding layers;

a sampling device for generating sample audio data from a set of layers smaller than the reference set of layers;

a comparison processor for comparing the sample audio data to the reference audio data, the comparison reflecting a difference between a first decoded signal corresponding to the sample audio data and a second decoded signal corresponding to the reference audio data;

an output device for determining whether the comparison meets a criterion and

if so, generating the output audio data signal to not include audio data from a first layer, the first layer being a layer of the reference set not included in the smaller set of layers;

and otherwise, generating the output audio data signal to include audio data from the first wherein the comparison is based on a perceptual model,

wherein the comparison processor is configured to:

generate a first perceptual indication by applying the perceptual model to the reference audio data; and

generate a second perceptual indication by applying the perceptual model to the sample audio data; and

the output device is arranged to determine whether the comparison meets the criterion in response to a comparison of the first perceptual indication and the second perceptual indication,

wherein the perceptual model is configured to:

determine an energy measure for each of a plurality of critical bands;

apply a loudness compensation to the energy measure of each of the plurality of critical bands to generate a perceptual indication comprising loudness compensated energy measures for each of the critical bands; and

the output device is further arranged to determine whether the comparison meets the criterion in response to a comparison of the loudness compensated energy measures for each of the critical bands for the reference audio data and the sample audio data.

2. The apparatus of claim 1 wherein the reference audio data corresponds to a frequency domain representation of an audio signal represented by the audio data of layers of the reference set, and the sample audio data corresponds to a frequency domain representation of an audio signal represented by the audio data of layers of the smaller set of layers.

3. The apparatus of claim 2 wherein the frequency domain representation is an internal frequency domain representation of an encoding protocol of the input encoded audio data signal.

4. The apparatus of claim 1 arranged to generate the output audio data from a minimum number of layers required in the smaller set of layers for the comparison to meet the criterion.

5. The apparatus of claim 1 wherein the loudness compensation comprises determining a loudness compensated energy measure for a critical band as a function of:

(

a

+

b

P

P

R

)

γ

where a is a design parameter with a value in the interval [0.25;0.75]; b is a design parameter with a value in the interval [0.25;0.75]; P R is a reference energy value, P is an energy value for the critical band, and γ is a design parameter with a value in the interval [0.1;0.3].

6. The apparatus of claim 1 wherein:

the reference unit is arranged to generate the reference audio data as a time domain audio signal by decoding the audio data of the reference set of layers; and

the reference unit is arranged to generate the sample audio data as a time domain audio signal by decoding the audio data of the first subset of layers.

7. The apparatus of claim 1 wherein output device is arranged to generate the output audio data signal to include audio data from all layers of the plurality of encoding layers if the comparison does not meet the criterion.

8. The apparatus of claim 1 wherein the base layer comprises parametrically encoded speech data based on a speech model, and at least one layer of the reference set of layers not included in the smaller set of layers comprises waveform encoded audio data.

9. The apparatus of claim 1 wherein input encoded audio data signal is encoded in accordance with an International Telecommunication Union Telecommunication Standardization Sector, ITU-T, G.718 protocol.

10. A communication system including a network entity which comprises:

a receiving device for receiving an input encoded audio data signal comprising a plurality of encoding layers including a base layer and a plurality of enhancement layers;

a reference unit for generating reference audio data from a reference set of layers of the plurality of encoding layers;

a sampling device for generating sample audio data from a set of layers smaller than the reference set of layers;

a comparison processor for comparing the sample audio data itself to the reference audio data itself, the comparison reflecting a difference between a first decoded signal corresponding to the sample audio data and a second decoded signal corresponding to the reference audio data;

an output device for determining whether the comparison meets a criterion and

if so, generating the output audio data signal to not include audio data from a first layer, the first layer being a layer of the reference set not included in the smaller set of layers;

and otherwise, generating the output audio data signal to include audio data from the first layer,

wherein the comparison is based on a perceptual model,

wherein the comparison processor is configured to:

generate a first perceptual indication by applying the perceptual model to the reference audio data;

generate a second perceptual indication by applying the perceptual model to the sample audio data; and

the output device is arranged to determine whether the comparison meets the criterion in response to a comparison of the first perceptual indication and the second perceptual indication,

wherein the perceptual model is configured to:

determine an energy measure for each of a plurality of critical bands;

apply a loudness compensation to the energy measure of each of the plurality of critical bands to generate a perceptual indication comprising loudness compensated energy measures for each of the critical bands; and

the output device is further arranged to determine whether the comparison meets the criterion in response to a comparison of the loudness compensated energy measures for each of the critical bands for the reference audio data and the sample audio data.

11. The communication system of claim 10 wherein the network entity is a Radio Access Network network element of a cellular communication system.

12. The communication system of claim 11 further comprising an allocating unit for allocating an air interface resource in response to a set of layers included in the output audio data signal.

13. A method for generating an output audio data signal, the method comprising:

receiving an input encoded audio data signal comprising a plurality of encoding layers including a base layer and a plurality of enhancement layers;

generating reference audio data from a reference set of layers of the plurality of encoding layers;

generating sample audio data from a set of layers smaller than the reference set of layers;

comparing the sample audio data itself to the reference audio data itself, the comparison reflecting a difference between a first decoded signal corresponding to the sample audio data and a second decoded signal corresponding to the reference audio data;

determining whether the comparison meets a criterion and

if so, generating the output audio data signal to not include audio data from a first layer, the first layer being a layer of the reference set not included in the smaller set of layers;

and otherwise, generating the output audio data signal to include audio data from the first layer,

wherein the comparison is based on a perceptual model,

wherein the comparison step further comprises:

generating a first perceptual indication by applying the perceptual model to the reference audio data;

generating a second perceptual indication by applying the perceptual model to the sample audio data;

the method further comprising determining whether the comparison meets the criterion in response to a comparison of the first perceptual indication and the second perceptual indication,

wherein the perceptual model is configured to:

determine an energy measure for each of a plurality of critical bands; and

apply a loudness compensation to the energy measure of each of the plurality of critical bands to generate a perceptual indication comprising loudness compensated energy measures for each of the critical bands; and

the method further comprises determining whether the comparison meets the criterion in response to a comparison of the loudness compensated energy measures for each of the critical bands for the reference audio data and the sample audio data.

Assignments (4)
CORRECTIVE ASSIGNMENT TO CORRECT THE REMOVE INCORRECT PATENT NO. 8577046 AND REPLACE WITH CORRECT PATENT NO. 8577045 PREVIOUSLY RECORDED ON REEL 034286 FRAME 0001. ASSIGNOR(S) HEREBY CONFIRMS THE ASSIGNMENT. Recorded Dec 3, 2014
From: MOTOROLA MOBILITY LLC
To: GOOGLE TECHNOLOGY HOLDINGS LLC
Reel/Frame 034538/0001 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Nov 18, 2014
From: MOTOROLA MOBILITY LLC
To: GOOGLE TECHNOLOGY HOLDINGS LLC
Reel/Frame 034286/0001 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Aug 21, 2012
From: MOTOROLA MOBILITY, INC.
To: MOTOROLA MOBILITY LLC
Reel/Frame 028829/0856 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Sep 28, 2011
From: FRANCOIS, HOLLY L; GIBBS, JONATHAN A
To: MOTOROLA MOBILITY, INC.
Reel/Frame 026984/0087 →
Priority Claims (1)
EP 09157046 · Apr 1, 2009 · regional
Continuity (1)
Related Publication 20120116560A1 · May 10, 2012