IP Library Granted Patent US 9,654,076
Granted Patent B2
US 9,654,076 · App. 14/613,203 · Granted May 16, 2017

Metadata for ducking control

Inventors: Tomlinson M. Holman (Cupertino, CA); Frank M. Baumgarte (Sunnyvale, CA); Eric A. Allamanche (Sunnyvale, CA)
Assignee: Apple Inc.
H03G3/3089G10L19/008H04N21/4396H04N21/8106H04R27/02H04R2227/003H04R2227/009
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 9,654,076
App. No.
14/613,203
Granted
May 16, 2017
Kind
B2
Abstract

An audio encoding device and an audio decoding device are described herein. The audio encoding device may examine a set of audio channels/channel groups representing a piece of sound program content and produce a set of ducking values to associate with one of the channels/channel groups. During playback of the piece of sound program content, the ducking values may be applied to all other channels/channel groups. Application of these ducking values may cause (1) the reduction in dynamic range of ducked channels/channel groups and/or (2) movement of channels/channel groups in the sound field. This ducking may improve intelligibility of audio in the non-ducked channel/channel group. For instance, a narration channel/channel group may be more clearly heard by listeners through the use of selective ducking of other channels/channel groups during playback.

Claims (42)

1. A method for playing back audio content, the method comprising:

receiving an audio asset representing a piece of sound program content comprising (i) a first channel group, object or stem that represents a first type of audio, (ii) a second channel group, object or stem that represents a second type of audio, (iii) a third channel group, object or stem that represents a third type of audio, and (iv) a first set of ducking values and a second set of ducking values associated with the first channel group, object or stem, wherein the first, second, and third types of audio are different, wherein the first set of ducking values is different than the second set of ducking values;

extracting the first and second sets of ducking values along with the (i) first channel group, object or stem, (ii) second channel group, object or stem, and (iii) third channel group, object or stem from the audio asset; and

during playback of the piece of sound program content through a plurality of loudspeakers

applying the first set of ducking values to the second channel group, object or stem; and

applying the second set of ducking values to the third channel group, object or stem, wherein applying the first and second sets of ducking values deemphasizes the second channel group, object or stem differently than the third channel group, object or stem.

2. The method of claim 1 , wherein application of the first and second sets of ducking values deemphasize by reducing dynamic range of the (i) first channel group, object or stem and (ii) second channel group, object or stem during playback.

3. The method of claim 2 , further comprising:

applying a scale factor to the first and second sets of ducking values prior to application of the first and second sets of ducking values.

4. The method of claim 1 further comprising producing a set of drive signals based on the first, second, and third channel groups, objects or stems to drive the plurality of loudspeakers to render sound in a sound field, wherein applying the first set of ducking values causes the rendering location of the second channel group, object or stem in the sound field to move to a different rendering location in the sound field during playback.

5. The method of claim 1 , wherein the audio asset is associated with video content, wherein the first type of audio is a narration of the video content, such that the first channel group, object or stem comprises audio content that describes actions taking place in the video content.

6. The method of claim 5 , wherein the second type of audio is dialogue of the video content and the third type of audio is music and effects of the video content.

7. An audio device for playing back audio content, the audio device comprising:

a hardware processor, and

a memory unit storing instructions to be executed by the hardware processor that cause the audio device to:

receive an audio asset representing a piece of sound program content comprising (i) a first channel group, object or stem that represents a first type of audio, (ii) a second channel group, object or stem that represents a second type of audio, (iii) a third channel group, object or stem that represents a third type of audio, and (iv) a first set of ducking values and a second set of ducking values associated with the first channel group or object or stem, wherein the first, second, and third types of audio are different, and wherein the first set of ducking values is different than the second set of ducking values;

extract the first and second sets of ducking values along with the first channel group, object or stem, the second channel group, object or stem, and the third channel group, object or stem, from the audio asset; and

apply, during playback of the piece of sound program content through a plurality of loudspeakers, (i) the first set of ducking values to the second channel group, object or stem and (ii) the second set of ducking values to the third channel group, object or stem, wherein application of the first and second sets of ducking values is to deemphasize the second channel group, object or stem differently than the third channel group, object or stem.

8. The audio device of claim 7 , wherein application of the first and second sets of ducking values deemphasize by reducing dynamic range of the (i) first channel group, object or stem and (ii) second channel group, object or stem during playback.

9. The audio device of claim 8 , wherein the memory unit includes further instructions, which when executed by the hardware processor cause the audio device to:

apply a scale factor to the first and second sets of ducking values prior to application of the first and second sets of ducking values.

10. The audio device of claim 7 , wherein the memory unit includes further instructions, which when executed by the hardware processor cause the audio device to

produce a set of drive signals based on the first, second, and third channel groups, objects or stems, to drive the plurality of loudspeakers to render sound in a sound field, wherein the instructions to apply the first set of ducking values cause a rendering location of the second channel group, object or stem in the sound field to move to a different location in the sound field during playback.

11. The audio device of claim 7 , wherein the audio asset is associated with video content, wherein the first type of audio is narration of the video content, such that the first channel group, object or stem comprises visually descriptive audio content that describes actions taking place in the video content.

12. The audio device of claim 11 , wherein the second type of audio is dialogue of the video content and the third type of audio is music and effects of the video content.

13. A method for playing back audio content, the method comprising:

receiving a piece of sound program content comprising (i) a first channel group, object or stem that represents a first type of audio, (ii) a second channel group, object or stem that represents a second type of audio, and (iii) ducking values that are associated with the first channel group, object or stem, wherein the first and second types of audio are different;

producing a set of drive signals based on the piece of sound program content to drive a plurality of loudspeakers to render sound in a sound field, such that (i) sound of the first channel group, object or stem is rendered and (ii) sound of the second channel group, object or stem is rendered at an original location within the sound field;

applying the ducking values to the second channel group, object or stem; and

based on applying the ducking values, adjusting the set of drive signals to cause the sound of the second channel group, object or stem to be rendered at a different location within the sound field.

14. The method of claim 13 wherein the ducking values indicate that the rendering of the second channel group, object or stem is to be moved from front loudspeakers to surround loudspeakers during playback.

15. The method of claim 14 , wherein the movement of the second channel group, object or stem is only when speech activity is detected in the first channel group, object or stem.

16. The method of claim 13 , wherein the audio asset is associated with video content, wherein (i) the first type of audio is narration of the video content, such that the first channel group, object or stem comprises visually descriptive audio content that describes actions taking place in the video content and (ii) the second type of audio is dialogue of the video content.

17. An audio device for playing back audio content, the audio device comprising:

a hardware processor, and

a memory unit storing instructions to be executed by the hardware processor that cause the audio device to:

receive a piece of sound program content comprising (i) a first channel group, object or stem that represents a first type of audio, (ii) a second channel group, object or stem that represents a second type of audio, and (iii) ducking values that are associated with the first channel group, object or stem, wherein the first and second types of audio are different;

produce a set of drive signals based on the piece of sound program content to drive a plurality of loudspeakers to render sound in a sound field, such that (i) sound of the first channel group, object or stem is rendered and (ii) sound of the second channel group, object or stem is rendered at an original location within the sound field;

apply the ducking values to the second channel group, object or stem; and

based on applying the ducking values, adjust the set of drive signals to cause the sound of the second channel group, object or stem to be rendered at a different location within the sound field.

18. The audio device of claim 17 , wherein the ducking values indicate that the rendering of the second channel group, object or stem is to be moved from front loudspeaker to surround loudspeakers during playback.

19. The audio device of claim 17 wherein the memory unit has further instructions stored therein that when executed by the hardware processor move the rendering location of the second channel group, object or stem only when speech activity is detected in the first channel group, object or stem.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Feb 3, 2015
From: HOLMAN, TOMLINSON M.; BAUMGARTE, FRANK M.; ALLAMANCHE, ERIC A.
To: APPLE INC.
Reel/Frame 034880/0557 →
Continuity (2)
Provisional Application 61970284 · Mar 25, 2014
Related Publication 20150280676A1 · Oct 1, 2015