IP Library Granted Patent US 9,930,402
Granted Patent B2
US 9,930,402 · App. 14/785,745 · Granted Mar 27, 2018

Automated audio adjustment

Inventor: Junhua Hou (Shanghai, CN)
Assignee: Verizon Patent and Licensing Inc.
H04N21/439H04N21/4223H04N21/42203H04N21/44204H04N21/44218H04N21/4532
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 9,930,402
App. No.
14/785,745
Granted
Mar 27, 2018
Kind
B2
Abstract

In embodiments, apparatuses, methods and storage media are described that are associated with adjusting audio during content presentation. While content is being presented, one or more persons may be identified that are consuming the content. The persons may be identified via various techniques, such as voice recognition, facial recognition, and distance detection. When persons may be uniquely identified, user audio preferences may be retrieved and used to adjust audio. Audio may be adjusted when persons are not uniquely identified, such as based on a number of persons or their location relative to a content consumption device. Audio adjustment may include volume adjustment and/or adjustment of audio, such as through application of audio effects, as it being presented. Other embodiments may be described and claimed.

Claims (68)

1. A method for automatically adjusting audio content of media content received from a network device via a network, comprising:

identifying, by a presentation engine of a content consumption device, one or more persons consuming the media content presented via the content consumption device;

determining, by the presentation engine, a physical location of the one or more persons relative to a location of the content consumption device;

determining, by the presentation engine, when one of the one or more persons attempts to input a voice command to control actions of the content consumption device; and

automatically adjusting, by the presentation engine, the audio content via the content consumption device, wherein automatically adjusting the audio content comprises:

lowering an output volume of the audio content to facilitate capture, by a microphone associated with the content consumption device, of the input and accurate recognition of the voice command by the presentation engine, and

modifying, based on the physical location of the one or more persons relative to the location of the content consumption device, equalizer settings, balance settings, or fader settings for the audio content.

2. The method of claim 1 , wherein identifying one or more persons comprises uniquely identifying one or more persons.

3. The method of claim 2 , wherein automatically adjusting the audio content further comprises:

retrieving one or more preferences for the one or more identified persons, based at least in part on a result of the identifying; and

automatically adjusting the audio content based at least in part on the retrieved one or more preferences.

4. The method of claim 3 , wherein automatically adjusting the audio content based at least in part on the one or more preferences comprises:

determining multiple individual audio content adjustments based on the one or more preferences for the one or more identified persons; and

determining a combined audio adjustment based on the multiple individual audio content adjustments, wherein the combined audio content adjustment is generated based on each of the multiple individual audio content adjustments.

5. The method of claim 1 , wherein identifying one or more persons comprises:

capturing an image proximate to the content consumption device; and

visually identifying the one or more persons based at least in part on the image.

6. The method of claim 5 , wherein visually identifying one or more persons from the image comprises performing facial detection on the image to identify faces of the one or more persons in the image.

7. The method of claim 1 , wherein identifying one or more persons comprises:

capturing audio proximate to the content consumption device; and

performing voice recognition of the captured audio to identify voices for the one or more persons.

8. The method of claim 7 , wherein identifying voices comprises identifying the captured audio falling in a human vocal range.

9. The method of claim 8 , wherein automatically adjusting the audio content further comprises, in response to identifying voices, lowering the output volume of the audio content to reduce interference with the identified voices.

10. The method of claim 1 , wherein identifying one or more persons comprises detecting one or more distances of the one or more persons from the content presentation device.

11. The method of claim 1 , further comprising:

capturing audio proximate to the content consumption device;

identifying, based on the captured audio, background noise present during presentation of the media content, wherein the background noise comprises audio outside a human vocal range; and

wherein automatically adjusting the audio content comprises increasing an output volume of the audio content to compensate for the background noise.

12. The method of claim 1 , further comprising:

determining, based on the identifying, an increase in a number of persons consuming the media content presented via the content consumption device; and

increasing the output volume of the audio content to compensate for the increased number of persons.

13. The method of claim 1 , further comprising:

determining, based on the physical location, a spatial distribution of the one or more persons consuming the media content presented via the content consumption device; and

adjusting a left/right balance based on the determined spatial distribution.

14. An apparatus for automatically adjusting audio content of a media content received from a network device via a network, the apparatus comprising:

one or more computer processors;

one or more microphones for capturing audio proximate to a content consumption device;

one or more identification modules configured to operate on the one or more computer processors to:

identify one or more persons consuming the media content via the content consumption device, and

determine a physical location of the one or more persons relative to a location of the content consumption device; and

an audio adjustment manager configured to operate on the one or more computer processors to determine when one of the one or more persons attempts to input a voice command to control actions of the content consumption device,

wherein the audio adjustment manager is further configured to operate on the one or more computer processors to automatically adjust the audio content via the content consumption device, and

wherein automatically adjusting the audio content comprises:

lowering an output volume of the audio content to facilitate capture, by the one or more microphones, of the input over the audio content and accurate recognition of the voice command by the one or more identification modules, and

modifying, based on the physical location of the one or more persons relative to the location of the content consumption device, equalizer settings, balance settings, or fader settings for the audio content.

15. The apparatus of claim 14 , wherein the audio adjustment manager configured to automatically adjust the audio content is further configured to:

retrieve one or more preferences for the one or more identified persons, based at least in part on a result of the one or more identification modules; and

automatically adjust the audio content based at least in part on the one or more preferences.

16. The apparatus of claim 15 , wherein the audio adjustment manager configured to adjust the audio content based at least in part on the one or more preferences is further configured to:

determine multiple individual audio content adjustments based on the one or more preferences for the one or more identified persons; and

determine a combined audio content adjustment based on the multiple individual audio content adjustments, wherein the combined audio content adjustment is generated based on each of the multiple individual audio content adjustments.

17. The apparatus of claim 14 , wherein the one or more identification modules configured to identify the one or more persons are further configured to:

capture an image proximate to the content consumption device;

perform facial detection on the image to identify faces of the one or more persons in the image; and

visually identify the one or more persons from the image based on the facial detection.

18. The apparatus of claim 14 , wherein the one or more identification modules configured to identify the one or more persons are further configured to:

capture audio proximate to the content consumption device; and

perform voice recognition of the captured audio to identify voices falling in a human vocal range for the one or more persons; and

wherein the audio adjustment manager configured to adjust the audio content is further configured to:

lower the output volume of the audio content in response to the identified voices.

19. At least one non-transitory computer-readable storage medium, comprising a plurality of instructions, which when executed by at least one processor of a content consumption device cause the at least one processor to:

identify one or more persons consuming media content received from a network device via a network;

determine a physical location of the one or more persons relative to a location of the content consumption device;

determine when one of the one or more persons speaks an input associated with a voice command to control actions of the content consumption device; and

automatically adjust the audio content via the content consumption device based least in part on identification of the one or more persons,

wherein the instructions that cause the at least one processor to automatically adjust the audio content further cause the at least one processor to:

lower an output volume of the audio content to facilitate capture, by a microphone associated with the content consumption device, of the input over the audio content and accurate recognition of the voice command, and

modify, based on the physical location of the one or more persons relative to the location of the content consumption device, equalizer settings, balance settings, or fader settings for the audio content.

Assignments (3)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Oct 20, 2015
From: HOU, JUNHUA
To: INTEL CORPORATION
Reel/Frame 036834/0676 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Oct 20, 2015
From: INTEL CORPORATION
To: MCI COMMUNICATIONS SERVICES, INC.
Reel/Frame 036834/0693 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Oct 20, 2015
From: MCI COMMUNICATIONS SERVICES, INC.
To: VERIZON PATENT AND LICENSING INC.
Reel/Frame 036834/0709 →
Continuity (1)
Related Publication 20160073153A1 · Mar 10, 2016