IP Library Granted Patent US 12,412,420
Granted Patent B2
US 12,412,420 · App. 17/358,894 · Granted Sep 9, 2025

Apparatus, systems, and methods for microphone gain control for electronic user devices

Inventor: Tigi Thomas (Bangalore, IN)
Assignee: Intel Corporation
G06V40/168G06F3/005G06F3/013G10L21/0356
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,412,420
App. No.
17/358,894
Granted
Sep 9, 2025
Kind
B2
Abstract

Apparatus, systems, and methods for microphone gain control for electronic user devices are disclosed. An example apparatus includes at least one memory, instructions in the apparatus, and processor circuitry to execute the instructions to determine a proximity of a user to a user device based on an image generated via one or more cameras associated with the user device, determine an amount of gain to be applied to an audio signal generated by a microphone associated with the user device based on the proximity of the user, and cause an amplifier of the user device to apply the gain to the audio signal.

Claims (83)

1. An apparatus comprising:

at least one memory;

machine-readable instructions; and

at least one processor circuit to execute the machine-readable instructions to:

determine a first proximity value representing a first distance of a user to a user device based on a first image generated via one or more cameras associated with the user device, the first image corresponding to a first time;

determine a second proximity value representing a second distance of the user to the user device based on a second image generated via the one or more cameras, the second image corresponding to a second time after the first time;

apply a first filtering parameter to the second proximity value to generate a weighted second proximity value;

apply a second filtering parameter to the first proximity value to generate a weighted first proximity value, a value of the second filtering parameter adjusted based on the first filtering parameter;

determine a filtered proximity value corresponding to the second time based on (a) the weighted first proximity value and (b) the weighted second proximity value;

determine an amount of gain to be applied to an audio signal generated by a microphone associated with the user device based on the filtered proximity value; and

cause an amplifier of the user device to apply the gain to the audio signal.

2. The apparatus of claim 1 , wherein one or more of the at least one processor circuit is to:

detect a face of the user in the first image;

identify a facial landmark associated with the face in the first image; and

determine the first proximity value based on the facial landmark.

3. The apparatus of claim 2 , wherein the facial landmark includes a first facial landmark and a second facial landmark and one or more of the at least one processor circuit is to:

determine a distance between the first facial landmark and the second facial landmark based on first pixel coordinate data associated with the first facial landmark and second pixel coordinate data associated with the second facial landmark; and

determine the first proximity value based on the distance between the first facial landmark and the second facial landmark.

4. The apparatus of claim 3 , wherein the first facial landmark includes a first eye point associated with a first eye of the user and the second facial landmark includes a second eye point associated with the second facial landmark.

5. The apparatus of claim 2 , wherein the face is a first face, the user is a first user, and one or more of the at least one processor circuit is to:

detect a second face in the first image, the second face associated with a second user;

generate a first bounding box for the first face and a second bounding box for the second face;

perform a comparison of a first area of the first bounding box to a second area of the second bounding box; and

select the first face for determining the first proximity value based on the comparison.

6. The apparatus of claim 5 , wherein one or more of the at least one processor circuit is to:

detect a third face in the first image, the third face associated with a third user;

generate a third bounding box for the third face;

assign a first confidence level to the first bounding box, a second confidence level to the second bounding box, and a third confidence level to the third bounding box; and

refrain from selecting the third face for determining the first proximity value based on the third confidence level.

7. The apparatus of claim 1 , wherein one or more of the at least one processor circuit is to:

determine a signal-to-noise ratio of the audio signal;

perform a comparison of the signal-to-noise ratio to a threshold; and

determine the amount of the gain to be applied to the audio signal based on the comparison.

8. The apparatus of claim 1 , wherein one or more of the at least one processor circuit is to cause the amplifier to apply the gain to a voice band of the audio signal.

9. The apparatus of claim 8 , wherein one or more of the at least one processor circuit is to:

filter the audio signal to (a) permit first frequencies associated with the voice band and (b) remove second frequencies associated with ambient noise; and

cause the amplifier to apply the gain to the filtered audio signal.

10. An apparatus comprising:

means for detecting a face of a user in a first video frame corresponding to a first time and in a second video frame corresponding to a second time after the first time, the first video frame and the second video frame associated with a video stream generated by a camera, the camera associated with a user device;

means for identifying first facial landmarks in the face in the first video frame and second facial landmarks in the face in the second video frame;

means for determining proximity to:

determine a first proximity value representing a first distance of the user relative to the camera based on the first facial landmarks;

determine a second proximity value representing a second distance of the user relative to the camera based on the second facial landmarks;

apply a first filtering parameter to the second proximity value to generate a weighted second proximity value;

apply a second filtering parameter to the first proximity value to generate a weighted first proximity value, a value of the second filtering parameter adjusted based on the first filtering parameter; and

determine a filtered proximity value corresponding to the second time based on (a) the weighted first proximity value and (b) the weighted second proximity value;

means for determining a gain to be applied to an audio signal based on the filtered proximity value, the audio signal to be generated by a microphone associated with the user;

means for filtering the audio signal to (a) permit first frequencies associated with a voice band and (b) remove second frequencies associated with ambient noise; and

means for controlling to cause an amplifier of the user device to apply the gain to the filtered audio signal.

11. The apparatus of claim 10 , wherein the first facial landmarks include a third facial landmark and a fourth facial landmark, the facial landmark identifying means is to determine a distance between the third facial landmark and the fourth facial landmark, the proximity determining means to determine the first proximity value based on the distance.

12. The apparatus of claim 10 , wherein the face detecting means is to generate a bounding box in response to detecting the face.

13. The apparatus of claim 10 , further including means for calculating a signal-to-noise ratio of the audio signal, the gain determining means to determine the gain to be applied based on the signal-to-noise ratio.

14. At least one non-transitory computer readable storage medium comprising machine-readable instructions to cause at least one processor circuit to at least:

determine a first proximity value representing a first distance of a user to a user device based on a first image generated via one or more cameras, the first image associated with a first time;

determine a second proximity value representing a second distance of the user to the user device based on a second image generated via the one or more cameras, the second image associated with a second time, the second time after the first time;

apply a first filtering parameter to the second proximity value to generate a weighted second proximity value;

apply a second filtering parameter to the first proximity value to generate a weighted first proximity value, a value of the second filtering parameter adjusted based on the first filtering parameter;

determine a filtered proximity value corresponding to the second time based on the weighted first proximity value and the weighted second proximity value;

determine an amount of gain to be applied to an audio signal generated by a microphone associated with the user device based on the filtered proximity value;

filter the audio signal to (a) permit first frequencies associated with a voice band and (b) remove second frequencies associated with ambient noise; and

cause an amplifier of the user device to apply the gain to the filtered audio signal.

15. The at least one non-transitory computer readable storage medium of claim 14 , wherein the machine-readable instructions are to cause one or more of the at least one processor circuit to:

detect a face of the user in the first image;

identify a facial landmark associated with the face in the first image; and

determine the first proximity value of the user based on the facial landmark.

16. The at least one non-transitory computer readable storage medium of claim 15 , wherein the facial landmark includes a first facial landmark and a second facial landmark and the machine-readable instructions are to cause one or more of the at least one processor circuit to:

determine a distance between the first facial landmark and the second facial landmark based on first pixel coordinate data associated with the first facial landmark and second pixel coordinate data associated with the second facial landmark; and

determine the first proximity value of the user based on the distance between the first facial landmark and the second facial landmark.

17. The at least one non-transitory computer readable storage medium of claim 16 , wherein the first facial landmark includes a first eye point associated with a first eye of the user and the second facial landmark includes a second eye point associated with the second facial landmark.

18. The at least one non-transitory computer readable storage medium of claim 15 , wherein the face is a first face, the user is a first user, and the machine-readable instructions are to cause one or more of the at least one processor circuit to:

detect a second face in the first image, the second face associated with a second user;

generate a first bounding box for the first face and a second bounding box for the second face;

perform a comparison of a first area of the first bounding box to a second area of the second bounding box; and

select the first face for determining the first proximity value based on the comparison.

19. The at least one non-transitory computer readable storage medium of claim 18 , wherein the machine-readable instructions are to cause one or more of the at least one processor circuit to:

detect a third face in the first image, the third face associated with a third user;

generate a third bounding box for the third face;

assign a first confidence level to the first bounding box, a second confidence level to the second bounding box, and a third confidence level to the third bounding box; and

refrain from selecting the third face for determining the first proximity value based on the third confidence level.

20. The at least one non-transitory computer readable storage medium of claim 14 , wherein the machine-readable instructions are to cause one or more of the at least one processor circuit to:

determine a signal-to-noise ratio of the audio signal;

perform a comparison of the signal-to-noise ratio to a threshold; and

determine the amount of the gain to be applied to the filtered audio signal based on the comparison.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jul 14, 2021
From: THOMAS, TIGI
To: INTEL CORPORATION
Reel/Frame 056855/0398 →
Continuity (1)
Related Publication 20210318850A1 · Oct 14, 2021
References Cited (72)
US 4811404A · Vilmur · 1989 [cited by examiner]
US 6363344B1 · Higuchi · 2002 [cited by examiner]
US 6593956B1 · Potts · 2003 [cited by applicant]
US 8681203B1 · Yin et al. · 2014 [cited by applicant]
US 9197974B1 · Clark · 2015 [cited by examiner]
US 9298974B1 · Kuo · 2016 [cited by applicant]
US 10027883B1 · Kuo · 2018 [cited by examiner]
US 10027888B1 · Mackraz · 2018 [cited by examiner]
US 10440324B1 · Lichtenberg · 2019 [cited by applicant]
US 10819950B1 · Lichtenberg et al. · 2020 [cited by applicant]
US 10820795B1 · Weise · 2020 [cited by examiner]
US 11252374B1 · Lichtenberg et al. · 2022 [cited by applicant]
US 20030035001A1 · Van Geest · 2003 [cited by applicant]
US 20080002948A1 · Murata · 2008 [cited by applicant]
US 20080136895A1 · Mareachen · 2008 [cited by applicant]
US 20080260131A1 · Akesson · 2008 [cited by applicant]
US 20080279425A1 · Tang · 2008 [cited by examiner]
US 20090015658A1 · Enstad · 2009 [cited by applicant]
US 20100123785A1 · Chen · 2010 [cited by applicant]
US 20110161074A1 · Pance et al. · 2011 [cited by applicant]
US 20110316996A1 · Abe · 2011 [cited by applicant]
US 20120150333A1 · De Luca · 2012 [cited by applicant]
US 20120183161A1 · Agevik · 2012 [cited by applicant]
US 20130156209A1 · Visser · 2013 [cited by examiner]
US 20130202130A1 · Zurek · 2013 [cited by examiner]
US 20140169627A1 · Gupta · 2014 [cited by examiner]
US 20140337016A1 · Herbig · 2014 [cited by examiner]
US 20140341547A1 · Shenoy · 2014 [cited by applicant]
US 20150131803A1 · Weksler et al. · 2015 [cited by applicant]
US 20150264299A1 · Leech · 2015 [cited by examiner]
US 20150296319A1 · Shenoy · 2015 [cited by examiner]
US 20160073054A1 · Balasaygun et al. · 2016 [cited by applicant]
US 20160195856A1 · Spero · 2016 [cited by examiner]
US 20170127035A1 · Kon et al. · 2017 [cited by applicant]
US 20170186441A1 · Wenus · 2017 [cited by applicant]
US 20170201825A1 · Whyte · 2017 [cited by applicant]
US 20180084365A1 · Ugur · 2018 [cited by applicant]
US 20180144746A1 · Mishra et al. · 2018 [cited by applicant]
US 20180310114A1 · Eronen · 2018 [cited by applicant]
US 20190172462A1 · Mishra · 2019 [cited by applicant]
US 20190313014A1 · Welbourne · 2019 [cited by examiner]
US 20190394606A1 · Tammi · 2019 [cited by applicant]
US 20200028884A1 · Childers et al. · 2020 [cited by applicant]
US 20200110572A1 · Lenke et al. · 2020 [cited by applicant]
US 20200137489A1 · Sheaffer · 2020 [cited by applicant]
US 20200145617A1 · Zavesky et al. · 2020 [cited by applicant]
US 20200184987A1 · Kupryjanow et al. · 2020 [cited by applicant]
US 20200314483A1 · Rakshit et al. · 2020 [cited by applicant]
US 20200393159A1 · Takayanagi · 2020 [cited by applicant]
US 20210092517A1 · Kulkarni et al. · 2021 [cited by applicant]
US 20210116983A1 · Hallberg et al. · 2021 [cited by applicant]
US 20210321212A1 · Li · 2021 [cited by applicant]
US 20220059071A1 · Pearce et al. · 2022 [cited by applicant]
US 20220060525A1 · Chavez et al. · 2022 [cited by applicant]
US 20220159403A1 · Sporer · 2022 [cited by applicant]
US 20220239848A1 · Swierk et al. · 2022 [cited by applicant]
US 20220295012A1 · Swierk et al. · 2022 [cited by applicant]
US 20220319063A1 · Mizobuchi et al. · 2022 [cited by applicant]
US 20230103060A1 · Chaudhuri · 2023 [cited by applicant]
CN 111754990A · 2020 [cited by applicant]
WO 2011123757A1 · 2011 [cited by applicant]
WO 2022067293A1 · 2022 [cited by applicant]
Feng et al., “Face Detection, Bounding Box Aggregation and Pose Estimation for Robust Facial Landmark Localisation in the Wild,” 2017 IEEE Conference on Computer Vision and Pattern Recognition Workshops (CVPRW), 2017, p… [cited by applicant]
United States Patent and Trademark Office, “Final Office Action,” issued in connection with U.S. Appl. No. 17/483,552, dated Feb. 13, 2025, 12 pages. [cited by applicant]
European Patent Office, “Communication pursuant to Article 94(3) EPC,” issued in connection with European Patent Application No. 22191774.3, Mar. 10, 2025, 7 pages. [cited by applicant]
“A Method and Apparatus for Smart Auto-Mute of Conference Participants,” IP.com Prior Art Database Technical Disclosure, May 1, 2014, 12 pages. [cited by applicant]
European Patent Office, “European Search Report,” issued in connection with European patent application No. 22191774.3, Feb. 20, 2023, 8 pages. [cited by applicant]
United States Patent and Trademark Office, “Non-Final Office Action,” issued in connection with U.S. Appl. No. 17/483,552, dated Oct. 25, 2024, 9 pages. [cited by applicant]
Warusfel, “Listen HRTF Database,” dated Sep. 6, 2002, updated May 25, 2003, retrieved from http://recherche.ircam.fr/equipes/salles/listen/, 1 page. [cited by applicant]
United States Patent and Trademark Office, “Non-Final Office Action,” issued in connection with U.S. Appl. No. 17/475,235, dated Oct. 30, 2024, 15 pages. [cited by applicant]
United States Patent and Trademark Office, “Final Office Action,” issued in connection with U.S. Appl. No. 17/475,235, dated Feb. 11, 2025, 19 pages. [cited by applicant]
United States Patent and Trademark Office, 0“Notice of Allowance and Fee(s) Due,” issued in connection with U.S. Appl. No. 17/483,552, dated Jul. 1, 2025, 9 pages. [cited by applicant]