IP Library Granted Patent US 12,651,590
Granted Patent B2
US 12,651,590 · App. 18/346,085 · Granted Jun 9, 2026

Headphone speech listening based on ambient noise

Inventors: Yang Lu (San Jose, CA); Carlos M. Avendano (Campbell, CA); Tony S. Verma (San Francisco, CA)
Assignee: Apple Inc.
G10K11/17854G10K11/17823G10K11/17827G10K11/17879G10K11/17881H04R1/1083G10L2021/02166H04R2460/13
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,651,590
App. No.
18/346,085
Granted
Jun 9, 2026
Kind
B2
Abstract

Microphone signals of a primary headphone are processed and either a first transparency mode of operation is activated or a second transparency mode of operation. In another aspect, a processor enters different configurations in response to estimated ambient acoustic noise being lower or higher than a threshold, wherein in a first configuration a transparency audio signal is adapted via target voice and wearer voice processing (TVWVP) of a microphone signal to boost detected speech frequencies in the transparency audio signal, and in a second configuration the TVWVP is controlled to, as the estimated ambient acoustic noise increases, reduce boosting of, or not boost at all, the detected speech frequencies in the transparency audio signal. Other aspects are also described and claimed.

Claims (45)

1 . A method for digital audio processing, the method comprising:

estimating ambient acoustic noise;

entering a first configuration in response to the estimated ambient acoustic noise being lower than a first threshold, wherein in the first configuration a transparency audio signal is adapted via target voice and wearer voice processing (TVWVP) of a first microphone signal to boost detected speech frequencies in the transparency audio signal;

entering a second configuration in response to the estimated ambient acoustic noise being higher than the first threshold, wherein in the second configuration the TVWVP is controlled to, as the estimated ambient acoustic noise increases, reduce boosting of, or not boost at all, the detected speech frequencies in the transparency audio signal; and

entering a third configuration in response to the estimated ambient acoustic noise being lower than a second threshold, the second threshold being lower than the first threshold, wherein in the third configuration the TVWVP is controlled to, as the estimated ambient acoustic noise decreases, reduce boosting of, or not boost at all, the detected speech frequencies in the transparency audio signal.

2 . The method of claim 1 wherein in the second configuration, and not in the first configuration, producing the transparency audio signal comprises sound pickup beamforming of the first microphone signal and a second microphone signal.

3 . The method of claim 1 further comprising

producing the transparency audio signal by processing the first microphone signal;

producing an anti-noise signal by processing the first microphone signal using feedforward acoustic noise cancellation; and

combining a weighted version of the transparency audio signal with a weighted version of the anti-noise signal, to drive a speaker.

4 . The method of claim 3 further comprising

entering a third configuration in response to the estimated ambient acoustic noise being lower than a second threshold, the second threshold being lower than the first threshold, wherein in the third configuration the TVWVP is controlled to, as the estimated ambient acoustic noise decreases, reduce boosting of, or not boost at all, the detected speech frequencies in the transparency audio signal.

5 . The method of claim 3 further comprising producing the weighted version of the transparency audio signal and the weighted version of the anti-noise signal, by producing a transparency gain vector that flattens a gain frequency response experienced in an ear canal of a wearer of a headphone in which the first microphone signal is generated.

6 . A digital audio processor comprising:

a transparency digital filter path through which a microphone signal is to be filtered to produce a transparency audio signal;

a feedforward acoustic noise cancellation digital filter path through which the microphone signal is filtered to produce an anti-noise signal; and

the digital audio processor being configured to:

combine a weighted version of the transparency audio signal with a weighted version of the anti-noise signal, to drive a speaker;

estimate ambient acoustic noise using the microphone signal;

enter a first configuration in response to the estimated ambient acoustic noise being lower than a first threshold, wherein in the first configuration the processor adapts the transparency digital filter path via target voice and wearer voice processing (TVWVP) of at least the microphone signal to boost detected speech frequencies in the transparency audio signal; and

enter a second configuration in response to the estimated ambient acoustic noise being higher than the first threshold, wherein in the second configuration the processor controls the TVWVP to, as the estimated ambient acoustic noise increases, reduce boosting of, or not boost at all, the detected speech frequencies in the transparency audio signal.

7 . The processor of claim 6 wherein in the second configuration, and not in the first configuration, the processor is configured to produce the transparency audio signal by sound pickup beamforming of the microphone signal and another microphone signal.

8 . The processor of claim 7 further configured to enter a third configuration in response to the estimated ambient acoustic noise being lower than a second threshold, the second threshold being lower than the first threshold, wherein in the third configuration the TVWVP is controlled to, as the estimated ambient acoustic noise decreases, reduce boosting of, or not boost at all, the detected speech frequencies in the transparency audio signal.

9 . The processor of claim 8 further configured to produce the weighted version of the transparency audio signal and the weighted version of the anti-noise signal, by producing a transparency gain vector that flattens a gain frequency response experienced in an ear canal of a wearer.

10 . The processor of claim 6 further configured to enter a third configuration in response to the estimated ambient acoustic noise being lower than a second threshold, the second threshold being lower than the first threshold, wherein in the third configuration the TVWVP is controlled to, as the estimated ambient acoustic noise decreases, reduce boosting of, or not boost at all, the detected speech frequencies in the transparency audio signal.

11 . The processor of claim 10 further configured to produce the weighted version of the transparency audio signal and the weight version of the anti-noise signal, by producing a transparency gain vector that flattens a gain frequency response experienced in an ear canal of a headphone wearer.

12 . The processor of claim 6 configured to perform the TVWVP by:

processing a bone conduction sensor signal and the microphone signal to separate effects of wearer voice from effects of target voice.

13 . An article of manufacture comprising a memory having stored therein instructions that configure a processor to:

estimate ambient acoustic noise;

enter a first configuration in response to the estimated ambient acoustic noise being lower than a first threshold, wherein in the first configuration a transparency audio signal is adapted via target voice and wearer voice processing (TVWVP) of a first microphone signal to boost detected speech frequencies in the transparency audio signal; and

enter a second configuration in response to the estimated ambient acoustic noise being higher than the first threshold, wherein in the second configuration the TVWVP is controlled to, as the estimated ambient acoustic noise increases, reduce boosting of, or not boost at all, the detected speech frequencies in the transparency audio signal; and

enter a third configuration in response to the estimated ambient acoustic noise being lower than a second threshold, the second threshold being lower than the first threshold, wherein in the third configuration the TVWVP is controlled to, as the estimated ambient acoustic noise decreases, reduce boosting of, or not boost at all, the detected speech frequencies in the transparency audio signal.

14 . The article of manufacture of claim 13 wherein in the second configuration, and not in the first configuration, producing the transparency audio signal comprises sound pickup beamforming of the first microphone signal and a second microphone signal.

15 . The article of manufacture of claim 13 wherein the instructions configure the processor to:

produce the transparency audio signal by processing the first microphone signal;

produce an anti-noise signal by processing the first microphone signal using feedforward acoustic noise cancellation; and

combine a weighted version of the transparency audio signal with a weighted version of the anti-noise signal, to drive a speaker.

16 . The article of manufacture of claim 15 wherein the instructions configure the processor to produce the weighted version of the transparency audio signal and the weighted version of the anti-noise signal, by producing a transparency gain vector that flattens a gain frequency response experienced in an ear canal of a wearer of a headphone in which the first microphone signal is generated.

17 . The article of manufacture of claim 13 wherein the instructions configure the processor to:

produce the transparency audio signal by processing the first microphone signal;

produce an anti-noise signal by processing the first microphone signal using feedforward acoustic noise cancellation; and

produce a weighted version of the transparency audio signal and a weighted version of the anti-noise signal by producing a transparency gain vector that flattens a gain frequency response experienced in an ear canal of a wearer of a headphone in which the first microphone signal is generated.

18 . The article of manufacture of claim 17 wherein the instructions configure the processor to:

enter a third configuration in response to the estimated ambient acoustic noise being lower than a second threshold, the second threshold being lower than the first threshold, wherein in the third configuration the TVWVP is controlled to, as the estimated ambient acoustic noise decreases, reduce boosting of, or not boost at all, the detected speech frequencies in the transparency audio signal.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jul 12, 2023
From: LU, YANG; AVENDANO, CARLOS M; VERMA, TONY S
To: APPLE INC.
Reel/Frame 064230/0261 →
Continuity (2)
Provisional Application 63357475 · Jun 30, 2022
Related Publication 20240005903A1 · Jan 4, 2024
References Cited (20)
US 8638961B2 · Pontoppidan · 2014 [cited by applicant]
US 9633671B2 · Giacobello et al. · 2017 [cited by applicant]
US 10034092B1 · Nawfal et al. · 2018 [cited by applicant]
US 11227623B1 · Vitt · 2022 [cited by examiner]
US 11276384B2 · Woodruff et al. · 2022 [cited by applicant]
US 11393486B1 · Woodruff et al. · 2022 [cited by applicant]
US 20100183164A1 · Elmedyb et al. · 2010 [cited by applicant]
US 20150264469A1 · Murata et al. · 2015 [cited by applicant]
US 20170194020A1 · Miller et al. · 2017 [cited by applicant]
US 20200380945A1 · Woodruff · 2020 [cited by examiner]
US 20210250683A1 · Han · 2021 [cited by examiner]
US 20210390972A1 · Lu et al. · 2021 [cited by applicant]
US 20210397407A1 · Eubank et al. · 2021 [cited by applicant]
US 20220270629A1 · Avendano et al. · 2022 [cited by applicant]
EP 3412036A1 · 2018 [cited by applicant]
WO 2018141464A1 · 2018 [cited by applicant]
Wang et al., “Supervised Speech Separation Based on Deep Learning: An Overview”, received from https://ieeexplore.ieee.org/document/8369155, May 30, 2018, 27 pages. [cited by applicant]
Swider et al., “Compatison of Delayless Digital Filtering Algorithms and Their Application to Multi-Sensor Signal Processing”, received from https://journals.sagepub.com/doi/10.1177/0142331218799148, Jun. 29, 2018, 16 p… [cited by applicant]
Koutrouvelis et al., “Binaural Speech Enhancement with Spatial Cue Preservation Utilising Simultaneous Masking” received from https://www.eurasip.org/Proceedings/Eusipco/Eusipco2017/papers/1570347090.pdf, 2017, 5 pages. [cited by applicant]
Dr. Harry Levitt, “Noise Reduction in Hearing Aids: A Review”, received from https://www.rehab.research.va.gov/jour/01/38/1/pdf/levitt.pdf, 2001, 12 pages. [cited by applicant]