IP Library Granted Patent US 12,567,434
Granted Patent B2
US 12,567,434 · App. 18/062,778 · Granted Mar 3, 2026

Audio system, audio device, and method for speaker extraction

Inventors: Rasmus Kongsgaard Olsson (Ballerup, DK); Clément Laroche (Ballerup, DK)
Assignee: GN AUDIO A/S
G10L25/78G10L21/0208G10L25/30G10L2021/02082
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,567,434
App. No.
18/062,778
Granted
Mar 3, 2026
Kind
B2
Abstract

A method for speech extraction in an audio device is disclosed. The method comprises obtaining a microphone input signal from one or more microphones including a first microphone. The method comprises applying an extraction model to the microphone input signal for provision of an output. The method comprises extracting a near speaker component in the microphone input signal according to the output of the extraction model being a machine-learning model for provision of a speaker output. The method comprises outputting the speaker output.

Claims (24)

1 . A method for speech extraction in an audio device, the method comprising

obtaining input signal from one or more microphones including a first microphone;

performing a short-time Fourier transformation on the input signal to obtain a frequency domain representation of the input signal;

performing a power normalizing on the frequency domain representation of the input signal to obtain a power normalized input signal;

feeding the power normalized input signal into an extraction model,

wherein the extraction model is a machine-learning model; and

wherein the extraction model is trained based on clean speech signals and a set of reverberant speech signals;

determining one or more mask parameters based on output of the extraction model, wherein the one or more mask parameters comprise first mask parameters and second mask parameters;

extract a near speaker component based on the first mask parameters;

extract a far speaker component based on the second mask parameters;

wherein the far speaker component originates from a far speaker at a distance larger than approximately 30 cm from the first microphone, and the near speaker component originates from a near speaker at a distance less than approximately 30 cm from the first microphone.

2 . The method according to claim 1 , the method further comprising:

determining a near speaker signal based on the near speaker component, and outputting the near speaker signal as a speaker output.

3 . The method according to claim 2 , wherein the method further comprising performing inverse short-time Fourier transformation on the speaker output for provision of an electrical output signal.

4 . The method according to claim 1 , wherein the extracting of the near speaker component in the input signal comprises:

determining one or more mask parameters including a first mask parameter based on the output of the extraction model.

5 . The method according to claim 1 , wherein the machine-learning model is an off-line trained neural network.

6 . The method according to claim 1 , wherein the extraction model comprises deep neural network.

7 . The method according to claim 1 , wherein the obtaining of the input signal comprises performing short-time Fourier transformation on the input signal from one or more microphones for provision of the input signal.

8 . The method according to claim 1 , wherein the method further comprising extracting an ambient noise component in the input signal according to the output of the extraction model.

9 . The method according to claim 1 ,

wherein the obtaining of the input signal from one or more microphones including the first microphone comprises obtaining one or more of a first microphone input signal, a second microphone input signal, and a combined microphone input signal based on the first microphone input signal and second microphone input signal,

wherein the input signal is based on one or more of the first microphone input signal, the second microphone input signal, and the combined microphone input signal.

10 . An audio device comprising a processor, an interface, a memory, and one or more transducers, wherein the audio device is configured to perform the method according to claim 1 .

Assignments (2)
MERGER Recorded Mar 30, 2026
From: GN AUDIO A/S
To: GN HEARING A/S
Reel/Frame 075299/0225 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Dec 7, 2022
From: OLSSON, RASMUS KONGSGAARD; LAROCHE, CLÉMENT
To: GN AUDIO A/S
Reel/Frame 062011/0957 →
Priority Claims (1)
EP 21217565 · Dec 23, 2021 · regional
Continuity (1)
Related Publication 20230206941A1 · Jun 29, 2023
References Cited (84)
US 6032072A · Greenwald · 2000 [cited by examiner]
US 6449593B1 · Valve · 2002 [cited by examiner]
US 7016488B2 · He · 2006 [cited by examiner]
US 8160273B2 · Visser · 2012 [cited by examiner]
US 8620672B2 · Visser · 2013 [cited by examiner]
US 8898056B2 · Chan · 2014 [cited by examiner]
US 9094771B2 · Tsingos · 2015 [cited by examiner]
US 10735887B1 · McElveen · 2020 [cited by examiner]
US 11081788B1 · Hahn, III · 2021 [cited by examiner]
US 11201643B1 · Aljohani · 2021 [cited by examiner]
US 11620983B2 · Zhang · 2023 [cited by examiner]
US 20020048376A1 · Ukita · 2002 [cited by examiner]
US 20030123674A1 · Boland · 2003 [cited by examiner]
US 20040252652A1 · Berestesky · 2004 [cited by examiner]
US 20050060142A1 · Visser et al. · 2005 [cited by applicant]
US 20070021958A1 · Visser et al. · 2007 [cited by applicant]
US 20080208538A1 · Visser · 2008 [cited by examiner]
US 20090175466A1 · Elko · 2009 [cited by examiner]
US 20090299742A1 · Toman · 2009 [cited by examiner]
US 20100324890A1 · Adeney · 2010 [cited by examiner]
US 20110038489A1 · Visser · 2011 [cited by examiner]
US 20110066428A1 · Yang · 2011 [cited by examiner]
US 20110264447A1 · Visser · 2011 [cited by examiner]
US 20120016669A1 · Endo · 2012 [cited by examiner]
US 20120114126A1 · Thiergart · 2012 [cited by examiner]
US 20120237037A1 · Ninan · 2012 [cited by examiner]
US 20130275128A1 · Claussen · 2013 [cited by examiner]
US 20130317783A1 · Tennant · 2013 [cited by examiner]
US 20140064476A1 · Mani · 2014 [cited by examiner]
US 20140278445A1 · Eddington, Jr. · 2014 [cited by examiner]
US 20140337021A1 · Kim · 2014 [cited by examiner]
US 20150112671A1 · Johnston · 2015 [cited by examiner]
US 20150235125A1 · Krishnan · 2015 [cited by examiner]
US 20160193732A1 · Breazeal · 2016 [cited by examiner]
US 20170100550A1 · Van De Laar · 2017 [cited by examiner]
US 20180053512A1 · Cilingir · 2018 [cited by examiner]
US 20180197557A1 · Guo · 2018 [cited by examiner]
US 20180220007A1 · Sun · 2018 [cited by examiner]
US 20190086508A1 · Isberg · 2019 [cited by examiner]
US 20190096419A1 · Giacobello · 2019 [cited by examiner]
US 20190137045A1 · Jalilian · 2019 [cited by examiner]
US 20190172480A1 · Kaskari · 2019 [cited by examiner]
US 20190228791A1 · Sun · 2019 [cited by examiner]
US 20190362711A1 · Nosrati · 2019 [cited by examiner]
US 20190394598A1 · Moore · 2019 [cited by examiner]
US 20200005806A1 · Seo · 2020 [cited by examiner]
US 20200075033A1 · Hijazi et al. · 2020 [cited by applicant]
US 20200105287A1 · Chang · 2020 [cited by examiner]
US 20200162274A1 · Lyer · 2020 [cited by examiner]
US 20200221244A1 · Peeler · 2020 [cited by examiner]
US 20200312345A1 · Fazeli · 2020 [cited by examiner]
US 20200312346A1 · Fazeli · 2020 [cited by examiner]
US 20200388299A1 · Kawai · 2020 [cited by examiner]
US 20210006900A1 · Ohashi · 2021 [cited by examiner]
US 20210042796A1 · Khoury · 2021 [cited by examiner]
US 20210043198A1 · Tanaka · 2021 [cited by examiner]
US 20210116532A1 · Yang · 2021 [cited by examiner]
US 20210204059A1 · Trestain · 2021 [cited by examiner]
US 20220036903A1 · Cilingir · 2022 [cited by examiner]
US 20220086592A1 · McElveen · 2022 [cited by examiner]
US 20220115021A1 · Ukai · 2022 [cited by examiner]
US 20220201421A1 · McElveen · 2022 [cited by examiner]
US 20220279305A1 · Sheaffer · 2022 [cited by examiner]
US 20220301582A1 · Wang · 2022 [cited by examiner]
US 20220319498A1 · Caroselli, Jr. · 2022 [cited by examiner]
US 20220345845A1 · Tsingos · 2022 [cited by examiner]
US 20230005488A1 · Hiroe · 2023 [cited by examiner]
US 20230029048A1 · Hahn, III · 2023 [cited by examiner]
US 20230077621A1 · Ono · 2023 [cited by examiner]
US 20230162757A1 · Zheng · 2023 [cited by examiner]
US 20230190140A1 · Tiron · 2023 [cited by examiner]
US 20230403506A1 · Zhu · 2023 [cited by examiner]
US 20240005942A1 · Liu · 2024 [cited by examiner]
US 20240085935A1 · Yu · 2024 [cited by examiner]
US 20240195916A1 · Schiøler · 2024 [cited by examiner]
US 20240303621A1 · Schäfer · 2024 [cited by examiner]
US 20250071505A1 · McElveen · 2025 [cited by examiner]
The extended European search report issued in European Application No. 21217565.7, dated Jun. 10, 2022. [cited by applicant]
Maciejewski et al., “Analysis of Robustness of Deep Single-Channel Speech Separation Using Corpora Constructed From Multiple Domains”, 2019 IEEE Workshop On Applications of Signal Processing To Audio and Acoustics (WASP… [cited by applicant]
Higuchi et al., “Adversarial training for data-driven speech enhancement without parallel corpus”, 2017 IEEE Automatic Speech Recognition and Understanding Workshop (ASRU), IEEE, Dec. 16, 2017, pp. 40-47. [cited by applicant]
Zhao et al., “Monaural Speech Dereverberation Using Temporal Convolutional Networks With Self Attention”, IEEE/ACM Transactions On Audio, Speech, and Language Processing, IEEE, USA, vol. 28, May 18, 2020, pp. 1598-1607. [cited by applicant]
Hussain et al., “Ensemble Hierarchical Extreme Learning Machine for Speech Dereverberation”, IEEE Transactions On Cognitive and Developmental Systems, IEEE, vol. 12, No. 4, Nov. 19, 2019, pp. 744-758. [cited by applicant]
Szoke et al., Building and Evaluation of a Real Room Impulse Response Dataset, IEEE Journal of Selected Topics in Signal Processing, IEEE, US, vol. 13, No. 4, Aug. 1, 2019, pp. 863-876. [cited by applicant]
Wuth et al., “Non causal deep learning based dereverberation”, Arxiv.org, Cornell University Library, 201 Olin Library Cornell University Ithaca, NY 14853, Sep. 7, 2020. [cited by applicant]