IP Library Granted Patent US 12,418,756
Granted Patent B2
US 12,418,756 · App. 18/097,154 · Granted Sep 16, 2025

System and method for enhancing speech of target speaker from audio signal in an ear-worn device using voice signatures

Inventors: Igor Lovchinsky (New York, NY); Andrew J. Casper (Inver Grove Heights, MN); Nicholas Morris (Brooklyn, NY); Matthew de Jonge (Brooklyn, NY); Jonathan Macoskey (Pittsburgh, PA)
Assignee: Chromatic Inc.
H04R25/507G06F3/165G10L17/18G10L21/0272G10L21/028G10L21/0316G10L21/0364G10L25/78H04R25/558H04R25/604H04R25/609H04R2225/43H04R2225/55H04R2430/01
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,418,756
App. No.
18/097,154
Granted
Sep 16, 2025
Kind
B2
Abstract

An ear-worn device is provided that operates to isolate and individually treat the received speech of a target speaker or multiple target speakers from an audio input signal detected in a multi-speaker environment. The ear-worn device uses a machine learning model that receives a voice signature of each of one or more target speakers as input signals, to identify and isolate the component of the audio input signal attributable to the target speaker(s). Once isolated, the target speaker's speech may be enhanced, de-emphasized, or otherwise processed in a manner desired by the wearer of the ear-worn device. The wearer may use an external electronic device, e.g., a phone, to select one or more target speakers in a conversation and/or configure various settings associated with processing the speech on the ear-worn device.

Claims (31)

1. A method of selectively processing, with an ear-worn device including a processor and a microphone coupled to the processor, a target speaker's speech from an audio signal comprising temporally overlapping speech components from multiple speakers, the method comprising:

detecting the audio signal with the microphone of the ear-worn device;

providing the audio signal detected by the microphone of the ear-worn device to the processor of the ear-worn device; and

isolating, with the processor of the ear-worn device, a component of the audio signal representing the target speaker's speech from among the temporally overlapping speech components from multiple speakers by:

providing the audio signal as a first input to a machine learning model and providing a voice signature of the target speaker as a second input to the machine learning model; and

processing the audio signal with the machine learning model using the voice signature of the target speaker.

2. The method of claim 1 , wherein the ear-worn device is a hearing aid or a headphone.

3. The method of claim 1 , wherein the voice signature of the target speaker is a feature vector obtained from another machine learning model trained to discriminate between voices of speakers, and wherein providing the voice signature as the second input to the machine learning model comprises providing the feature vector as an input signal to the machine learning model.

4. The method of claim 1 , wherein the audio signal is a second audio signal, and wherein the method further comprises:

detecting a first audio signal with the microphone of the ear-worn device; and

extracting the voice signature of the target speaker from the first audio signal before processing the second audio signal with the machine learning model using the voice signature of the target speaker.

5. The method of claim 1 , wherein processing the audio signal with the machine learning model comprises processing the audio signal with a neural network.

6. The method of claim 5 , wherein processing the audio signal with a neural network comprises processing the audio signal with a recurrent neural network or a convolutional neural network.

7. The method of claim 1 , wherein the target speaker comprises a wearer of the ear-worn device, and wherein the method further comprises, after isolating the component of the audio signal representing the target speaker's speech from among the temporally overlapping speech components, suppressing the component of the audio signal representing the target speaker's speech.

8. The method of claim 1 , further comprising amplifying the component of the audio signal representing the target speaker's speech after isolating the component of the audio signal representing the target speaker's speech.

9. The method of claim 1 , wherein the temporally overlapping speech components from multiple speakers comprises the component of the audio signal representing the target speaker's speech and a component of the audio signal representing a secondary speaker's speech, and wherein the method further comprises suppressing the component of the audio signal representing the secondary speaker's speech after isolating the component of the audio signal representing the target speaker's speech.

10. The method of claim 1 , wherein the ear-worn device further comprises a speaker device, and wherein the method further comprises outputting from the speaker device of the ear-worn device the component of the audio signal representing the target speaker's speech after isolating the component of the audio signal representing the target speaker's speech from among the temporally overlapping speech components from multiple speakers.

11. An ear-worn device comprising:

one or more microphones configured to detect an audio signal including speech from one or more speakers in a conversation;

at least one processor configured to perform the method of claim 1 to generate an output audio signal; and

one or more speaker devices configured to output the output audio signal.

12. A system, comprising:

the ear-worn device of claim 11 ; and

a second ear-worn device comprising:

one or more second microphones configured to detect a second audio signal;

at least one second processor configured to:

generate a second output audio signal at least in part by:

isolating a component of the second audio signal representing the target speaker's speech from among the temporally overlapping speech components from multiple speakers by processing the second audio signal with a second machine learning model;

receive, from the ear-worn device, the output audio signal; and

combine the output audio signal and the second output audio signal to generate a combined audio signal; and

one or more second speaker devices configured to output the combined audio signal.

Assignments (3)
CHANGE OF NAME Recorded Oct 9, 2025
From: CHROMATIC INC.
To: FORTELL RESEARCH INC.
Reel/Frame 073057/0966 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Aug 31, 2023
From: LOVCHINSKY, IGOR; CASPER, ANDREW J.; MORRIS, NICHOLAS; DE JONGE, MATTHEW
To: CHROMATIC INC.
Reel/Frame 064766/0204 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Mar 7, 2023
From: MACOSKEY, JONATHAN
To: CHROMATIC INC.
Reel/Frame 062901/0080 →
Continuity (7)
Continuation In Part 17576718 · Jan 14, 2022
Continuation In Part 17576746 · Jan 14, 2022
Continuation In Part 17576893 · Jan 14, 2022
Continuation In Part 17576899 · Jan 14, 2022
Continuation In Part PCTUS2022012567 · Jan 14, 2022
Provisional Application 63305676 · Feb 1, 2022
Related Publication 20230306982A1 · Sep 28, 2023
References Cited (82)
US 7804973B2 · De Vries et al. · 2010 [cited by applicant]
US 9716939B2 · Censo et al. · 2017 [cited by applicant]
US 9881631B2 · Erdogan et al. · 2018 [cited by applicant]
US 10199047B1 · Clark · 2019 [cited by applicant]
US 10516934B1 · Solbach · 2019 [cited by applicant]
US 10536775B1 · Sen et al. · 2020 [cited by applicant]
US 10659893B2 · Pedersen et al. · 2020 [cited by applicant]
US 10721571B2 · Crow et al. · 2020 [cited by applicant]
US 10805748B2 · Fichtl et al. · 2020 [cited by applicant]
US 10812915B2 · Santos et al. · 2020 [cited by applicant]
US 10957301B2 · Hoby et al. · 2021 [cited by applicant]
US 11245993B2 · Andersen et al. · 2022 [cited by applicant]
US 11270198B2 · Ah · 2022 [cited by applicant]
US 11330378B1 · Jelcicová et al. · 2022 [cited by applicant]
US 11375325B2 · Froehlich et al. · 2022 [cited by applicant]
US 11445307B2 · Pandey et al. · 2022 [cited by applicant]
US 11553286B2 · Sabin et al. · 2023 [cited by applicant]
US 11620977B2 · Birmingham et al. · 2023 [cited by applicant]
US 11647344B2 · Chen et al. · 2023 [cited by applicant]
US 11678120B2 · Nyayate et al. · 2023 [cited by applicant]
US 11696079B2 · Jelcicova et al. · 2023 [cited by applicant]
US 11812225B2 · Casper et al. · 2023 [cited by applicant]
US 11818523B2 · Lovchinsky et al. · 2023 [cited by applicant]
US 11818547B2 · Casper et al. · 2023 [cited by applicant]
US 11832061B2 · Casper et al. · 2023 [cited by applicant]
US 11877125B2 · Casper et al. · 2024 [cited by applicant]
US 11950056B2 · Casper et al. · 2024 [cited by applicant]
US 12075215B2 · Casper et al. · 2024 [cited by applicant]
US 20070172087A1 · Olsen · 2007 [cited by applicant]
US 20100027820A1 · Kates · 2010 [cited by applicant]
US 20140064529A1 · Jang · 2014 [cited by applicant]
US 20150078575A1 · Selig et al. · 2015 [cited by applicant]
US 20170229117A1 · Van Der Made et al. · 2017 [cited by applicant]
US 20200043499A1 · Basye et al. · 2020 [cited by applicant]
US 20200204928A1 · Fichtl · 2020 [cited by applicant]
US 20210105565A1 · Pedersen et al. · 2021 [cited by applicant]
US 20210274296A1 · Rohde et al. · 2021 [cited by applicant]
US 20210281958A1 · Diehl et al. · 2021 [cited by applicant]
US 20210289299A1 · Durrieu · 2021 [cited by applicant]
US 20220095061A1 · Diehl et al. · 2022 [cited by applicant]
US 20220124444A1 · Andersen et al. · 2022 [cited by applicant]
US 20220159403A1 · Sporer et al. · 2022 [cited by applicant]
US 20220223161A1 · Fuchs et al. · 2022 [cited by applicant]
US 20220230048A1 · Li et al. · 2022 [cited by applicant]
US 20220232321A1 · Wexler et al. · 2022 [cited by applicant]
US 20220232331A1 · Jelcicová et al. · 2022 [cited by applicant]
US 20220256294A1 · Diehl et al. · 2022 [cited by applicant]
US 20230037356A1 · Pontoppidan et al. · 2023 [cited by applicant]
US 20230087486A1 · Pennies-Hochmuth et al. · 2023 [cited by applicant]
US 20230209283A1 · Wagner et al. · 2023 [cited by applicant]
US 20230232169A1 · Casper et al. · 2023 [cited by applicant]
US 20230232170A1 · Casper et al. · 2023 [cited by applicant]
US 20230232171A1 · Casper et al. · 2023 [cited by applicant]
US 20230232172A1 · Casper et al. · 2023 [cited by applicant]
US 20230254650A1 · Lovchinsky et al. · 2023 [cited by applicant]
US 20230254651A1 · Casper et al. · 2023 [cited by applicant]
US 20230292074A1 · Marquardt et al. · 2023 [cited by applicant]
US 20230319492A1 · Corey et al. · 2023 [cited by applicant]
US 20230388725A1 · Casper et al. · 2023 [cited by applicant]
US 20230402055A1 · Kulasekaran et al. · 2023 [cited by applicant]
US 20240048922A1 · Casper et al. · 2024 [cited by applicant]
US 20240056747A1 · Casper et al. · 2024 [cited by applicant]
US 20240129674A1 · Casper et al. · 2024 [cited by applicant]
US 20240194213A1 · Wichern et al. · 2024 [cited by applicant]
US 20240221769A1 · Philipsson et al. · 2024 [cited by applicant]
US 20240292165A1 · Lovchinsky et al. · 2024 [cited by applicant]
US 20240381039A1 · Casper et al. · 2024 [cited by applicant]
CN 105611477A · 2016 [cited by applicant]
EP 0357212A2 · 1990 [cited by applicant]
KR 102316626B1 · 2021 [cited by applicant]
WO WO2020079485A1 · 2020 [cited by applicant]
WO WO2022079848A1 · 2022 [cited by applicant]
WO WO2022107393A1 · 2022 [cited by applicant]
WO WO2022191879A1 · 2022 [cited by applicant]
WO WO2023010014A1 · 2023 [cited by applicant]
WO WO2023110836A1 · 2023 [cited by applicant]
International Search Report and Written Opinion mailed Apr. 28, 2023 in connection with International Application No. PCT/US2023/010837. [cited by applicant]
Giri et al., Personalized Percepnet: Real-time, Low-complexity Target Voice Separation and Enhancement. Amazon Web Service, Jun. 8, 2021, arXiv preprint arXiv:2106.04129. 5 pages. [cited by applicant]
International Search Report and Written Opinion mailed Jun. 16, 2022 in connection with International Application No. PCT/US2022/012567. [cited by applicant]
Gerlach et al., A Survey on Application Specific Processor Architectures for Digital Hearing Aids. Journal of Signal Processing Systems. Mar. 20, 2021;94:1293-1308. https://link.springer.com/rticle/10.1007/s11265-021-01… [cited by applicant]
International Preliminary Report on Patentability mailed Jul. 25, 2024 in connection with International Application No. PCT/US2023/010837. [cited by applicant]
International Preliminary Report on Patentability mailed Jul. 25, 2024 in connection with International Application No. PCT/US2022/012567. [cited by applicant]
Cited By (4)
US 12,574,691 US 12,610,200 US 12,634,642 US 12,713,188