IP Library Granted Patent US 11,818,523
Granted Patent B2
US 11,818,523 · App. 18/137,951 · Granted Nov 14, 2023

System and method for enhancing speech of target speaker from audio signal in an ear-worn device using voice signatures

Inventors: Igor Lovchinsky (New York, NY); Andrew J. Casper (Inver Grove Heights, MN); Nicholas Morris (Brooklyn, NY); Matthew de Jonge (Brooklyn, NY); Jonathan Macoskey (Pittsburgh, PA)
Assignee: Chromatic Inc.
H04R25/507G06F3/165G10L17/18G10L21/028G10L21/0364H04R25/558H04R25/604H04R25/609H04R2225/43H04R2225/55H04R2430/01
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,818,523
App. No.
18/137,951
Granted
Nov 14, 2023
Kind
B2
Abstract

An ear-worn device is provided that operates to isolate and individually treat the received speech of a target speaker or multiple target speakers from an audio input signal detected in a multi-speaker environment. The ear-worn device uses a machine learning model that receives a voice signature of each of one or more target speakers as input signals, to identify and isolate the component of the audio input signal attributable to the target speaker(s). Once isolated, the target speaker's speech may be enhanced, de-emphasized, or otherwise processed in a manner desired by the wearer of the ear-worn device. The wearer may use an external electronic device, e.g., a phone, to select one or more target speakers in a conversation and/or configure various settings associated with processing the speech on the ear-worn device.

Claims (39)

1. A hearing aid system, comprising:

an ear-worn device including:

a microphone configured to receive an audible signal and output an electrical signal representing the audible signal;

front-end circuitry coupled to the microphone and configured to receive the electrical signal representing the audible signal, digitize the electrical signal, and output a digitized version of the audible signal;

a controller configured to:

receive the digitized version of the audible signal; and

selectively output the digitized version of the audible signal to either a digital signal processor (DSP) or a neural network engine comprising a voice isolation and classification neural network;

the neural network engine, wherein the neural network engine is coupled to an output of the controller and configured to:

output from the voice isolation and classification neural network a de-noised version of the digitized version of the audible signal;

output from the voice isolation and classification neural network an embedding of the digitized version of the audible signal;

compare the embedding of the digitized version of the audible signal to a reference embedding;

separate the digitized version of the audible signal into multiple source signals using the voice isolation and classification neural network;

apply gains to multiple source signals of the de-noised version of the digitized version of the audible signal based at least in part on a result of the comparison of the embedding of the digitized version of the audible signal to the reference embedding;

create a combined signal by recombining the multiple source signals after application of the gains; and

provide the combined signal to the DSP;

the DSP, wherein the DSP is coupled to the output of the controller and an output of the neural network engine, and wherein the DSP is configured to, upon receiving the combined signal from the neural network engine, filter the combined signal to generate an output signal; and

a speaker, coupled to an output of the DSP and configured to playback the output signal in audible form; and

an electronic device, comprising a voice isolation and classification neural network, wherein the electronic device is configured to:

receive a speech sample from a target speaker and generate the reference embedding; and

provide the reference embedding to the ear-worn device.

2. The hearing aid system of claim 1 , wherein the voice isolation and classification neural network of the ear-worn device and the voice isolation and classification neural network of the electronic device are of a same type.

3. The hearing aid system of claim 2 , wherein the voice isolation and classification neural network of the ear-worn device is a recurrent neural network.

4. The hearing aid system of claim 2 , wherein the electronic device is configured to provide a graphical user interface (GUI) to a user and to provide the reference embedding to the ear-worn device upon selection by the user of the target speaker on the GUI.

5. The hearing aid system of claim 1 , wherein the electronic device is configured to provide a graphical user interface (GUI) to a user and to provide the reference embedding to the ear-worn device upon selection by the user of the target speaker on the GUI.

6. The hearing aid system of claim 1 , wherein the reference embedding is a first reference embedding, and wherein the electronic device is configured to received multiple speech samples from respective target speakers and to generate respective reference embeddings including the first reference embedding using the voice isolation and classification neural network of the electronic device, wherein the electronic device is further configured to provide the respective reference embeddings to the ear-worn device, and wherein the neural network engine of the ear-worn device is further configured to compare the embedding of the digitized version of the audible signal to the respective reference embeddings.

7. The hearing aid system of claim 6 , wherein the neural network engine of the ear-worn device is configured to compare the embedding of the digitized version of the audible signal to the respective reference embeddings at least in part by calculating a respective cosine similarity between the embedding of the digitized version of the audible signal and a respective reference embedding.

8. The hearing aid system of claim 7 , wherein the neural network engine is configured to apply the gains to multiple source signals of the de-noised version of the digitized version of the audible signal based at least in part on results of the comparisons of the embedding of the digitized version of the audible signal to the respective reference embeddings.

9. The hearing aid system of claim 1 , wherein the reference embedding represents a voice of a wearer of the ear-worn device, and wherein, when the comparison of the embedding of the digitized version of the audible signal to the reference embedding indicates that the wearer's voice is present in the digitized version of the audible signal, the neural network engine is configured to reduce a component of the digitized version of the audible signal representing the wearer's voice.

10. The hearing aid system of claim 1 , wherein the electronic device is a smartphone.

11. The hearing aid system of claim 1 , wherein the neural network engine of the ear-worn device is configured to compare the embedding of the digitized version of the audible signal to the reference embedding at least in part by calculating a cosine similarity between the embedding of the digitized version of the audible signal and the reference embedding.

12. The hearing aid system of claim 1 , wherein the front-end circuitry, controller, neural network engine, and DSP are implemented on a system-on-chip.

13. The hearing aid system of claim 1 , wherein the controller is further configured to determine a heuristic of the digitized version of the audible signal, and wherein the controller is further configured to selectively output the digitized version of the audible signal to either the DSP or the neural network engine depending at least in part on the heuristic of the digitized version of the audible signal.

14. The hearing aid system of claim 1 , wherein the controller is further configured to receive a user-selected mode and to selectively output the digitized version of the audible signal to either the DSP or the neural network engine depending at least in part on the user-selected mode.

15. The hearing aid system of claim 14 , wherein the electronic device is further configured to provide to the ear-worn device the user-selected mode.

16. The hearing aid system of claim 1 , wherein the neural network engine is further configured to receive an indication of a user-selected directionality and to select the gains based at least in part on the user-selected directionality.

17. The hearing aid system of claim 1 , wherein the DSP is configured to apply frequency-dependent filtering including the application of non-linear gains to different frequency bands of the combined signal.

18. The hearing aid system of claim 1 , wherein the controller is configured to provide the digitized version of the audible signal to the neural network engine in segments, and wherein the neural network engine is configured to process a segment of the digitized version of the audible signal in a time less than or equal to a duration of the segment.

19. The hearing aid system of claim 1 , wherein the voice isolation and classification neural network of the ear-worn device is a recurrent neural network.

20. The hearing aid system of claim 1 , wherein the voice isolation and classification neural network of the ear-worn device is a convolutional neural network.

Assignments (2)
CHANGE OF NAME Recorded Oct 9, 2025
From: CHROMATIC INC.
To: FORTELL RESEARCH INC.
Reel/Frame 073057/0966 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Sep 5, 2023
From: LOVCHINSKY, IGOR; CASPER, ANDREW J.; MORRIS, NICHOLAS; DE JONGE, MATTHEW; MACOSKEY, JONATHAN
To: CHROMATIC INC.
Reel/Frame 064797/0566 →
Continuity (8)
Continuation 18097154 · Jan 13, 2023
Continuation In Part 17576899 · Jan 14, 2022
Continuation In Part 17576718 · Jan 14, 2022
Continuation In Part PCTUS2022012567 · Jan 14, 2022
Continuation In Part 17576746 · Jan 14, 2022
Continuation In Part 17576893 · Jan 14, 2022
Provisional Application 63305676 · Feb 1, 2022
Related Publication 20230254650A1 · Aug 10, 2023
Cited By (9)
US 12,356,153 US 12,356,154 US 12,356,156 US 12,363,489 US 12,418,756 US 12,574,691 US 12,610,200 US 12,634,642 US 12,713,188