IP Library Granted Patent US 12711977
Granted Patent B2
US 12711977 · App. 18/719,087 · Granted Aug 18, 2026

Method of operating an audio device system and an audio device system

Inventors: Rasmus Malik Hoeegh Lindrup (Berkeley, CA); Jens Brehm Bagger Nielsen (Fredensborg, DK)
Assignee: WIDEX A/S
G10L21/0216G10L15/1815G10L21/0364G10L25/30G10L2021/02166
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12711977
App. No.
18/719,087
Filed
Jun 12, 2024
Granted
Aug 18, 2026
Kind
B2
Art Unit
2653
USPC
704/202
Abstract

A method of operating an audio device system in order to provide at least one of improved noise reduction and speech intelligibility and an audio device system adapted to carry out the method.

Claims (57)

1 . A method of operating an audio device system comprising the steps of:

a) providing a plurality of sound source signals each from a sound source of a present sound environment;

b) selecting a first sound signal comprising speech;

c) comparing speech content of said first sound signal with speech content of the provided plurality of sound source signals,

wherein said comparing comprises:

(i) generating, by a natural language processing model, semantic and/or syntactic numerical representations of the speech content of the first sound signal and of each of the plurality of sound source signals, and

(ii) determining, by a language model, for each of the plurality of sound source signals a probability that the speech content thereof constitutes a conversational response to the speech content of the first sound signal;

d) selecting, based on said comparison, as output signal the sound source signal that most likely is part of a conversation that an audio device system user is paying attention to, wherein said selecting comprises selecting the sound source signal having the highest predicted probability of being such a conversational response and/or having a highest semantic and/or syntactic similarity with the first sound signal; and

e) providing an audio output based on said output signal, wherein the contribution to the audio output from the remaining sound source signals is suppressed compared to the contribution from the output signal.

2 . The method according to claim 1 , wherein the step of providing the plurality of sound source signals each from the sound source of the present sound environment comprises the further steps of:

using an encoder-decoder neural network that has been obtained by feeding a mixed audio signal comprising a plurality of speech signals and a plurality of noise signals to the neural network and subsequently train the neural network to provide only said plurality of speech signals; or

using a plurality of beam formers each adapted to point in a desired direction different from other beam formers.

3 . The method according to claim 2 , wherein the step of using the plurality of beam formers each adapted to point in the desired direction different from the other beam formers comprises the further step of:

determining that a beam former is pointing in a desired direction if speech is detected in the beam former output signal.

4 . The method according to claim 1 , wherein the step of selecting the first sound signal is carried out by at least one of the steps of:

detecting a sound from a source that is positioned directly in front of the user, detecting a sound from a source that the user is looking at, detecting the audio device system user's own voice, by identifying the sound source signal that exhibits the highest similarity with an EEG signal of the user; and by subsequently carrying out at least one of the steps of:

selecting the first sound signal from at least one of said detected sounds in response to a predetermined interaction between the user and the audio device system wherein said predetermined interaction comprises at least one: making a specific head movement, tapping an audio device of the audio device system, operating an audio device control means, speaking a control word, or operating a graphical user interface of the audio device system.

5 . The method according to claim 1 , wherein said step of comparing the speech content of said first sound signal with the speech content of the provided plurality of sound source signals comprises at least one of:

assigning a numerical representation to at least some of the words comprised in the first sound signal and in said plurality of provided sound source signals, and

providing a word embedding similarity measure in order to estimate a similarity between the first sound signal and said plurality of provided sound source signals; and

determining timing of speech ending for the first sound signal and determining timing of speech onset for said plurality of provided sound source signals and subsequently identifying at least one sound source signals with speech onset within a predetermined duration after speech ending for the first sound signal;

assigning a numerical representation to at least one of syntactic and semantic information comprised in the first sound signal and in said plurality of provided sound source signals; or

providing at least one of a syntactic similarity and a semantic similarity between the first sound signal and said plurality of provided sound source signals.

6 . The method according to claim 1 , wherein the step of selecting, based on said comparison, as the output signal the sound source signal that the audio device system user is most likely paying attention to comprises at least one of the steps of:

selecting as the output signal the sound source signal having a word embedding similarity measure that is most similar with a word embedding similarity measure of the first sound source signal;

selecting as the output signal the sound source signal having a speech onset within a predetermined duration after speech ending of the first sound signal;

selecting as the output signal the sound source signal having the highest score of at least one of the semantic similarity measure and the syntactic similarity measure; or

selecting as the output signal the sound source signal having a highest combined score, wherein the combined score is obtained by combining at least some of:

the word embedding similarity measure score, the semantic similarity measure, the syntactic similarity measure, a sound pressure level score reflecting the strength of the signal, a previous participant score reflecting whether the speaker representing the sound source signal has previously participated in the conversation and having a speech onset within said predetermined duration after speech ending of the first sound signal.

7 . The method according to claim 1 , wherein the step of providing the audio output based on said output signal comprises at least one of the steps of:

suppressing the contribution to the audio output from the remaining sound source signals such that the combined level of the remaining sound source signals is in the range between 3 and 24 dB or between 6 and 18 dB below the output signal level; or

enabling the user to control the ratio between the output signal level and the combined level of the remaining sound source signals.

8 . The method according to claim 1 , comprising the further steps of:

replacing the first sound signal with the output signal; and

triggering that steps c), d) and e) are carried out in response to a detection of said output signal reaching a speech ending.

9 . The method according to claim 1 , comprising the further step of:

triggering that steps c), d) and e) are carried out in response to a detection of said output signal reaching a speech ending.

10 . The method according to claim 1 , wherein the first sound signal is associated with a first speaker and a sound source signal of the plurality of sound source signals is associated with a second speaker different from the first speaker.

11 . The method according to claim 1 , wherein said step of comparing the speech content of said first sound signal with the speech content of the provided plurality of sound source signals comprises:

assigning a numerical representation to at least some of the words comprised in the first sound signal and in said plurality of provided sound source signals, and

providing a word embedding similarity measure in order to estimate a similarity between the first sound signal and said plurality of provided sound source signals; and

determining timing of speech ending for the first sound signal and determining timing of speech onset for said plurality of provided sound source signals and subsequently identifying at least one sound source signals with speech onset within a predetermined duration after speech ending for the first sound signal;

assigning a numerical representation to at least one of syntactic and semantic information comprised in the first sound signal and in said plurality of provided sound source signals; and

providing at least one of a syntactic similarity and a semantic similarity between the first sound signal and said plurality of provided sound source signals.

12 . The method according to claim 1 , wherein the step of providing the audio output based on said output signal comprises the steps of:

suppressing the contribution to the audio output from the remaining sound source signals such that the combined level of the remaining sound source signals is in the range between 3 and 24 dB or between 6 and 18 dB below the output signal level; and

enabling the user to control the ratio between the output signal level and the combined level of the remaining sound source signals.

13 . An audio device system comprising at least one audio device, wherein said at least one audio device comprises an acoustical-electrical input transducer and an electrical-acoustical output transducer, and wherein said audio device system further comprises

a sound source signal separator adapted to receive an input signal from said acoustical-electrical input transducer and to provide a plurality of sound source signals each representing a sound source of a present sound environment;

a first sound signal selector adapted to enable a user to select a first sound signal comprising speech, wherein said first sound signal is based on said audio input signal;

a speech content comparator adapted to compare the speech content of said first sound signal with the speech content of the provided plurality of sound source signals, and adapted to select, based on said comparison, as output signal the sound source signal that most likely is part of a conversation that the audio device system user is paying attention to,

wherein said comparing comprises:

(i) generating, by a natural language processing model, semantic and/or syntactic numerical representations of the speech content of the first sound signal and of each of the plurality of sound source signals, and

(ii) determining, by a language model, for each of the plurality of sound source signals a probability that the speech content thereof constitutes a conversational response to the speech content of the first sound signal, and

wherein said selecting comprises selecting the sound source signal having the highest predicted probability of being such a conversational response and/or having a highest semantic and/or syntactic similarity with the first sound signal; and

a digital signal processor adapted to process the output signal, such that the contribution to the audio output from the remaining sound source signals is suppressed compared to the contribution from the output signal; wherein

said processed output signal is provided to the electrical-acoustical output transducer in order to provide the audio output, and wherein the acoustical-electrical input transducer provides an input signal representing the present sound environment and provides the input signal to the sound source signal separator and to the first sound signal selector.