IP Library Granted Patent US 12677103
Granted Patent B2
US 12677103 · App. 18/566,322 · Granted Jul 7, 2026

Method of operating an audio device system and audio device system

Inventors: Rasmus Malik Hoeegh Lindrup (Berkeley, CA); Jens Brehm Bagger Nielsen (Ballerup, DK); Asger Ougaard (Frederikssund, DK); Robert Scholes Lyck (Soeborg, DK)
Assignee: WIDEX A/S
H04R25/507G10L17/04G10L19/08G10L19/167G10L21/007G10L21/0208G10L21/0272G10L25/30G10L25/87H04R2225/41H04R2225/43
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12677103
App. No.
18/566,322
Granted
Jul 7, 2026
Kind
B2
Abstract

A method ( 100 ) of operating an audio device system in order to provide at least one of improved noise reduction and speech intelligibility and an audio device system ( 200, 300, 400 ) adapted to carry out the method.

Claims (53)

1 . A method of operating an audio device system comprising the steps of:

a) providing an audio signal;

b) passing a frame of the audio signal x t through a neural network encoder and hereby obtaining a latent encoding z t ;

c) manipulating said latent encoding z t to provide a transformed latent encoding z t ′ by using a model configured to transform said latent encoding z t , wherein said step of manipulating said latent encoding z t to provide a transformed latent encoding z t ′ is adapted to provide at least one of:

noise suppression;

change of at least one speech characteristic selected from a group comprising: spectral speech characteristics, temporal speech characteristics, enunciation, pronunciation, pitch and phoneme emphasis; and

at least one of sound source separation, removal, suppression and enhancement;

d) using a forecasting model to provide a prediction of a future transformed latent encoding z t+k ′ based at least on said transformed latent encoding z t ′;

e) passing said predicted future transformed latent encoding z t+k ′ through a neural network decoder and hereby providing a predicted future transformed electrical output signal {tilde over (s)} t+k ;

f) using an electrical-acoustical output transducer of the audio device to generate an acoustical output based on the predicted future electrical output signal {tilde over (s)} t+k .

2 . The method according to claim 1 , wherein said audio signal is selected from a group comprising: an audio signal derived from at least one acoustical-electrical input transducer accommodated in the audio device system, an audio signal being wirelessly transmitted to a device of the audio device system, and an audio signal generated internally by a device of the audio device system.

3 . The method according to claim 1 , wherein

said encoder and said decoder are both part of a variational autoencoder.

4 . The method according to claim 1 , comprising the further steps of:

using a probabilistic forecasting model to provide an estimate of the uncertainty of the forecasting; and

only using a prediction of a future transformed latent encoding z t+k ′ if the estimated uncertainty is below a given threshold.

5 . The method according to claim 1 , wherein

at least one of the steps b), c), d) and e) are carried out in at least one computing device of the audio device system and wherein

the processing delay resulting from at least one of: data transmission between the devices of the audio device system, the processing carried out in step b), the processing carried out in step c) and the processing carried out in step e)

have been at least partly compensated by generating an acoustical output signal based on the predicted future electrical output signal {tilde over (s)} t+k .

6 . The method according to claim 1 , comprising the further step of selecting a predicted future electrical output signal s t+k to be forwarded to the electrical-acoustical output transducer based on selecting the predicted future electrical output signal {tilde over (s)} t+k having a time stamp that matches the time stamp of the most recent frame of the audio signal x t .

7 . An audio device system comprising at least one audio device and at least one computing device, wherein

said at least one computing device is selected from a group of devices comprising: a personal computing device, a smart phone, a remote microphone system and a remote server, wherein

said at least one audio device is selected from a group of devices comprising: an ear level audio device, an earphone and a hearing aid, wherein

at least one computing device is adapted to carry out at least one of the following method steps:

passing a frame of an audio signal x t through a neural network encoder and hereby obtaining a latent encoding z t ,

manipulating said latent encoding z t to provide a transformed latent encoding z t ′ by using a model configured to transform said latent encoding z t , wherein said step of manipulating said latent encoding z t to provide a transformed latent encoding z t ′ is adapted to provide at least one of:

noise suppression;

change of at least one speech characteristic selected from a group comprising: spectral speech characteristics, temporal speech characteristics, enunciation, pronunciation, pitch and phoneme emphasis; and

at least one of sound source separation, removal, suppression and enhancement;

using a forecasting model to provide a prediction of a future transformed latent encoding z t+k ′ based at least on said transformed latent encoding z t ′;

passing said predicted future transformed latent encoding z t+k ′ through a neural network decoder and hereby providing a predicted future transformed electrical output signal {tilde over (s)} t+k ;

wherein

said at least one audio device is adapted to carry out the method step of using an electrical-acoustical output transducer of the audio device to generate an acoustical output based on the predicted future electrical output signal {tilde over (s)} t+k

f) of claim 1 , and wherein

the audio signal is provided by a computing device or an audio device.

8 . The audio device system ( 200 ) according to claim 7 , comprising a synchronization unit ( 212 ) adapted to select a predicted future electrical output signal {tilde over (s)} t+k to be forwarded to an electrical-acoustical output transducer ( 214 ) based on selecting the predicted future electrical output signal {tilde over (s)} t+k having a time stamp that matches the time stamp of the most recent frame of the audio signal x t .

9 . An internet server comprising a downloadable application that may be executed by a computing device, wherein the downloadable application is adapted to cause at least one of the following method steps to be carried out:

passing a frame of an audio signal x t through a neural network encoder and hereby obtaining a latent encoding z t ,

manipulating said latent encoding z t to provide a transformed latent encoding z t ′ by using a model configured to transform said latent encoding z t , wherein said step of manipulating said latent encoding Zi to provide a transformed latent encoding z t ′ is adapted to provide at least one of:

noise suppression;

change of at least one speech characteristic selected from a group comprising: spectral speech characteristics, temporal speech characteristics, enunciation, pronunciation, pitch and phoneme emphasis; and

at least one of sound source separation, removal, suppression and enhancement;

using a forecasting model to provide a prediction of a future transformed latent encoding z t+k ′ based at least on said transformed latent encoding z t ′;

passing said predicted future transformed latent encoding z t+k ′ through a neural network decoder and hereby providing a predicted future transformed electrical output signal {tilde over (s)} t+k .

10 . A non-transitory computer readable medium carrying instructions which, when executed by a computing device causes at least one of the following method steps to be carried out:

passing a frame of an audio signal x t through a neural network encoder and hereby obtaining a latent encoding z t ,

manipulating said latent encoding z t to provide a transformed latent encoding z t ′ by using a model configured to transform said latent encoding z t , wherein said step of manipulating said latent encoding z t to provide a transformed latent encoding z t ′ is adapted to provide at least one of:

noise suppression;

change of at least one speech characteristic selected from a group comprising: spectral speech characteristics, temporal speech characteristics, enunciation, pronunciation, pitch and phoneme emphasis; and

at least one of sound source separation, removal, suppression and enhancement;

using a forecasting model to provide a prediction of a future transformed latent encoding z t+k ′ based at least on said transformed latent encoding z t ′;

passing said predicted future transformed latent encoding z t+k ′ through a neural network decoder and hereby providing a predicted future transformed electrical output signal {tilde over (s)} t+k .