IP Library Granted Patent US 12706110
Granted Patent B2
US 12706110 · App. 18/884,978 · Granted Aug 11, 2026

Leveraging self-supervised speech representations for domain adaptation in speech enhancement

Inventors: Ching-Hua Lee (Mountain View, CA); Chouchang Yang (San Jose, CA); Rakshith Sharma Srinivasa (Sunnyvale, CA); Yashas Malur Saidutta (Menlo Park, CA); Jaejin Cho (Mountain View, CA); Yilin Shen (Mountain View, CA); Hongxia Jin (San Jose, CA)
Assignee: SAMSUNG ELECTRONICS CO., LTD.
G10L21/0208G10L15/063G10L21/02G10L21/0216G10L2021/02166
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12706110
App. No.
18/884,978
Granted
Aug 11, 2026
Kind
B2
Abstract

A method for generating a customized speech enhancement model includes obtaining noisy-clean speech data from a source domain, obtaining noisy speech data from a target domain; obtaining raw speech data, using the noisy-clean speech data, the noisy speech data, and the raw speech data, training the customized SE model based on at least one of self-supervised representation-based adaptation (SSRA), ensemble mapping, or self-supervised adaptation loss, generating the customized SE model by denoising the noisy speech data using the trained customized SE model, and providing the customized SE model to a user device to use the denoised noisy speech data.

Claims (41)

1 . A method for generating a customized speech enhancement (SE) model, performed by at least one processor of an electronic device, the method comprising:

obtaining noisy-clean speech data from a source domain;

obtaining noisy speech data from a target domain;

obtaining raw speech data;

using the noisy-clean speech data, the noisy speech data, and the raw speech data, training the customized SE model based on at least one of self-supervised representation-based adaptation (SSRA), ensemble mapping, or self-supervised adaptation loss;

generating the customized SE model by denoising the noisy speech data using the trained customized SE model; and

providing the customized SE model to a user device to use the denoised noisy speech data.

2 . The method of claim 1 , wherein the training the customized SE model comprises training the customized SE model based on the SSRA, and the training the customized SE model further comprises pre-training a self-supervised learning (SSL) encoder in a self-supervised manner, providing a target domain enhanced signal to the SSL encoder, and providing source domain clean signals to the SSL encoder.

3 . The method of claim 1 , wherein the training the customized SE model comprises training the customized SE model based on the ensemble mapping, and the training the customized SE model further comprises pseudo labeling the noisy speech data from the target domain.

4 . The method of claim 1 , wherein the training the customized SE model comprises training the customized SE model based on the self-supervised adaptation loss, and the training the customized SE model further comprises using a distance metric in an SSRA loss term.

5 . The method of claim 1 , wherein the noisy speech data is obtained from the user device in the target domain.

6 . The method of claim 5 , wherein the user device comprises at least one of a mobile phone, a refrigerator, a smart watch, glasses, or a television.

7 . The method of claim 1 , wherein the noisy speech data is obtained from a plurality of microphones corresponding to a plurality of user devices.

8 . A server device comprising:

a memory storing instructions; and

at least one processor,

wherein the instructions, when executed by the at least one processor, cause the server device to:

obtain noisy-clean speech data from a source domain;

obtain noisy speech data from a target domain;

obtain raw speech data;

using the noisy-clean speech data, the noisy speech data, and the raw speech data, train a customized SE model based on at least one of self-supervised representation-based adaptation (SSRA), ensemble mapping, or self-supervised adaptation loss;

generate the customized SE model by denoising the noisy speech data using the trained customized SE model; and

provide the customized SE model to a user device to use the denoised noisy speech data.

9 . The server device of claim 8 , wherein the instructions, when executed by the at least one processor, cause the server device to pre-train a self-supervised learning (SSL) encoder in a self-supervised manner, provide a target domain enhanced signal to the SSL encoder, and provide source domain clean signals to the SSL encoder.

10 . The server device of claim 8 , wherein the instructions, when executed by the at least one processor, cause the server device to train the customized SE model based on the ensemble mapping, and pseudo label the noisy speech data from the target domain.

11 . The server device of claim 8 , wherein the instructions, when executed by the at least one processor, cause the server device to train the customized SE model based on the self-supervised adaptation loss, and use a distance metric in an SSRA loss term.

12 . The server device of claim 8 , wherein the noisy speech data is obtained from the user device in the target domain.

13 . The server device of claim 12 , wherein the user device comprises at least one of a mobile phone, a refrigerator, a smart watch, glasses, or a television.

14 . The server device of claim 8 , wherein the noisy speech data is obtained from a plurality of microphones corresponding to a plurality of user devices.

15 . A non-transitory computer-readable recording medium configured to store instructions for generating a customized speech enhancement (SE) model, which, when executed by at least one processor of an electronic device, cause the at least one processor to perform a method comprising:

obtaining noisy-clean speech data from a source domain;

obtaining noisy speech data from a target domain;

obtaining raw speech data;

using the noisy-clean speech data, the noisy speech data, and the raw speech data, training the customized SE model based on at least one of self-supervised representation-based adaptation (SSRA), ensemble mapping, or self-supervised adaptation loss;

generating the customized SE model by denoising the noisy speech data using the trained customized SE model; and

providing the customized SE model to a user device to use the denoised noisy speech data.

16 . The non-transitory computer-readable recording medium of claim 15 , wherein the training the customized SE model comprises training the customized SE model based on the SSRA, and the training the customized SE model further comprises pre-training a self-supervised learning (SSL) encoder in a self-supervised manner, providing a target domain enhanced signal to the SSL encoder, and providing source domain clean signals to the SSL encoder.

17 . The non-transitory computer-readable recording medium of claim 15 , wherein the training the customized SE model comprises training the customized SE model based on the ensemble mapping, and the training the customized SE model further comprises pseudo labeling the noisy speech data from the target domain.

18 . The non-transitory computer-readable recording medium of claim 15 , wherein the training the customized SE model comprises training the customized SE model based on the self-supervised adaptation loss, and the training the customized SE model further comprises using a distance metric in an SSRA loss term.

19 . The non-transitory computer-readable recording medium of claim 15 , wherein the noisy speech data is obtained from the user device in the target domain.

20 . The non-transitory computer-readable recording medium of claim 19 , wherein the user device comprises at least one of a mobile phone, a refrigerator, a smart watch, glasses, or a television.