IP Library Granted Patent US 12,658,178
Granted Patent B2
US 12,658,178 · App. 18/587,238 · Granted Jun 16, 2026

Method for personalization of ASR models

Inventors: Pablo Peso Parada (Chertsey, GB); Mete Ozay (Chertsey, GB); Karthikeyan Saravanan (Chertsey, GB)
Assignee: Samsung Electronics Co., Ltd.
G10L15/075G10L15/02G10L15/04G10L15/063G10L15/07G10L15/1815G10L15/22G10L15/30G10L21/028G10L2015/223
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,658,178
App. No.
18/587,238
Granted
Jun 16, 2026
Kind
B2
Abstract

The disclosure generally relate to a method performed by a user device obtaining a pre-trained automatic speech recognition (ASR) model, obtaining a user data from a user database, analysing a distribution of the user data with respect to an acoustic characteristic, determining, using the distribution, whether data augmentation for the acoustic characteristic is to be applied, when it is determined that data augmentation is to be applied, dividing the user data into a training subset and a validation subset, based on an acoustic characteristic being less audible in the training subset than in the validation subset, applying data augmentation to add the acoustic characteristic to the user data in the training subset, and updating the pre-trained ASR model with the augmented training subset to generate a personalised local ML model.

Claims (73)

1 . A method, performed by a user device, for generating a local machine learning (ML) model which is personalised to a user of the user device, the method comprising:

obtaining a pre-trained automatic speech recognition (ASR) model;

obtaining a user data from a user database;

analysing a distribution of the user data with respect to an acoustic characteristic;

determining, using the distribution, whether data augmentation for the acoustic characteristic is to be applied;

when it is determined that data augmentation is to be applied, dividing the user data into a training subset and a validation subset, wherein the acoustic characteristic is less audible in the training subset than in the validation subset;

applying data augmentation to add the acoustic characteristic to the training subset; and

updating the pre-trained ASR model with the augmented training subset to generate a personalised local ML model.

2 . The method of claim 1 , wherein the applying of the data augmentation comprises applying the data augmentation to add the acoustic characteristic to the training subset in parallel.

3 . The method of claim 1 ,

wherein the acoustic characteristic is selected from environmental acoustic characteristics, user acoustic characteristics, and semantic and syntactic characteristics,

wherein the environmental acoustic characteristics includes noise, signal to noise ratio and reverberation,

wherein the user acoustic characteristics includes speed and spectral characteristics, and

wherein the semantic and syntactic characteristics includes meaning and structure of a text uttered by the user.

4 . The method of claim 1 , wherein the analysing of the distribution of the user data with respect to an acoustic characteristic comprises plotting a probability distribution P(θ|X) for user dataset X for acoustic characteristic θ.

5 . The method of claim 1 , wherein the determining, using the distribution, whether the data augmentation is to be applied, comprises selecting data augmentation parameters which maximise a probability of generating similar acoustic characteristics in the user data.

6 . The method of claim 1 ,

wherein the determining, using the distribution, whether the data augmentation is to be applied, comprises:

setting personalised settings, for the user, for a data augmentation parameter which corresponds to the acoustic characteristic, and

wherein the personalised settings comprise at least one statistical value calculated from the corresponding distribution.

7 . The method of claim 6 , further comprising:

comparing a value of at least one of the personalised settings determined from the corresponding distribution to a threshold; and

when the value of the at least one of the personalised settings meets the threshold, determining that the data augmentation is to be applied.

8 . The method of claim 7 , further comprising:

generating a signal having an acoustic characteristic which matches values of the at least one of the personalised settings determined from the corresponding distribution; and

adding the generated signal when applying the data augmentation.

9 . The method of claim 1 , further comprising:

obtaining, using at least one user data in the validation subset, a signal which matches the acoustic characteristic of the at least one user data; and

applying the data augmentation to add the obtained signal to user data in the training subset.

10 . The method of claim 9 , further comprising:

generating a noise signal to be applied the data augmentation,

wherein the applying of the data augmentation comprises adding the noise signal.

11 . The method of claim 10 , wherein the generating of the noise signal comprises:

extracting noise frames which are individual audio frames in the validation subset containing noise;

concatenating consecutive noise frames to form multiple noise segments; and

concatenating a subset of the noise segments to form a noise signal which is longer than the training subset.

12 . The method of claim 11 , further comprising:

comparing each of the multiple noise segments to a length threshold; and

rejecting at least one noise segment which are shorter than the length threshold.

13 . The method of claim 10 , further comprising:

retrieving a noise signal which is similar to the generated noise signal from a generic noise dataset; and

adding the retrieved noise signal when applying data augmentation.

14 . The method of claim 10 , wherein the generating of the noise signal to be applied the data augmentation comprises generating the noise signal to be applied by extracting background noise using a speech separation model.

15 . The method of claim 9 , wherein the applying of the data augmentation comprises applying a reverberation signal.

16 . The method of claim 15 , further comprising:

estimating a value of reverberation in the user data of the validation subset;

selecting the reverberation signal in a form of a room impulse response (RIR) with a matching value of reverberation; and

storing the selected reverberation signal in a personal dataset for the user.

17 . The method of claim 1 , further comprising:

obtaining, at the user device, input speech;

processing the input speech using the personalised local ML model to determine a voice command from the user; and

implementing the voice command on the user device to enable voice-based control of the user device.

18 . A user device comprising:

a memory storing one or more instructions; and

at least one processor configured to execute the one or more instructions to:

obtain a pre-trained automatic speech recognition (ASR) model,

obtain a user data from a user database,

analyse a distribution of the user data with respect to an acoustic characteristic,

determine, using the distribution, whether data augmentation for the acoustic characteristic is to be applied,

when it is determined that data augmentation is to be applied, divide the user data into a training subset and a validation subset, based on an acoustic characteristic being less audible in the training subset than in the validation subset,

apply data augmentation to add the acoustic characteristic in the training subset, and

update the pre-trained ASR model with the augmented training subset to generate a personalised local machine learning (ML) model.

19 . The user device of claim 18 , wherein the at least one processor is further configured to execute the one or more instructions to:

obtain, using at least one user data in the validation subset, a signal which matches the acoustic characteristic of the at least one user data, and

apply the data augmentation to add the obtained signal to user data in the training subset.

20 . A non-transitory data carrier carrying code which, when implemented on a processor of a user device, causes the user device to perform operations, the operations comprising:

obtaining a pre-trained automatic speech recognition (ASR) model;

obtaining a user data from a user database;

analysing a distribution of the user data with respect to an acoustic characteristic;

determining, using the distribution, whether data augmentation for the acoustic characteristic is to be applied;

when it is determined that data augmentation is to be applied, dividing the user data into a training subset and a validation subset, wherein the acoustic characteristic is less audible in the training subset than in the validation subset;

applying data augmentation to add the acoustic characteristic to the training subset; and

updating the pre-trained ASR model with the augmented training subset to generate a personalised local machine learning (ML) model.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Feb 26, 2024
From: PARADA, PABLO PESO; OZAY, METE; SARAVANAN, KARTHIKEYAN
To: SAMSUNG ELECTRONICS CO., LTD.
Reel/Frame 066561/0797 →
Priority Claims (2)
GB 2302913 · Feb 28, 2023 · national
GB 2309497 · Jun 23, 2023 · national
Continuity (1)
Related Publication 20250006183A1 · Jan 2, 2025
References Cited (58)
US 7346507B1 · Natarajan · 2008 [cited by examiner]
US 7533019B1 · Hakkani-Tur · 2009 [cited by examiner]
US 8135589B1 · Reding · 2012 [cited by examiner]
US 8682669B2 · Suendermann · 2014 [cited by examiner]
US 10102849B2 · Vozila · 2018 [cited by examiner]
US 11593588B2 · Kim et al. · 2023 [cited by applicant]
US 11694694B2 · Traynor · 2023 [cited by examiner]
US 12062369B2 · Kupryjanow · 2024 [cited by examiner]
US 12142298B1 · Kottur · 2024 [cited by examiner]
US 12190877B1 · Barber · 2025 [cited by examiner]
US 12340793B1 · Ljolje · 2025 [cited by examiner]
US 20020065656A1 · Reding · 2002 [cited by examiner]
US 20020065657A1 · Reding · 2002 [cited by examiner]
US 20070233487A1 · Cohen · 2007 [cited by examiner]
US 20070299664A1 · Peters · 2007 [cited by examiner]
US 20100268536A1 · Suendermann · 2010 [cited by examiner]
US 20160203817A1 · Formhals · 2016 [cited by examiner]
US 20170040016A1 · Cui · 2017 [cited by examiner]
US 20180090131A1 · Mangalath · 2018 [cited by examiner]
US 20180350348A1 · Fukuda · 2018 [cited by examiner]
US 20190034826A1 · Ahmad · 2019 [cited by examiner]
US 20190130896A1 · Zhou · 2019 [cited by examiner]
US 20190304480A1 · Narayanan · 2019 [cited by examiner]
US 20200005766A1 · Kim · 2020 [cited by examiner]
US 20200105262A1 · Abhinav · 2020 [cited by examiner]
US 20200327884A1 · Bui · 2020 [cited by examiner]
US 20200335086A1 · Paraskevopoulos · 2020 [cited by examiner]
US 20210035563A1 · Cartwright · 2021 [cited by examiner]
US 20210043186A1 · Nagano · 2021 [cited by examiner]
US 20210287660A1 · Sharma · 2021 [cited by examiner]
US 20210304737A1 · Han · 2021 [cited by examiner]
US 20220189461A1 · Zhao · 2022 [cited by examiner]
US 20220189463A1 · Oh et al. · 2022 [cited by applicant]
US 20220246132A1 · Zhang · 2022 [cited by examiner]
US 20220335937A1 · Thomas et al. · 2022 [cited by applicant]
US 20220392432A1 · Alphonso · 2022 [cited by examiner]
US 20230033768A1 · Chong · 2023 [cited by examiner]
US 20230056955A1 · Yao et al. · 2023 [cited by applicant]
US 20230061505A1 · Oh · 2023 [cited by examiner]
US 20230067305A1 · Assa · 2023 [cited by examiner]
US 20230107493A1 · Bijwadia · 2023 [cited by examiner]
US 20230116052A1 · Eskimez · 2023 [cited by examiner]
US 20230146945A1 · Lai · 2023 [cited by examiner]
US 20230197064A1 · Bekker · 2023 [cited by examiner]
US 20240055012A1 · Wang · 2024 [cited by examiner]
US 20250259641A1 · de la Rey · 2025 [cited by examiner]
US 20250329268A1 · Ji · 2025 [cited by examiner]
KR 1020220053475A · 2022 [cited by applicant]
WO 2021162536A1 · 2021 [cited by applicant]
Algolia; Voice search: the latest statistics and trends for 2022 and beyong; https://www.algolia.com/blog/product/voice-search-the-latest-statistics-andtrends-for-2022-and-beyond/; Mar. 2, 2023. [cited by applicant]
Graves et al.; Connectionist temporal classification: Labelling unsegmented sequence data with recurrent neural networks; 2006. [cited by applicant]
fireflies.ai; 15 Uses of Voice Recognition Software Today; https://fireflies.ai/blog/uses-of-voice-recognition-software; Jan. 20, 2020. [cited by applicant]
Lenzen; Motherjones; Hey Siri—Why Don't You Understand More People Like Me?; https://www.motherjones.com/media/2021/02/digital-assistants-accents-englishrace-google-siri-alexa/; Feb. 23, 2021. [cited by applicant]
Enge; Perficient; https://www.perficient.com/insights/research-hub/voice-usage-trends; Jun. 30, 2020. [cited by applicant]
Kim et al.; Improved Vocal Tract Length Perturbation for a State-of-the-Art End-to-End Speech Recognition System; ResearchGate; Interspeech 2019; Sep. 15-19, 2019; Graz, Austria. [cited by applicant]
Lam et al.; On-the-Fly Aligned Data Augmentation for Sequence-to-Sequence ASR; Jun. 9, 2021. [cited by applicant]
Ko et al.; A study on data augmentation of reverberant speech for robust speech recognition; 2017. [cited by applicant]
Harwell; The Accent Gap; The Washington Post; https://www.washingtonpost.com/graphics/2018/business/alexa-does-not-understand-your-accent/; Jul. 19, 2018. [cited by applicant]