IP Library › Granted Patent US 12,499,873
Granted Patent B2
US 12,499,873 · App. 18/405,666 · Granted Dec 16, 2025

Method for personalisation of ASR models

Inventors: Umberto Michieli (Staines, GB); Mete Ozay (Staines, GB); Edward Fish (Staines, GB)
Assignee: Samsung Electronics Co., Ltd.
G10L15/063G10L15/30
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,499,873
App. No.
18/405,666
Granted
Dec 16, 2025
Kind
B2
Abstract

The present techniques generally relate to a computer-implemented method for personalising automatic speech recognition (ASR) or other general-purpose models using mixed precision (MP) quantization. Each model is personalised on a user device or server to a target memory budget B using user data.

Claims (80)

1 . A computer-implemented method for generating a personalized local machine learning (ML) model data, the method comprising:

obtaining a general ML model in a form of a neural network comprising a plurality of layers with each layer having a set of general weights;

obtaining a memory budget for the local ML model;

obtaining a set of local data samples; and

for each layer in the plurality of layers of the general ML model:

detecting a sensitivity of the layer to quantization by applying the general ML model to the set of local data samples;

determining, using the detected sensitivity of the layer and the obtained memory budget, an optimum bit depth for the corresponding layer in the local ML model, so that the local ML model satisfies the memory budget;

defining at least one scaling factor for the layer to scale the set of general weights to the optimum bit depth; and

generating a set of quantized weights by quantizing the set of general weights for the layer using the at least one scaling factor,

wherein the set of quantized weights for each layer define the personalized local ML model, and

wherein there are multiple optimum bit depths which satisfy a uniformity constraint.

2 . The method of claim 1 , wherein the memory budget for the local ML model is between a tenth to a half of a memory budget required to store the general ML model.

3 . The method of claim 1 , wherein the general ML model is an automatic speech recognition model and the set of local data samples comprise audio data samples.

4 . The method of claim 1 , wherein detecting the sensitivity of the layer to quantization comprises:

detecting activation of each layer; and

calculating, using the detected activation, a statistic which is indicative of the sensitivity of the layer to quantization.

5 . The method of claim 4 , further comprising:

determining an optimum bit depth by:

ranking each layer using the calculated statistic; and

assigning lower optimum bit depths to the lower ranked layers.

6 . The method of any claim 1 , wherein determining an optimum bit depth comprises:

reducing a bit depth of a first layer in the plurality of layers;

computing a size of the local ML model with the reduced bit depth;

comparing the calculated size with the memory budget,

when the calculated size exceeds the memory budget, reducing a bit depth of a subsequent layer in the plurality of layers; and

repeating the computing, comparing and reducing a bit depth of a subsequent layer until the calculated size meets the memory budget.

7 . The method of claim 6 , wherein the first layer is the lowest ranked layer and a subsequent layer is a layer which is higher in the ranking.

8 . The method of claim 1 , wherein defining at least one scaling factor for the layer comprises:

detecting minimum and maximum values activation of each layer when applying the general ML model to the set of local data samples; and

defining a scaling factor for each layer using a difference between the minimum and maximum values activation of each layer.

9 . The method of claim 8 , wherein the scaling factor S l for each layer is defined by:

S

l

=

(

X

l

M

-

X

l

m

)

/

(

2

b

l

-

1

)

where X l m is the minimum activation value for each layer of the full precision model, X l M is the maximum activation value for each layer of the full precision model and b l is the number of bits in the layer.

10 . The method of claim 1 , wherein there is one scaling factor for each layer and each scaling factor is defined by minimizing a distance between a quantized output from each layer of the local ML model when using a calibration data set and an output from the corresponding layer in the general ML model when using the calibration data set.

11 . The method of claim 1 , wherein there are two scaling factors for each layer and the two scaling factors are defined using two quantization ranges per layer.

12 . The method of claim 11 , further comprising:

defining the two scaling factors by performing a linear search using a calibration data set.

13 . The method of claim 12 , further comprising:

using a Hessian-based calibration optimization in the linear search.

14 . A user device for generating a personalized local machine learning (ML) model, wherein the user device is configured to:

obtain a general ML model in the form of a neural network comprising a plurality of layers with each layer having a set of general weights;

obtain a memory budget for the local ML model;

obtain a set of local data samples; and

for each layer in the plurality of layers of the general ML model:

detect, using a sensitivity module, a sensitivity of the layer to quantization by applying the general ML model to the set of local data samples;

determine, using the detected sensitivity of the layer to quantization and the obtained memory budget, an optimum bit depth for the corresponding layer in the local ML model, so that the local ML model satisfies the memory budget;

define at least one scaling factor for the layer to scale the set of general weights to the optimum bit depth; and

generate, using a quantization module, a set of quantized weights by quantizing the set of general weights for the layer using the at least one scaling factor,

wherein the set of quantized weights for each layer define the personalized local ML model, and

wherein there are multiple optimum bit depths which satisfy a uniformity constraint.

15 . A server for generating personalized local machine learning (ML) model data to be used on a user device, wherein the server is configured to:

train a general ML model in the form of a neural network comprising a plurality of layers with each layer having a set of general weights;

obtain a memory budget for the local ML model to be stored on the user device;

obtain a set of local data samples from the user device; and

for each layer in the plurality of layers of the general ML model:

detect, using a sensitivity module, a sensitivity of the layer to quantization by applying the general ML model to the set of local data samples;

determine, using the detected sensitivity of the layer to quantization and the obtained memory budget, an optimum bit depth for the corresponding layer in the local ML model, so that the local ML model satisfies the memory budget;

define at least one scaling factor for the layer to scale the set of general weights to the optimum bit depth; and

generate, using a quantization module, a set of quantized weights by quantizing the set of general weights for the layer using the at least one scaling factor,

wherein the set of quantized weights for each layer define the personalized local ML model, and

wherein there are multiple optimum bit depths which satisfy a uniformity constraint.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jan 5, 2024
From: MICHIELI, UMBERTO; OZAY, METE; FISH, EDWARD
To: SAMSUNG ELECTRONICS CO., LTD.
Reel/Frame 066037/0695 →
Priority Claims (2)
GB 2301370 · Jan 31, 2023 · national
GB 2305633 · Apr 17, 2023 · national
Continuity (1)
Related Publication 20240257800A1 · Aug 1, 2024
References Cited (59)
US 9190057B2 · Hoffmeister et al. · 2015 [cited by applicant]
US 10354656B2 · Zhao et al. · 2019 [cited by applicant]
US 10964312B2 · Barton et al. · 2021 [cited by applicant]
US 11398238B2 · Kim et al. · 2022 [cited by applicant]
US 20170085867A1 · Baran · 2017 [cited by examiner]
US 20200234112A1 · Wang · 2020 [cited by examiner]
US 20210160499A1 · Wang · 2021 [cited by examiner]
US 20210203936A1 · Gish et al. · 2021 [cited by applicant]
US 20210264279A1 · Esser et al. · 2021 [cited by applicant]
US 20210279635A1 · Gadelrab et al. · 2021 [cited by applicant]
US 20210306578A1 · Xiong · 2021 [cited by examiner]
US 20210406690A1 · Ramachandran et al. · 2021 [cited by applicant]
US 20220044114A1 · Sriram et al. · 2022 [cited by applicant]
US 20220129736A1 · Shen et al. · 2022 [cited by applicant]
US 20220129759A1 · Yao et al. · 2022 [cited by applicant]
US 20220164411A1 · Jain · 2022 [cited by examiner]
US 20230072337A1 · Lee · 2023 [cited by examiner]
US 20230168921A1 · Kim · 2023 [cited by examiner]
US 20230267301A1 · El-Kurdi · 2023 [cited by examiner]
US 20230281423A1 · Bijalwan · 2023 [cited by examiner]
US 20240169180A1 · Yang · 2024 [cited by examiner]
CN 113903347A · 2022 [cited by applicant]
CN 114049530A · 2022 [cited by applicant]
WO 2022080790A1 · 2022 [cited by applicant]
WO 2022088063A1 · 2022 [cited by applicant]
Graves et al., Connectionist Temporal Classification: Labelling Unsegmented Sequence Data with Recurrent Neural Networks, In Proceedings of the International Conference on Machine Learning, 369-376, 2006. [cited by applicant]
Harwell, Alexa-does-not-understand-your-accent, https://www.washingtonpost.com/graphics/2018/business/alexa-does-not-understand-your-accent/; Jul. 19, 2018. [cited by applicant]
Rangarajan, Hey Siri—Why Don't You Understand More People Like Me?, https://www.motherjones.com/media/2021/02/digital-assistants-accents-english-race-google-siri-alexa/, Feb. 23, 2021. [cited by applicant]
Dee, Voice search: the latest statistics and trends for 2022 and beyond, https://www.algolia.com/blog/product/voice-search-the-latest-statistics-and-trends-for-2022-and-beyond/, 2022. [cited by applicant]
Enge, Mobile Voice Usage Trends in 2020, https://www.perficient.com/insights/research-hub/voice-usage-trends, Jun. 30, 2020. [cited by applicant]
Ravanelli et al., Speaker recognition from raw waveform with sincnet, 2018 IEEE Spoken Language Technology Workshop (SLT). IEEE, 2018. [cited by applicant]
Long et al., Fully Convolutional Networks for Semantic Segmentation, published in the Proceedings of the IEEE/CVF Computer Vision and Pattern Recognition (CVPR) conference, 2015. [cited by applicant]
Badrinarayanan et al., SegNet: A Deep Convolutional Encoder-Decoder Architecture for Image Segmentation, published in the Proceedings of the IEEE/CVF Computer Vision and Pattern Recognition (CVPR) conference, 2015. [cited by applicant]
Strudel et al., Segmenter: Transformer for Semantic Segmentation, published in the Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV), 2021. [cited by applicant]
Liu et al., Swin Transformer: Hierarchical Vision Transformer using Shifted Windows, published in the Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV), 2021. [cited by applicant]
Baevski et al., wav2vec 2.0: A framework for self-supervised learning of speech representations, published in NeurIPS, 2020. [cited by applicant]
Rahaman et al., On the Spectral Bias of Neural Networks, published in ICML, 2019. [cited by applicant]
Bernhardsson, Language Pitch, Feb. 1, 2017. [cited by applicant]
Fukushima, Cognitron: A self-organizing multilayered neural network, published in Biological Cybernetics, 20(3), 121-136, Feb. 4, 1975. [cited by applicant]
Maas et al., Rectifier nonlinearities improve neural network acoustic models, published in the International Conference on Machine Learning (ICML). vol. 30. No. 1., 2013. [cited by applicant]
Yu et al., Low-bit quantization needs good distribution, published in the Computer Vision and Pattern Recognition Worksops (CVPRW), 2020. [cited by applicant]
Gholamit et al., A survey of quantization methods for efficient neural network inference, publisehd in arXiv:2103.13630, 2021. [cited by applicant]
Stevens et al., Deep learning with PyTorch, Chapter 15, published by Manning Publications and on the Pytorch website at https://pytorch.org/docs/stable/quantization.html, 2020. [cited by applicant]
Wu et al., Easyquant: Post-training quantization via scale optimization, published in NeurIPS, 2020. [cited by applicant]
Yuan et al., Ptq4vit: Post-training quantization for vision transformers with twin uniform quantization, published in Springer, pp. 191 to 207, 2022. [cited by applicant]
Panayotov et al., Librispeech: An asr corpus based on public domain audio books, published in ICASSP, 2015. [cited by applicant]
Conneau et al., Fleurs: Few-shot learning evaluation of universal representations of speech, published in Spoken Language Technologies (SLT), 2022. [cited by applicant]
Radford et al., Robust speech recognition via large-scale weak supervision, published in arXiv:2212.04356, 2022. [cited by applicant]
Warden et al., Speech commands: A dataset for limited-vocabulary speech recognition, published in arXiv:1804.03209, 2018. [cited by applicant]
Dong et al., Hawq: Hessian aware quantization of neural networks with mixed-precision, published in ICCV, 2019. [cited by applicant]
Wang et al., FAIRSEQ S2t: Fast speech to text modelling with FAIRSEQ, published in AACL, 2020. [cited by applicant]
Liu et al., Post-training quantization for vision transformer, published in NeurIPS, 2021. [cited by applicant]
Eryilmaz et al., Understanding how orthogonality of parameters improves quantization of neural networks, published in IEEE TNNLS, 2022. [cited by applicant]
Chen et al., Towards Mixed-Precision Quantization of Neural Networks via Constrained Optimization, Oct. 13, 2021. [cited by applicant]
Gao et al., Extremely Low Footprint End-to-End ASR System for Smart Device, Jul. 7, 2021. [cited by applicant]
Yuan et al., PTQ4ViT: Post-Training Quantization for Vision Transformers with Twin Uniform Quantization, Jul. 27, 2022. [cited by applicant]
Yao et al., HAWQ-V3: Dyadic Neural Network Quantization, Jun. 23, 2021. [cited by applicant]
Tsuji et al., GPQ: Greedy Partial Quantization of Convolutional Neural Networks Inspired by Submodular Optimization, 2020. [cited by applicant]
Combined Search and Examination Report, dated Feb. 13, 2024, issued in Great Britain Application No. GB2305633.6. [cited by applicant]