IP Library › Granted Patent US 12,321,846
Granted Patent B2
US 12,321,846 · App. 17/116,117 · Granted Jun 3, 2025

Knowledge distillation using deep clustering

Inventor: Takashi Fukuda (Tokyo, JP)
Assignee: INTERNATIONAL BUSINESS MACHINES CORPORATION
G06N3/045G06F40/44G06N3/088
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,321,846
App. No.
17/116,117
Granted
Jun 3, 2025
Kind
B2
Abstract

Methods and systems for training a neural network include clustering a full set of training data samples into specialized training clusters. Specialized teacher neural networks are trained using respective specialized training clusters of the specialized training clusters. Soft labels are generated for the full set of training data samples using the specialized teacher neural networks. A student model is trained using the full set of training data samples, the specialized training clusters, and the soft labels.

Claims (142)

1. A computer-implemented method for training a neural network, comprising:

clustering, by a clustering neural network, a full set of training data samples into a plurality of specialized training clusters according to acoustic conditions estimated through a generated supervector using an average of phone alignment information, representing frame alignment of phonemes, by employing log likelihoods of the acoustic conditions concatenated with averaged log Mel-frequency features for context-dependent phonemes as features estimated by a general teacher neural network;

training a plurality of specialized teacher neural networks using respective specialized training clusters of the plurality of specialized training clusters;

generating soft labels for the full set of training data samples using the plurality of specialized teacher neural networks; and

training a student model using the full set of training data samples, the specialized training clusters, and the soft labels.

2. The method of claim 1 , wherein the training data samples include acoustic waveforms.

3. The method of claim 1 , wherein clustering includes unsupervised deep clustering.

4. The method of claim 1 , further comprising training a general teacher neural network using the full set of training data samples.

5. The method of claim 4 , wherein generating the soft labels further uses the general teacher neural network.

6. The method of claim 5 , wherein training the student model comprises minimizing a loss function:

ℒ

=

-

(

1

-

λ

)

⁢

∑

i

⁢

q

bl

⁡

(

i

|

x

)

⁢

log

⁢

p

⁡

(

i

|

x

)

-

λ

⁢

∑

i

⁢

q

sp

⁡

(

i

|

x

)

⁢

log

⁢

p

⁡

(

i

|

x

)

where λ is a weight hyperparameter, q bl is a set of soft labels from the general teacher neural network, q sp is a set of soft labels from the plurality of specialized teacher neural networks, and p(i|x) is a probability of a class from the student model, with i indicating an index of context-dependent phonemes and x indicating an input signal.

7. The method of claim 4 , wherein the student model is a neural network that has fewer parameters than the general teacher neural network.

8. The method of claim 1 , further comprising performing speech recognition on a new utterance using the trained student model.

9. The method of claim 8 , further comprising performing a natural language task on recognized speech from the new utterance.

10. A non-transitory computer readable storage medium comprising a computer readable program for training a neural network, wherein the computer readable program when executed on a computer causes the computer to:

cluster, by a clustering neural network, a full set of training data samples into a plurality of specialized training clusters according to acoustic conditions estimated through a generated supervector using an average of phone alignment information, representing frame alignment of phonemes, by employing log likelihoods of the acoustic conditions concatenated with averaged log Mel-frequency features for context-dependent phonemes as features estimated by a general teacher neural network;

train a plurality of specialized teacher neural networks using respective specialized training clusters of the plurality of specialized training clusters;

generate soft labels for the full set of training data samples using the plurality of specialized teacher neural networks; and

train a student model using the full set of training data samples, the specialized training clusters, and the soft labels.

11. The non-transitory computer readable storage medium of claim 10 , wherein the training data samples include acoustic waveforms.

12. The non-transitory computer readable storage medium of claim 10 , wherein computer readable program further causes the computer to perform unsupervised deep clustering on the full set of training data samples.

13. The non-transitory computer readable storage medium of claim 10 , wherein computer readable program further causes the computer to train a general teacher neural network using the full set of training data samples.

14. The non-transitory computer readable storage medium of claim 13 , wherein computer readable program further causes the computer to use the general teacher neural network to generate soft labels.

15. The non-transitory computer readable storage medium of claim 14 , wherein computer readable program further causes the computer to a loss function to train the student model:

ℒ

=

-

(

1

-

λ

)

⁢

∑

i

⁢

q

bl

⁡

(

i

|

x

)

⁢

log

⁢

p

⁡

(

i

|

x

)

-

λ

⁢

∑

i

⁢

q

sp

⁡

(

i

|

x

)

⁢

log

⁢

p

⁡

(

i

|

x

)

where λ is a weight hyperparameter, qui is a set of soft labels from the general teacher neural network, q sp is a set of soft labels from the plurality of specialized teacher neural networks, and p(i|x) is a probability of a class from the student model, with i indicating an index of context-dependent phonemes and x indicating an input signal.

16. The non-transitory computer readable storage medium of claim 13 , wherein the student model is a neural network that has fewer parameters than the general teacher neural network.

17. The non-transitory computer readable storage medium of claim 10 , wherein computer readable program further causes the computer to perform speech recognition on a new utterance using the trained student model.

18. A system for training a neural network, comprising:

a hardware processor; and

a memory that stores computer program code which, when executed by the hardware processor, implements:

a student model;

a clustering neural network that clusters a full set of training data samples into a plurality of specialized training clusters according to acoustic conditions estimated through a generated supervector using an average of phone alignment information, representing frame alignment of phonemes, by employing log likelihoods of the acoustic conditions concatenated with averaged log Mel-frequency features for context-dependent phonemes as features estimated by a general teacher neural network;

a plurality of specialized teacher neural networks that together generate soft labels for the full set of training data samples; and

a model trainer that trains the plurality of specialized teacher neural networks using respective specialized training clusters of the plurality of specialized training clusters, and that trains a student model using the full set of training data samples, the specialized training clusters, and the soft labels.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Dec 9, 2020
From: FUKUDA, TAKASHI
To: INTERNATIONAL BUSINESS MACHINES CORPORATION
Reel/Frame 054591/0339 →
Continuity (1)
Related Publication 20220180206A1 · Jun 9, 2022
References Cited (46)
US 5611019A · Nakatoh · 1997 [cited by examiner]
US 6067517A · Bahl · 2000 [cited by examiner]
US 7417983B2 · He · 2008 [cited by examiner]
US 9412361B1 · Geramifard · 2016 [cited by examiner]
US 11907845B2 · Fukuda et al. · 2024 [cited by applicant]
US 20030023436A1 · Eide · 2003 [cited by examiner]
US 20050137862A1 · Monkowski · 2005 [cited by examiner]
US 20080300875A1 · Yao · 2008 [cited by examiner]
US 20090222266A1 · Sakai · 2009 [cited by examiner]
US 20140074614A1 · Mehanian · 2014 [cited by examiner]
US 20160284347A1 · Sainath · 2016 [cited by examiner]
US 20170270100A1 · Audhkhasi · 2017 [cited by examiner]
US 20190080684A1 · Ichikawa · 2019 [cited by examiner]
US 20190378006A1 · Fukuda et al. · 2019 [cited by applicant]
US 20200034702A1 · Fukuda et al. · 2020 [cited by applicant]
US 20200034703A1 · Fukuda et al. · 2020 [cited by applicant]
US 20200186897A1 · Dareddy · 2020 [cited by examiner]
US 20210407090A1 · Li · 2021 [cited by examiner]
US 20220188622A1 · Nagano et al. · 2022 [cited by applicant]
US 20220188643A1 · Fukuda · 2022 [cited by applicant]
US 20220344049A1 · Hall · 2022 [cited by examiner]
CN 110033077A · 2019 [cited by applicant]
CN 110674880A · 2020 [cited by applicant]
CN 111950302A · 2020 [cited by applicant]
CN 110852426A · 2023 [cited by applicant]
CN 111933185A · 2024 [cited by applicant]
WO 2010130733A1 · 2010 [cited by applicant]
WO 2018169708A1 · 2018 [cited by applicant]
WO 2020194077A1 · 2020 [cited by applicant]
Asif et al. “Ensemble Knowledge Distillation for Learning Improved and Efficient Networks”, Sep. 2019, arXiv.org, <arxiv.org/abs/1909.08097v1> (Year: 2019). [cited by examiner]
Wu et al., “Learning to Teach with Dynamic Loss Functions”, Oct. 2018, arXiv.org, <arxiv.org/abs/1810.12081> (Year: 2018). [cited by examiner]
Y. Chebotar and A.Waters, “Distilling knowledge from ensembles of neural networks for speech recognition,” Proc. Interspeech, pp. 3439-3443, 2016 (Year: 2016). [cited by examiner]
J. Anden and S. Mallat, “Deep Scattering Spectrum,” in IEEE Transactions on Signal Processing, vol. 62, No. 16, pp. 4114-4128, Aug. 15, 2014, doi: 10.1109/TSP.2014.2326991 (Year: 2014). [cited by examiner]
Price et al., “Wise teachers train better DNN acoustic models”, EURASIP Journal on Audio, Speech, and Music Processing, Dec. 2016, pp. 1-19. [cited by applicant]
Xu et al., “Knowledge Distillation from Multilingual and Monolingual Teachers for End-to-End Multilingual Speech Recognition”, Proceedings of APSIPA Annual Summit and Conference 2019, Nov. 2019, pp. 844-849. [cited by applicant]
Fukuda et al., “Implicit Transfer of Privileged Acoustic Information in a Generalized Knowledge Distillation Framework”, Interspeech 2020, Oct. 2020, pp. 41-45. [cited by applicant]
UK Combined Search and Examination Report issued in corresponding Patent Application No. GB2116914.9 dated Sep. 23, 2022, 13 pages. [cited by applicant]
Asif et al., “Ensemble Knowledge Distillation for Learning Improved and Efficient Networks”, 24th European Conference on Artificial Intelligence. Aug. 29-Sep. 8, 2020. pp. 1-8. [cited by applicant]
Caron et al., “Deep Clustering for Unsupervised Learning of Visual Features”, arXiv:1807.05520v2 [cs.CV]. Mar. 18, 2019. pp. 1-30. [cited by applicant]
Fukuda et al., “Efficient Knowledge Distillation from an Ensemble of Teachers”, Interspeach 2017. Aug. 20-24, 2017. pp. 3697-3701. [cited by applicant]
Fukuda et al., “Mixed Bandwidth Acoustic Modeling Leveraging Knowledge Distillation”, IEEE Automatic Speech Recognition and Understanding Workshop. Dec. 14, 2019. pp. 1-7. [cited by applicant]
Jung et al., “Distilling the Knowledge of Specialist Deep Neural Networks in Acoustic Scene Classification”, 2019 Detection and Classification of Acoustic Scenes and Events. Oct. 25-26, 2019. pp. 114-118. [cited by applicant]
Mell et al., “The NIST Definition of Cloud Computing”, National Institute of Standards and Technology, Special Publication 800-145, Sep. 2011, pp. 1-7. [cited by applicant]
Niu et al., “GATCluster: Self-Supervised Gaussian-Attention Network for Image Clustering”, arXiv:2002.11863v1 [cs.CV]. Feb. 27, 2020. pp. 1-14. [cited by applicant]
Upadhyay, Ujjwal, “Knowledge Distillation”, Neural Machine—Medium. https://medium.com/neuralmachine/knowledge-distillation-dc241d7c2322. Downloaded on Aug. 28, 2020. pp. 1-10. [cited by applicant]
Xiang et al. “Learning from Multiple Experts: Self-paced Knowledge Distillation for Long-tailed Classification”, arXiv:2001.01536 [cs.CV], Sep. 21, 2020, 17 pages. [cited by applicant]