IP Library › Granted Patent US 12,217,741
Granted Patent B2
US 12,217,741 · App. 17/324,535 · Granted Feb 4, 2025

Large scale privacy-preserving speech recognition system using federated learning

Inventors: Sylvain Le Groux (Placerville, CA); Erwan Barry Tarik Zerhouni (Zürich, CH)
Assignee: CISCO TECHNOLOGY, INC.
G10L15/16G06N3/088G10L15/142G10L25/51
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,217,741
App. No.
17/324,535
Granted
Feb 4, 2025
Kind
B2
Abstract

A method for implementing a privacy-preserving automatic speech recognition system using federated learning. The method includes receiving, from respective client devices, at a cloud server, local acoustic model weights for a neural network-based acoustic model of a local automatic speech recognition system running on the respective client devices, wherein the local acoustic model weights are generated at the respective client devices without labelled data, updating a global automatic speech recognition system based on (a) the local acoustic model weights received from the respective client devices and (b) global acoustic model weights of the global automatic speech recognition system derived from labelled data to obtain an updated global automatic speech recognition system, and sending the updated global automatic speech recognition system to the respective client devices to operate as a new local automatic speech recognition system.

Claims (32)

1. A method comprising:

receiving, from respective client devices, at a cloud server, local acoustic model weights for a neural network-based acoustic model of a local automatic speech recognition system running on the respective client devices, wherein the local acoustic model weights are generated at the respective client devices without labelled data;

updating a global automatic speech recognition system based on (a) the local acoustic model weights received from the respective client devices and (b) global acoustic model weights of the global automatic speech recognition system derived from labelled data to obtain an updated global automatic speech recognition system, wherein the updating comprises controlling an influence of unsupervised data on the updated global automatic speech recognition system by multiplying a first balancing coefficient with the local acoustic model weights and multiplying a second balancing coefficient with the global acoustic model weights, and combining resulting multiplication products thereof, wherein the first balancing coefficient is equal to one minus the second balancing coefficient, and wherein the second balancing coefficient has a value of zero to one; and

sending the updated global automatic speech recognition system to the respective client devices to operate as a new local automatic speech recognition system.

2. The method of claim 1 , further comprising selecting the respective client devices from which to receive the local acoustic model weights.

3. The method of claim 1 , wherein the global automatic speech recognition system is a hybrid Deep Neural Network/Hidden Markov Model (DNN-HMM) automatic speech recognition system.

4. The method of claim 1 , further comprising executing a language model that is based on one of recurrent neural network modeling or efficient N-Gram statistics.

5. The method of claim 1 , further comprising receiving from the respective client devices an indication of a language processed by the respective client devices.

6. The method of claim 1 , wherein updating the global automatic speech recognition system comprises calculating differences between global acoustic model weights of the global automatic speech recognition system and respective local acoustic model weights received from the respective client devices.

7. The method of claim 6 , wherein updating the global automatic speech recognition system further comprises calculating a weighted average of the differences.

8. The method of claim 7 , wherein updating the global automatic speech recognition system further comprises calculating a weights update for the global automatic speech recognition system based on the weighted average of the differences and the global acoustic model weights of the global automatic speech recognition system derived from labelled data.

9. The method of claim 8 , wherein updating the global automatic speech recognition system further comprises controlling a balance of influence between the weighted average of the differences and the global acoustic model weights of the global automatic speech recognition system derived from labelled data.

10. The method of claim 1 , wherein the local acoustic model weights are derived without intervention from a user.

11. A device comprising:

a network interface;

a memory; and

one or more processors coupled to the network interface and the memory, and configured to:

receive, from respective client devices, local acoustic model weights for a neural network-based acoustic model of a local automatic speech recognition system running on the respective client devices, wherein the local acoustic model weights are generated at the respective client devices without labelled data;

update a global automatic speech recognition system based on (a) the local acoustic model weights received from the respective client devices and (b) global acoustic model weights of the global automatic speech recognition system derived from labelled data to obtain an updated global automatic speech recognition system, including controlling an influence of unsupervised data on the updated global automatic speech recognition system by multiplying a first balancing coefficient with the local acoustic model weights and multiplying a second balancing coefficient with the global acoustic model weights, and combining resulting multiplication products thereof, wherein the first balancing coefficient is equal to one minus the second balancing coefficient, and wherein the second balancing coefficient has a value of zero to one; and

send the updated global automatic speech recognition system to the respective client devices to operate as a new local automatic speech recognition system.

12. The device of claim 11 , wherein the one or more processors are configured to select the respective client devices from which to receive the local acoustic model weights.

13. The device of claim 11 , wherein the global automatic speech recognition system is a hybrid Deep Neural Network/Hidden Markov Model (DNN-HMM) automatic speech recognition system.

14. The device of claim 11 , further comprising a language model that is based on one of recurrent neural network modeling or efficient N-Gram statistics.

15. The device of claim 11 , wherein the one or more processors are configured to receive from the respective client devices an indication of a language processed by the respective client devices.

16. The device of claim 11 , wherein the one or more processors are configured to update the global automatic speech recognition system by calculating differences between global acoustic model weights of the global automatic speech recognition system and respective local acoustic model weights received from the respective client devices.

17. The device of claim 16 , wherein the one or more processors are further configured to update the global automatic speech recognition system by calculating a weighted average of the differences.

18. The device of claim 17 , wherein the one or more processors are configured to update the global automatic speech recognition system by calculating a weights update for the global automatic speech recognition system based on the weighted average of the differences and the global acoustic model weights of the global automatic speech recognition system derived from labelled data.

19. A non-transitory computer readable storage media encoded with instructions that, when executed by a processor, cause the processor to:

receive, from respective client devices, local acoustic model weights for a neural network-based acoustic model of a local automatic speech recognition system running on the respective client devices, wherein the local acoustic model weights are generated at the respective client devices without labelled data;

update a global automatic speech recognition system based on (a) the local acoustic model weights received from the respective client devices and (b) global acoustic model weights of the global automatic speech recognition system derived from labelled data to obtain an updated global automatic speech recognition system, including controlling an influence of unsupervised data on the updated global automatic speech recognition system by multiplying a first balancing coefficient with the local acoustic model weights and multiplying a second balancing coefficient with the global acoustic model weights, and combining resulting multiplication products thereof, wherein the first balancing coefficient is equal to one minus the second balancing coefficient, and wherein the second balancing coefficient has a value of zero to one; and

send the updated global automatic speech recognition system to the respective client devices to operate as a new local automatic speech recognition system.

20. The non-transitory computer readable storage media of claim 19 , wherein the global automatic speech recognition system is a hybrid Deep Neural Network/Hidden Markov Model (DNN-HMM) automatic speech recognition system.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded May 19, 2021
From: LE GROUX, SYLVAIN; ZERHOUNI, ERWAN BARRY TARIK
To: CISCO TECHNOLOGY, INC.
Reel/Frame 056288/0421 →
Continuity (1)
Related Publication 20220383857A1 · Dec 1, 2022
References Cited (17)
US 9098560B2 · Fitterer · 2015 [cited by examiner]
US 10176802B1 · Ladhak et al. · 2019 [cited by applicant]
US 20180039883A1 · Kurata · 2018 [cited by applicant]
US 20180240456A1 · Jeong · 2018 [cited by examiner]
US 20190371298A1 · Hannun et al. · 2019 [cited by applicant]
US 20220201021A1 · Joshi · 2022 [cited by examiner]
Ji et al., “Learning Private Neural Language Modeling with Attentive Aggregation,” arXiv:1812.07108v2 [cs.CL], Mar. 13, 2019, available at https://doi.org/10.48550/arXiv.1812.07108 (Year: 2019). [cited by examiner]
Cui et al., “Federated Acoustic Modeling for Automatic Speech Recognition,” IEEE, May 13, 2021, pp. 6748-6752, doi: 10.1109/ICASSP39728.2021.9414305 (Year: 2021). [cited by examiner]
Wahab et al., “Federated Machine Learning: Survey, Multi-Level Classification, Desirable Criteria and Future Directions in Communication and Networking Systems,” in IEEE Communications Surveys & Tutorials, vol. 23, No. … [cited by examiner]
Wang et al., “Addressing Class Imbalance in Federated Learning,” Association for the Advancement of Artificial Intelligence, arXiv:2008.06217v2 [cs.LG], Dec. 15, 2020, available at https://doi.org/10.48550/arXiv.2008.06… [cited by examiner]
Jiang, Di, et al. “A GDPR-compliant Ecosystem for Speech Recognition with Transfer, Federated, and Evolutionary Learning.” ACM Transactions on Intelligent Systems and Technology (2021). [cited by examiner]
H. Brendan McMahan et al., “Communication-Efficient Learning of Deep Networks from Decentralized Data”, Artificial Intelligence and Statistics, PMLR, Apr. 10, 2017, 10 pages. [cited by applicant]
Daniel Povey et al., “Parallel Training of Dnns With Natural Gradient and Parameter Averaging”, ICLR, arXiv:1410.7455v8 [cs.NE], Jun. 22, 2015, 28 pages. [cited by applicant]
Dhruv Guliani et al., “Training Speech Recognition Models With Federated Learning: A Quality/Cost Framework”, ICASSP 2021-2021 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), Jun. 6, 2… [cited by applicant]
Vimal Manohar et al., “Semi-Supervised Training of Acoustic Models Using Lattice-Free MMI”, 2018 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), Apr. 15, 2018, 5 pages. [cited by applicant]
Sheng Li et al. “Semi-Supervised Ensemble DNN Acoustic Model Training”, 2017 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), Mar. 5, 2017, 5 pages. [cited by applicant]
Xiaodong Cui et al., “Federated Acoustic Modeling for Automatic Speech Recognition”, ICASSP 2021-2021 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), Jun. 6, 2021, 5 pages. [cited by applicant]