IP Library › Granted Patent US 12,633,284
Granted Patent B2
US 12,633,284 · App. 18/371,233 · Granted May 19, 2026

Patched multi-condition training for robust speech recognition

Inventors: Pablo Peso Parada (Staines, GB); Agnieszka Dobrowolska (Staines, GB); Karthikeyan Saravanan (Staines, GB); Mete Ozay (Staines, GB)
Assignee: SAMSUNG ELECTRONICS CO., LTD.
G10L15/063G10L21/0216
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,633,284
App. No.
18/371,233
Granted
May 19, 2026
Kind
B2
Abstract

A method of obtaining a patched signal for training a model for use in at least one of a speech and an audio recognition is disclosed. The method comprises obtaining a first signal, wherein the first signal is at least one of a speech and an audio signal, modifying the first signal to obtain at least one second signal, dividing the first signal and the at least one second signal respectively into a plurality of first patches and a plurality of second patches, wherein each one of the plurality of first patches comprises a respective part of the first signal and each one of the plurality of second patches comprises a respective part of the at least one second signal and mixing selected ones of the plurality of first patches and the plurality of second patches to obtain a patched signal.

Claims (22)

1 . A method implemented by a processor of obtaining a patched signal for training a model for use in at least one of a speech and an audio recognition, the method comprising:

obtaining a first signal, wherein the first signal is at least one of a speech and an audio signal;

modifying the first signal to obtain at least one second signal;

dividing the first signal and the at least one second signal respectively into a plurality of first patches and a plurality of second patches, wherein each one of the plurality of first patches comprises a respective part of the first signal and each one of the plurality of second patches comprises a respective part of the at least one second signal;

mixing selected ones of the plurality of first patches and the plurality of second patches to obtain a patched signal; and

training the model for use in at least one of the speech and the audio recognition using the patched signal,

wherein mixing the selected ones of the plurality of first patches and the plurality of second patches comprises randomly selecting respective ones of the first and second patches, and combining the randomly selected first and second patches to obtain the patched signal,

wherein the ones of the first and second patches are randomly selected based on at least one probability threshold corresponding to an identity of the current user.

2 . The method of claim 1 , wherein the ones of the first and second patches are randomly selected according to the at least one probability threshold, such that the at least one probability threshold defines a probability of one of the plurality of first patches being selected for a given part of the patched signal and a probability of one of the plurality of second patches being selected for the given part of the patched signal.

3 . The method of claim 2 , wherein the at least one probability threshold is set in dependence on the identity of the current user, such that different values of the at least one probability threshold may be set for different users.

4 . The method of claim 3 , wherein the at least one probability threshold is set in dependence on historical data associated with the identity of the current user.

5 . The method of claim 1 , wherein modifying the first signal comprises convolving the first signal with a distortion function.

6 . The method of claim 5 , wherein the distortion function is a room impulse response, RIR, function.

7 . The method of claim 6 , wherein the RIR function introduces a delay in the at least one second signal relative to the first signal, the method further comprising:

removing the delay from the at least one second signal prior to obtaining the patched signal.

8 . The method of claim 1 , wherein modifying the first signal comprises adding noise into the first signal.

9 . The method of claim 8 , wherein the noise comprises random noise.

10 . The method of claim 8 , wherein the noise comprises one or more recordings of background noise samples.

11 . The method of claim 8 , wherein the noise comprises one or more synthesized noise samples.

12 . The method of claim 1 , wherein the at least one second signal comprises a plurality of second signals, and modifying the first signal to obtain the plurality of second signals comprises applying different modifications to the first signal to obtain respective ones of the plurality of second signals.

13 . The method of claim 1 , wherein each of the first and second patches have the same length.

14 . The method of claim 1 , wherein each of the first and second patches has a randomly determined length.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Sep 21, 2023
From: PESO PARADA, PABLO; DOBROWOLSKA, AGNIESZKA; SARAVANAN, KARTHIKEYAN; OZAY, METE
To: SAMSUNG ELECTRONICS CO., LTD.
Reel/Frame 064987/0377 →
Priority Claims (2)
GB 2203733 · Mar 17, 2022 · national
GB 2300844 · Jan 19, 2023 · national
Continuity (2)
Continuation PCTKR2023002236 · Feb 16, 2023
Related Publication 20240013775A1 · Jan 11, 2024
References Cited (70)
US 8880410B2 · Nagel et al. · 2014 [cited by applicant]
US 10014000B2 · Nagel et al. · 2018 [cited by applicant]
US 10529317B2 · Lee et al. · 2020 [cited by applicant]
US 10971142B2 · Sriram et al. · 2021 [cited by applicant]
US 20080215322A1 · Fischer et al. · 2008 [cited by applicant]
US 20100318354A1 · Seltzer et al. · 2010 [cited by applicant]
US 20110238416A1 · Seltzer et al. · 2011 [cited by applicant]
US 20130090934A1 · Nagel et al. · 2013 [cited by applicant]
US 20150170663A1 · Disch et al. · 2015 [cited by applicant]
US 20160358602A1 · Krishnaswamy et al. · 2016 [cited by applicant]
US 20170125012A1 · Kanthak et al. · 2017 [cited by applicant]
US 20170200446A1 · Cui et al. · 2017 [cited by applicant]
US 20180350379A1 · Wung et al. · 2018 [cited by applicant]
US 20180350387A1 · Nagel et al. · 2018 [cited by applicant]
US 20180367674A1 · Schalk-Schupp et al. · 2018 [cited by applicant]
US 20190206394A1 · Ichikawa et al. · 2019 [cited by applicant]
US 20200051580A1 · Seo et al. · 2020 [cited by applicant]
US 20200135179A1 · Yang et al. · 2020 [cited by applicant]
US 20210035563A1 · Cartwright et al. · 2021 [cited by applicant]
US 20210043186A1 · Nagano et al. · 2021 [cited by applicant]
US 20210065681A1 · Soni et al. · 2021 [cited by applicant]
US 20210142815A1 · Bryan · 2021 [cited by examiner]
US 20210343274A1 · Kang et al. · 2021 [cited by applicant]
US 20220101828A1 · Fukutomi et al. · 2022 [cited by applicant]
CN 109741736A · 2019 [cited by applicant]
EP 3166105A1 · 2017 [cited by applicant]
JP 2021105684A · 2021 [cited by applicant]
JP 2021135314A · 2021 [cited by applicant]
WO 2020219971A1 · 2020 [cited by applicant]
Kim et al. “SpecMix : A Mixed Sample Data Augmentation method for Training with Time-Frequency Domain Features”, maskarXiv: 2108.03020v1 [cs.SD] Aug. 6, 2021 (Year: 2021). [cited by examiner]
Smyth et al. “A Virtual Acoustic Film Dubbing Stage”, Spatial Audio, Sense the Space of Sound, AES 40th Conference, Tokyo, 2010 (Year: 2010). [cited by examiner]
Communication dated Jul. 19, 2023, issued by the United Kingdom Patent Office for United Kingdom Patent Application No. 2300844.4. [cited by applicant]
E. Tsunoo, K. Shibata, C. Narisetty, Y. Kashiwagi, and S. Watanabe, “Data augmentation methods for end-to-end speech recognition on distant-talk scenarios”, INTERSPEECH 2021, 2021, 5 pages. [cited by applicant]
C. Richey, M. A. Barrios, Z. Armstrong, C. Bartels, H. Franco, M. Graciarena, A. Lawson, M. K. Nandwana, A. R. Stauffer, J. van Hout, P. Gamble, J. Hetherly, C. Stephenson, and K. Ni, “Voices obscured in complex environ… [cited by applicant]
D. S. Park, W. Chan, Y. Zhang, C. Chiu, B. Zoph, E. D. Cubuk, and Q. V. Le, “Specaugment: A simple data augmentation method for automatic speech recognition”, in INTERSPEECH 2019, ISCA, 2019, pp. 2613-2617. [cited by applicant]
T. Ko, V. Peddinti, D. Povey, and S. Khudanpur, “Audio augmentation for speech recognition”, in INTERSPEECH, ISCA, 2015, pp. 3586-3589. [cited by applicant]
N. Jaitly and G. E. Hinton, “Vocal tract length perturbation (VTLP) improves speech recognition”, Proceedings of the 30 th International Conference on Machine Learning, Speech and Language, vol. 28, 2013, 5 pages. [cited by applicant]
C. Kim, M. Shin, A. Garg, and D. Gowda, “Improved vocal tract length perturbation for a state-of-the-art end-to-end speech recognition system”, in Interspeech 2019, ISCA, 2019, pp. 739-743, http://dx.doi.org/10.21437/In… [cited by applicant]
V. Peddinti, D. Povey, and S. Khudanpur, “A time delay neural network architecture for efficient modeling of long temporal contexts”, in INTERSPEECH, ISCA, 2015, pp. 3214-3218. [cited by applicant]
A. Jain, P. R. Samala, D. Mittal, P. Jyothi, and M. Singh, “SPLICEOUT: A simple and efficient audio augmentation method”, CoRR, vol. abs/2110.00046, Oct. 13, 2021, 24 pages, arXiv:2110.00046v2 [cs.SD]. [cited by applicant]
H. Wang, Y. Zou, and W. Wang, “Specaugment++: A hidden space data augmentation method for acoustic scene classification”, CoRR, vol. abs/2103.16858, Jun. 15, 2021, 5 pages, arXiv:2103.16858v3 [eess.AS]. [cited by applicant]
X. Song, Z. Wu, Y. Huang, D. Su, and H. Meng, “SpecSwap: A simple data augmentation method for end-to-end speech recognition”, in INTERSPEECH 2020, ISCA, 2020, pp. 581-585, http://dx.doi.org/10.21437/Interspeech.2020-22… [cited by applicant]
A. N. Carr, Q. Berthet, M. Blondel, O. Teboul, and N. Zeghidour, “Self-supervised learning of audio representations from permutations with differentiable ranking”, IEEE Signal Process. Lett., vol. 28, pp. 708-712, Mar. … [cited by applicant]
L. Meng, J. Xu, X. Tan, J. Wang, T. Qin, and B. Xu, “Mixspeech: Data augmentation for low-resource automatic speech recognition,” in ICASSP. IEEE, Feb. 25, 2021, pp. 7008-7012, arXiv:2102.12664v1 [cs.CL]. [cited by applicant]
H. Zhang, M. Cissé, Y. N. Dauphin, and D. Lopez-Paz, “mixup: Beyond empirical risk minimization”, in ICLR 2018, Open-Review.net, 2018, 13 pages. [cited by applicant]
Y. Tokozume, Y. Ushiku, and T. Harada, “Learning from between-class examples for deep sound recognition”, in ICLR 20108, 2018, 13 pages, https://github.com/mil-tokyo/bc_learning_sound/. [cited by applicant]
T. K. Lam, M. Ohta, S. Schamoni, and S. Riezler, “On-the-fly aligned data augmentation for sequence-to-sequence ASR”, INTERSPEECH 2021, ISCA, 2021, 5 pages, http://dx.doi.org/10.21437/Interspeech.2021-1679. [cited by applicant]
T. Nguyen, S. Stüker, J. Niehues, and A. Waibel, “Improving sequence-to-sequence speech recognition training with on-the-fly data augmentation”, in ICASSP. IEEE, Feb. 3, 2020, pp. 7689-7693, arXiv:1910.13296v2 [eess.AS]. [cited by applicant]
D. Yu, M. L. Seltzer, J. Li, J. Huang, and F. Seide, “Feature learning in deep neural networks—Studies on speech recognition tasks”, in ICLR, Jan. 16, 2013, 9 pages, arXiv:1301.3605v1 [cs.LG]. [cited by applicant]
A. Y. Hannun, C. Case, J. Casper, B. Catanzaro, G. Diamos, E. Elsen, R. Prenger, S. Satheesh, S. Sengupta, A. Coates, and A. Y. Ng, “Deep speech: Scaling up end-to-end speech recognition”, CoRR, vol. abs/1412.5567, Dec.… [cited by applicant]
V. A. Trinh, H. S. Kavaki, and M. I. Mandel, “Importantaug: a data augmentation agent for speech”, Proc. ICASSP 2022, Feb. 19, 2022, 5 pages, arXiv:2112.07156v2 [eess.AS]. [cited by applicant]
A. Sriram, H. Jun, Y. Gaur, and S. Satheesh, “Robust speech recognition using generative adversarial networks”, in ICASSP. IEEE, Nov. 5, 2018, pp. 5639-5643, imarXiv: 1711.01567v1 [cs.CL]. [cited by applicant]
E. Tsunoo, K. Shibata, C. Narisetty, Y. Kashiwagi, and S. Watanabe, “Data augmentation methods for end-to-end speech recognition on distant-talk scenarios”, CoRR, vol. abs/2106.03419, Jun. 7, 2021, 5 pages, arXiv:2106.0… [cited by applicant]
J. B. Allen and D. A. Berkley, “Image method for efficiently simulating small-room acoustics”, The Journal of the Acoustical Society of America, vol. 65, No. 4, pp. 943-950, Apr. 1979. [cited by applicant]
E. A. Lehmann, A. M. Johansson, and S. Nordholm, “Reverberation-time prediction method for room impulse responses simulated with the image-source model”, in 2007 IEEE Workshop on Applications of Signal Processing to Aud… [cited by applicant]
T. Ko, V. Peddinti, D. Povey, M. L. Seltzer, and S. Khudanpur, “A study on data augmentation of reverberant speech for robust speech recognition”, in 2017 IEEE International Conference on Acoustics, Speech and Signal Pr… [cited by applicant]
C. Kim, A. Misra, K. Chin, T. Hughes, A. Narayanan, T. Sainath, and M. Bacchiani, “Generation of large-scale simulated utterances in virtual rooms to train deep-neural networks for far-field speech recognition in google… [cited by applicant]
J. Li, L. Deng, Y. Gong, and R. Haeb-Umbach, “An overview of noise-robust automatic speech recognition”, IEEE Acm Trans. Audio Speech Lang. Process., vol. 22, No. 4, pp. 745-777, Apr. 2014. [cited by applicant]
D. S. Park, Y. Zhang, C.-C. Chiu, Y. Chen, B. Li, W. Chan, Q. V. Le, and Y. Wu, “Specaugment on large scale datasets”, in ICASSP 2020-2020 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP… [cited by applicant]
J. Kim, M. Kumar, D. Gowda, A. Garg, and C. Kim, “A comparison of streaming models and data augmentation methods for robust speech recognition”, Nov. 19, 2021, 7 pages, augmenarXiv:2111.10043v1 [eess.AS]. [cited by applicant]
S. Barreda, “Perceptual validation of vowel normalization methods for variationist research”, Language Variation and Change, vol. 33, No. 1, p. 27-53, 2021, doi: 10.1017=S0954394521000016. [cited by applicant]
C. Fefferman, S. Mitter, and H. Narayanan, “Testing the manifold hypothesis,” Journal of the American Mathematical Society, vol. 29, No. 4, pp. 983-1049, Oct. 2016, http://dx.doi.org/10.1090/jams/852. [cited by applicant]
C. Vaz and S. Narayanan, “Learning a speech manifold for signal subspace speech denoising”, in Interspeech 2015, Dresden, Germany, Sep. 2015, pp. 1735-1739. [cited by applicant]
A. Sinha, K. Ayush, J. Song, B. Uzkent, H. Jin, and S. Ermon, “Negative data augmentation”, in International Conference on Learning Representations, Feb. 9, 2021, 17 pages, arXiv:2102.05113v1 [cs.CV]. [cited by applicant]
M. Ravanelli, T. Parcollet, P. Plantinga, A. Rouhe, S. Cornell, L. Lugosch, C. Subakan, N. Dawalatabad, A. Heba, J. Zhong, J.-C. Chou, S.-L. Yeh, S.-W. Fu, C.-F. Liao, E. Rastorgueva, F. Grondin, W. Aris, H. Na, Y. Gao,… [cited by applicant]
A. Gulati, J. Qin, C.-C. Chiu, N. Parmar, Y. Zhang, J. Yu, W. Han, S. Wang, Z. Zhang, Y. Wu, and R. Pang, “Conformer: Convolution-augmented transformer for speech recognition”, May 16, 2020, 5 pages, improvearXiv:2005.0… [cited by applicant]
V. Panayotov, G. Chen, D. Povey, and S. Khudanpur, “Librispeech: an ASR corpus based on public domain audio books”, in 2015 IEEE international conference on acoustics, speech and signal processing (ICASSP), IEEE, 2015, … [cited by applicant]
P. Peso Parada, D. Sharma, J. Lainez, D. Barreda, T. v. Waterschoot, and P. A. Naylor, “A Single-Channel Non-Intrusive C50 Estimator Correlated With Speech Recognition Performance”, IEEE/ACM Transactions on Audio, Speec… [cited by applicant]
D. Snyder, G. Chen, and D. Povey, “MUSAN: A Music, Speech, and Noise Corpus”, Center for Language and Speech Processing, The Johns Hopkins University, Oct. 28, 2015, 4 pages, arXiv:1510.08484v1 [cs.SD]. [cited by applicant]
Alex Graves et al., “Connectionist temporal classification: Labelling unsegmented sequence data with recurrent neural networks”, In Proceedings of the International Conference on Machine Learning, ICML 2006, 2006, pp. 3… [cited by applicant]