IP Library › Granted Patent US 12,451,125
Granted Patent B2
US 12,451,125 · App. 17/652,823 · Granted Oct 21, 2025

Speech recognition apparatus and method

Inventors: Daichi Hayakawa (Inzai, JP); Takehiko Kagoshima (Yokohama, JP)
Assignee: KABUSHIKI KAISHA TOSHIBA
G10L15/18G10L15/063G10L15/20
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,451,125
App. No.
17/652,823
Granted
Oct 21, 2025
Kind
B2
Abstract

According to one embodiment, a speech recognition apparatus includes processing circuitry. The processing circuitry generates a plurality of augmented speech data, based on input speech data, generates a plurality of acoustic scores, based on the plurality of augmented speech data and an acoustic model, generates a plurality of adjusted acoustic scores by resampling the acoustic scores, generates an integrated acoustic score by integrating the adjusted acoustic scores, generates an integrated lattice, based on the integrated acoustic score, a pronunciation dictionary, and a language model, and searches a speech recognition result with a highest likelihood from the integrated lattice.

Claims (48)

1. A speech recognition apparatus comprising processing circuitry configured to:

generate a plurality of augmented speech data, based on input speech data, the plurality of augmented speech data each corresponding to an entirety of the input speech data,

wherein at least one of the plurality of augmented speech data is generated by applying a signal processing operation selected from:

(i) a speech speed conversion in which the input speech data is resampled at a sampling rate different from an original sampling rate and then restored to the original sampling rate using interpolation,

(ii) a sound volume conversion in which an amplitude of a waveform of the input speech data is multiplied by a gain coefficient, and

(iii) a voice quality conversion in which a pitch of the input speech data is altered using a Pitch Synchronous Overlap and Add (PSOLA) method;

generate a plurality of acoustic scores, based on the plurality of augmented speech data and an acoustic model;

generate a plurality of adjusted acoustic scores by resampling the acoustic scores, wherein the resampling comprises, for each acoustic score sequence of K-dimensional vectors over T n time frames, resampling each of the K dimensions independently using interpolation at a sampling rate of T/T n , where T n is a number of time frames of the plurality of augmented speech data, and where T is a number of time frames of the input speech data, thereby converting each acoustic score sequence into a corresponding sequence over T time frames;

generate an integrated acoustic score by calculating, for each time frame, an average value, median, or maximum value across the adjusted acoustic scores;

generate an integrated lattice, based on the integrated acoustic score, a pronunciation dictionary, and a language model;

search a speech recognition result with a highest likelihood from the integrated lattice;

estimate a conversion parameter based on the input speech data, wherein the conversion parameter is at least one of (i) a speech speed conversion parameter based on a comparison between a number of morae per unit time and a reference speech speed, (ii) a sound volume conversion parameter based on a power spectrum comparison, and (iii) a voice quality conversion parameter based on a pitch mean comparison; and

generate at least one of the plurality of augmented speech data by executing a conversion process in which the estimated conversion parameter is uniformly applied to the input speech data using the corresponding signal processing operation.

2. The speech recognition apparatus according to claim 1 , wherein the integrated lattice is a word lattice in which candidate words by speech recognition are nodes, and likelihoods of the candidate words are edges.

3. The speech recognition apparatus according to claim 1 , wherein the processing circuitry is further configured to auto-determine a conversion parameter relating to the conversion process, based on the input speech data.

4. The speech recognition apparatus according to claim 1 , wherein the plurality of augmented speech data include the input speech data.

5. The speech recognition apparatus according to claim 1 , wherein the acoustic model is a single model that is trained to output a posterior probability corresponding to an acoustic score by inputting speech data in units of at least one of a phoneme, a syllable, a character, a word-piece, and a word.

6. The speech recognition apparatus according to claim 1 , wherein the processing circuitry is further configured to generate an adapted acoustic model in which the acoustic model is adapted to a speaker of the input speech data, based on the input speech data and the speech recognition result corresponding to the input speech data.

7. A speech recognition apparatus comprising processing circuitry configured to:

generate a plurality of augmented speech data, based on input speech data, the plurality of augmented speech data each corresponding to an entirety of the input speech data,

wherein at least one of the plurality of augmented speech data is generated by applying a signal processing operation selected from:

(i) a speech speed conversion in which the input speech data is resampled at a sampling rate different from an original sampling rate and then restored to the original sampling rate using interpolation,

(ii) a sound volume conversion in which an amplitude of a waveform of the input speech data is multiplied by a gain coefficient, and

(iii) a voice quality conversion in which a pitch of the input speech data is altered using a Pitch Synchronous Overlap and Add (PSOLA) method;

generate a plurality of acoustic scores, based on the plurality of augmented speech data and an acoustic model;

generate a plurality of adjusted acoustic scores by resampling the acoustic scores, wherein the resampling comprises, for each acoustic score sequence of K-dimensional vectors over T n time frames, resampling each of the K dimensions independently using interpolation at a sampling rate of T/T n , where T n is a number of time frames of the plurality of augmented speech data, and where T is a number of time frames of the input speech data, thereby converting each acoustic score sequence into a corresponding sequence over T time frames;

generate a plurality of lattices, based on the acoustic scores, a pronunciation dictionary, and a language model;

generate an integrated lattice by integrating the lattices;

search a speech recognition result with a highest likelihood from the integrated lattice;

estimate a conversion parameter based on the input speech data, wherein the conversion parameter is at least one of (i) a speech speed conversion parameter based on a comparison between a number of morae per unit time and a reference speech speed, (ii) a sound volume conversion parameter based on a power spectrum comparison, and (iii) a voice quality conversion parameter based on a pitch mean comparison; and

generate at least one of the plurality of augmented speech data by executing a conversion process in which the estimated conversion parameter is uniformly applied to the input speech data using the corresponding signal processing operation.

8. The speech recognition apparatus according to claim 7 , wherein each of the lattices is a word lattice in which candidate words by speech recognition are nodes, and likelihoods of the candidate words are edges.

9. The speech recognition apparatus according to claim 8 , wherein the processing circuitry is further configured to generate the integrated lattice by connecting start points of the lattices and connecting end points of the lattices, and integrating common parts of the candidate words.

10. The speech recognition apparatus according to claim 7 , wherein the processing circuitry is further configured to auto-determine a conversion parameter relating to the conversion process, based on the input speech data.

11. The speech recognition apparatus according to claim 7 , wherein the processing circuitry is further configured to generate an adapted acoustic model in which the acoustic model is adapted to a speaker of the input speech data, based on the input speech data and the speech recognition result corresponding to the input speech data.

12. A speech recognition method comprising:

generating a plurality of augmented speech data, based on input speech data, the plurality of augmented speech data each corresponding to an entirety of the input speech data,

wherein at least one of the plurality of augmented speech data is generated by applying a signal processing operation selected from:

(i) a speech speed conversion in which the input speech data is resampled at a sampling rate different from an original sampling rate and then restored to the original sampling rate using interpolation,

(ii) a sound volume conversion in which an amplitude of a waveform of the input speech data is multiplied by a gain coefficient, and

(iii) a voice quality conversion in which a pitch of the input speech data is altered using a Pitch Synchronous Overlap and Add (PSOLA) method;

generating a plurality of acoustic scores, based on the plurality of augmented speech data and an acoustic model;

generating a plurality of adjusted acoustic scores by resampling the acoustic scores, wherein the resampling comprises, for each acoustic score sequence of K-dimensional vectors over T n time frames, resampling each of the K dimensions independently using interpolation at a sampling rate of T/T n , where T n is a number of time frames of the plurality of augmented speech data, and where T is a number of time frames of the input speech data, thereby converting each acoustic score sequence into a corresponding sequence over T time frames;

generating an integrated acoustic score by calculating, for each time frame, an average value, median, or maximum value across the adjusted acoustic scores;

generating an integrated lattice, based on the integrated acoustic score, a pronunciation dictionary, and a language model; and

searching a speech recognition result with a highest likelihood from the integrated lattice;

estimating a conversion parameter based on the input speech data, wherein the conversion parameter is at least one of (i) a speech speed conversion parameter based on a comparison between a number of morae per unit time and a reference speech speed, (ii) a sound volume conversion parameter based on a power spectrum comparison, and (iii) a voice quality conversion parameter based on a pitch mean comparison; and

generating at least one of the plurality of augmented speech data by executing a conversion process in which the estimated conversion parameter is uniformly applied to the input speech data using the corresponding signal processing operation.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Apr 5, 2022
From: HAYAKAWA, DAICHI; KAGOSHIMA, TAKEHIKO
To: KABUSHIKI KAISHA TOSHIBA
Reel/Frame 059506/0629 →
Priority Claims (1)
JP 2021-091236 · May 31, 2021 · national
Continuity (1)
Related Publication 20220383860A1 · Dec 1, 2022
References Cited (59)
US 5606644A · Chou · 1997 [cited by examiner]
US 7761296B1 · Bakis · 2010 [cited by examiner]
US 9966066B1 · Corfield · 2018 [cited by examiner]
US 10115393B1 · Kumar · 2018 [cited by examiner]
US 10679621B1 · Sundaram · 2020 [cited by examiner]
US 11164592B1 · Wu · 2021 [cited by examiner]
US 20010010037A1 · Imai · 2001 [cited by examiner]
US 20030004721A1 · Zhou · 2003 [cited by examiner]
US 20060293883A1 · Endo · 2006 [cited by examiner]
US 20070073540A1 · Hirakawa · 2007 [cited by examiner]
US 20070100618A1 · Lee · 2007 [cited by examiner]
US 20090018833A1 · Kozat · 2009 [cited by examiner]
US 20130297311A1 · Yamaguchi · 2013 [cited by examiner]
US 20130325456A1 · Takagi · 2013 [cited by examiner]
US 20140372120A1 · Harsham · 2014 [cited by examiner]
US 20160240188A1 · Seto · 2016 [cited by examiner]
US 20170025119A1 · Song · 2017 [cited by examiner]
US 20180166071A1 · Lee et al. · 2018 [cited by applicant]
US 20180350352A1 · Song · 2018 [cited by examiner]
US 20180359580A1 · Aran · 2018 [cited by examiner]
US 20190318742A1 · Srivastava · 2019 [cited by examiner]
US 20190385628A1 · Nakashika · 2019 [cited by examiner]
US 20200066260A1 · Hayakawa · 2020 [cited by examiner]
US 20210020175A1 · Shaq et al. · 2021 [cited by applicant]
US 20210135644A1 · Chon · 2021 [cited by examiner]
US 20240038213A1 · Kanagawa · 2024 [cited by examiner]
CN 102013253A · 2011 [cited by applicant]
CN 109182197A · 2019 [cited by applicant]
CN 112185342A · 2021 [cited by applicant]
JP 9325798A · 1997 [cited by applicant]
JP 2004139049A · 2004 [cited by applicant]
JP 2005221678A · 2005 [cited by applicant]
JP 2007309979A · 2007 [cited by applicant]
JP 201982865A · 2019 [cited by applicant]
JP 202012928A · 2020 [cited by applicant]
JP 2012118441 · 2021 [cited by examiner]
WO WO2015075789A1 · 2015 [cited by applicant]
WO WO2017037830A1 · 2017 [cited by applicant]
Fujimura, H., Ding, N., Hayakawa, D., & Kagoshima, T. (2020). Simultaneous Flexible Keyword Detection and Text-dependent Speaker Recognition for Low-resource Devices. In ICPRAM (pp. 297-307). (Year: 2020). [cited by examiner]
Wegmann, S., & Gillick, L. (2010). Why has (reasonably accurate) Automatic Speech Recognition been so hard to achieve ?. arXiv preprint arXiv:1003.0206. (Year: 2010). [cited by examiner]
Rybach, D., Riley, M., & Schalkwyk, J. (Dec. 2017). On lattice generation for large vocabulary speech recognition. In 2017 IEEE Automatic Speech Recognition and Understanding Workshop (ASRU) (pp. 228-235). IEEE. (Year: … [cited by examiner]
Miranda, J., Neto, J. P., & Black, A. W. (Dec. 2012). Recovery of acronyms, out-of-lattice words and pronunciations from parallel multilingual speech. In 2012 IEEE Spoken Language Technology Workshop (SLT) (pp. 348-353)… [cited by examiner]
Akbacak, M., & Hansen, J. H. (May 2006). Spoken proper name retrieval in audio streams for limited-resource languages via lattice based search using hybrid representations. In 2006 IEEE International Conference on Acous… [cited by examiner]
Akbacak, M., & Hansen, J. H. (2009). Spoken proper name retrieval for limited resource languages using multilingual hybrid representations. IEEE transactions on audio, speech, and language processing, 18(6), 1486-1495. … [cited by examiner]
Ko, T., Peddinti, V., Povey, D., & Khudanpur, S. (Sep. 2015). Audio augmentation for speech recognition. In Interspeech (vol. 2015, p. 3586). (Year: 2015). [cited by examiner]
Japan Office Action issued May 21, 2024 in Japanese Patent Application No. 2021-151465 (with unedited computer-generated English Translation), 6 pages. [cited by applicant]
Japanese Decision to Grant issued May 21, 2024 in Japanese Patent Application No. 2021-091236 (with unedited computer-generated English Translation), 5 pages. [cited by applicant]
Vanis, et al., “Employing Bayesian Networks and Conditional Probability Functions for Determining Dependences in Road Traffic Accidents Data”, 2017 Smart City Symposium Prague (SCSP), 2017, 5 pages. [cited by applicant]
Xu et al., “An Improved Consensus-Like Method For Minimum Bayes Risk Decoding and Lattice Combination”, in Proceedings of ICASSP, 2010, 4 Pages. [cited by applicant]
Qian et al. “Very Deep Convolutional Neural Networks for Robust Speech Recognition”, arXiv:1610.00277v1, 2016, 8 Pages. [cited by applicant]
Rybach et al., “On Lattice Generation for Large Vocabulary Speech Recognition”, IEEE Automatic Speech Recognition and Understanding Workshop (ASRU), 2017, 8 Pages. [cited by applicant]
Le et al., “Word/Sub-Word Lattices Decomposition and Combination for Speech Recognition”, IEEE International Conference on Acoustics, Speech and Signal Processing, 2008, 5 Pages. [cited by applicant]
Moulines et al., “Pitch-Synchronous Waveform Processing Techniques for Text-to-Speech Synthesis Using Diphones”, Speech Communication 9 North-Holland, 1990, 15 Pages. [cited by applicant]
Lahat et al., “A Spectral Autocorrelation Method for Measurement of the Fundamental Frequency of Noise-Corrupted Speech”, in IEEE Transactions on Acoustics, Speech, and Signal Processing, vol. ASSP-35, No. 6, Jun. 1987,… [cited by applicant]
Werbos, “Backpropagation Through Time: What It Does and How To Do It”, Proceedings of the IEEE, vol. 78, No. 10, Oct. 1990, 11 Pages. [cited by applicant]
Lee et al., “Real-Time Word Confidence Scoring Using Local Posterior Probabilities on Tree Trellis Search”, ICASSP, 2004, 4 Pages. [cited by applicant]
Kastanos et al., “Confidence Estimation for Black Box Automatic Speech Recognition Systems Using Lattice Recurrent Neural Networks”, ICASSP, arXiv:1910.11933v2, 2020, 5 Pages. [cited by applicant]
Maekawa, “Corpus of Spontaneous Japanese: Its Design and Evaluation”, In Proceedings ISCA and IEEE Workshop on Spontaneous Speech Processing and Recognition, SSPR, 2003, 6 pages. [cited by applicant]
Combined Chinese Office Action and Search Report issued Apr. 1, 2025 in Chinese Patent Application No. 202210168336.X (with unedited computer-generated English translation), 21 pages. [cited by applicant]