IP Library › Granted Patent US 10,529,337
Granted Patent B2
US 10,529,337 · App. 16/240,928 · Granted Jan 7, 2020

Symbol sequence estimation in speech

Inventors: Kenneth W. Church (Yorktown Heights, NY); Gakuto Kurata (Tokyo, JP); Bhuvana Ramabhadran (Yorktown Heights, NY); Abhinav Sethy (Yorktown Heights, NY); Masayuki Suzuki (Tokyo, JP); Ryuki Tachibana (Tokyo, JP)
Assignee: INTERNATIONAL BUSINESS MACHINES CORPORATION
G10L15/26
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 10,529,337
App. No.
16/240,928
Granted
Jan 7, 2020
Kind
B2
Abstract

Symbol sequences are estimated using a computer-implemented method including detecting one or more candidates of a target symbol sequence from a speech-to-text data, extracting a related portion of each candidate from the speech-to-text data, detecting repetition of at least a partial sequence of each candidate within the related portion of the corresponding candidate, labeling the detected repetition with a repetition indication, and estimating whether each candidate is the target symbol sequence, using the corresponding related portion including the repetition indication of each of the candidates.

Claims (55)

1. A computer-implemented method comprising:

detecting repetitions of at least a partial sequence of related portions of one or more candidates of a target symbol sequence from a speech-to-text data;

labeling the detected repetitions with a repetition indication; and

estimating, by employing a trained estimation model with a plurality of Long Short-Term Memory (LSTM) neural networks, whether each candidate is the target symbol sequence; using the corresponding related portion including the repetition indication of each of the candidates, wherein each candidate is divided into a plurality of word groups, with each of the plurality of word groups being fed into one of a plurality of LSTM neural networks to generate a plurality of outputs, with the plurality of outputs being fed into a softmax layer.

2. The method of claim 1 , wherein the detecting the one or more candidates of the target symbol sequence from the speech-to-text data includes

extracting two or more symbol sequences that constitute each of the candidates, from the speech-to-text data, wherein the two or more symbol sequences are separate from each other in the speech-to-text data.

3. The method of claim 2 , wherein detecting repetition of at least a partial sequence of each candidate within the related portion of the corresponding candidate includes

detecting at least one of the two or more symbol sequences that constitute the corresponding candidate within the related portion of the corresponding candidates.

4. The method of claim 2 , wherein the extracting two or more symbol sequences are performed by extracting the predetermined number of symbol sequences, the two or more symbol sequences do not overlap, and the concatenation of the two or more symbol sequences forms each of the candidates.

5. The method of claim 1 , wherein the related portion of each of the candidates includes a portion adjacent to the each of the candidates.

6. The method of claim 5 , wherein the estimating whether each candidate is the target symbol sequence, based on the repetition indication of each corresponding candidate includes

estimating a probability that each candidate is the target symbol sequence by inputting the related portion of each candidate with the repetition indication into a recurrent neural network.

7. The method of claim 6 , wherein the estimating whether each candidate is the target symbol sequence, based on the repetition indication of each corresponding candidate further includes

determining which candidate outputs the highest probability from the recurrent neural network among the candidates.

8. The method of claim 6 , wherein

the extracting a related portion for each candidate from the speech-to-text data includes extracting a plurality of the related portions of the candidates from the speech-to-text data,

wherein the estimating a probability that each candidate is the target symbol sequence by inputting the related portion of each of the candidates with labelled repetition into a recurrent neural network includes inputting each of the plurality of the related portions of each of the candidates with labelled repetition into a recurrent neural network among a plurality of recurrent neural networks, and

wherein each of the plurality of the related portions of each of the candidates with repetition indications is input into a recurrent neural network among the plurality of recurrent neural networks in a direction depending on a location of each of the plurality of the related portions to the each of the candidates.

9. The method of claim 6 , further comprising:

requiring additional speech-to-text data in response to determining that the probabilities for the candidates are below a threshold.

10. The method of claim 1 , wherein words being recursively fed includes every word of each candidate being fed into each respective LSTM neural network to generate an output and to have each output fed into each respective LSTM neural network along with the word following the word employed to generate the output.

11. The method of claim 1 , wherein the labeling the detected repetition with the repetition indication includes

labeling the detected repetition with an indication of a symbol length of the detected repetition.

12. The method of claim 1 , wherein the labeling the detected repetition with the repetition indication includes

labeling the detected repetition with an indication of a location of the detected repetition in the each candidate.

13. The method of claim 1 , further comprising:

detecting a similar portion that is similar to at least a partial sequence of each of the candidates from the related portion of each of the candidates, and

labeling the detected similar portion with information indicating similarity, and

wherein the estimating whether each candidate is the target symbol sequence, using the corresponding related portion including the repetition indication of each of the candidates includes estimating whether each of the candidates is the target symbol sequence, based on the repetition indication and the similar portion of the each candidate.

14. An apparatus comprising:

a processor; and

one or more computer readable mediums collectively including instructions that, when executed by the processor, cause the processor to perform operations comprising:

detecting repetitions of at least a partial sequence of related portions of one or more candidates of a target symbol sequence from a speech-to-text data, and

labeling the detected repetitions with a repetition indication,

estimating, by employing a trained estimation model with a plurality of Long Short-Term Memory (LSTM) neural networks, whether each candidate is the target symbol sequence, based on the corresponding related portion including the repetition indication of each of the candidates, wherein each candidate is divided into a plurality of word groups, with each of a plurality of word groups being fed into one of a plurality of LSTM neural networks to generate a plurality of outputs, with the plurality of outputs being fed into a softmax layer.

15. The apparatus of claim 14 ,

wherein the detecting the one or more candidates of the target symbol sequence from the speech-to-text data includes

extracting two or more symbol sequences that constitute each of the candidates, from the speech-to-text data, wherein the two or more symbol, sequences are separate from each other in the speech-to-text data.

16. The apparatus of claim 15 , wherein detecting repetition of at least a partial sequence of each candidate within the related portion of the corresponding candidate includes

detecting at least one of the two or more symbol sequences that constitute the corresponding candidate within the related portion of the corresponding candidates.

17. The apparatus of claim 15 , wherein the extracting two or more symbol sequences are performed by extracting the predetermined number of symbol sequences, the two or more symbol sequences do not overlap, and the concatenation of the two or more symbol sequences forms each of the candidates.

18. The apparatus of claim 17 , wherein the related portion of each of the candidates includes a portion adjacent to the each of the candidates.

19. The apparatus of claim 18 , wherein the estimating whether each candidate is the target symbol sequence, based on the repetition indication of each corresponding candidate includes

estimating a probability that each candidate is the target symbol sequence by inputting the related portion of each candidate with the repetition indication into a recurrent neural network.

20. A non-transitory computer readable storage medium having instructions embodied therewith, the instructions executable by a processor or programmable circuitry to cause the processor or programmable circuitry to perform operations comprising:

detecting repetitions of at least a partial sequence of related portions of one or more candidates of a target symbol sequence from a speech-to-text data, and

labeling the detected repetitions with a repetition indication,

estimating, by employing a trained estimation model with a plurality of Long Short-Term Memory (LSTM) neural networks, whether each candidate is the target symbol sequence, based on the corresponding related portion including the repetition indication of each of the candidates, wherein each candidate is divided into a plurality of word groups, with each of the plurality of word groups being fed into one of a plurality of LSTM neural networks to generate a plurality of outputs, with the plurality of outputs being fed into a softmax layer.

21. The non-transitory computer readable storage medium of claim 20 , wherein the detecting the one or more candidates of the target symbol sequence from the speech-to-text data includes

extracting two or more symbol sequences that constitute each of the candidates, from the speech-to-text data, wherein the two or more symbol sequences are separate from each other in the speech-to-text data.

22. The non-transitory computer readable storage medium of claim 21 , wherein detecting repetition of at least a partial sequence of each candidate within the related portion of the corresponding candidate includes

detecting at least one of the two or more symbol sequences that constitute the corresponding candidate within the related portion of the corresponding candidates.

23. The non-transitory computer readable storage medium of claim 21 , wherein the extracting two or more symbol sequences are performed by extracting the predetermined number of symbol sequences, the two or more symbol sequences do not overlap, and the concatenation of the two or more symbol sequences forms each of the candidates.

24. The non-transitory computer readable storage medium of claim 23 wherein the related portion of each of the candidates includes a portion adjacent to the each of the candidates.

25. The non-transitory computer readable storage medium of claim 24 , wherein the estimating whether each candidate is the target symbol sequence, based on the repetition indication of each corresponding candidate includes estimating a probability that each candidate is the target symbol sequence by inputting the related portion of each candidate with the repetition indication into a recurrent neural network.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jan 7, 2019
From: CHURCH, KENNETH W.; KURATA, GAKUTO; RAMABHADRAN, BHUVANA; SETHY, ABHINAV; SUZUKI, MASAYUKI; TACHIBANA, RYUKI
To: INTERNATIONAL BUSINESS MACHINES CORPORATION
Reel/Frame 047916/0148 →
Continuity (2)
Continuation 15409126 · Jan 18, 2017
Related Publication 20190139550A1 · May 9, 2019