IP Library › Granted Patent US 11,302,331
Granted Patent B2
US 11,302,331 · App. 16/750,274 · Granted Apr 12, 2022

Method and device for speech recognition

Inventors: Dhananjaya N. Gowda (Suwon-si, KR); Kwangyoun Kim (Suwon-si, KR); Abhinav Garg (Suwon-si, KR); Chanwoo Kim (Suwon-si, KR)
Assignee: Samsung Electronics Co., Ltd.
G10L15/26G10L15/02G10L15/28G10L17/04G10L17/20
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,302,331
App. No.
16/750,274
Granted
Apr 12, 2022
Kind
B2
Abstract

Provided are an electronic device for recognizing speech of a user, and a method, performed by the electronic device, of recognizing speech. The method includes obtaining an audio signal based on a speech input based on the audio signal being input, obtaining an output value of a first automatic speech recognition (ASR) model that outputs a character string at a first level; obtaining an output value of a second ASR model that outputs a character string at a second level corresponding to the audio signal based on the output value of the first ASR model based on the audio signal being input; and recognizing the speech from the output value of the second ASR model.

Claims (35)

1. A method, performed by an electronic device, of recognizing speech of a user, the method comprising:

obtaining an audio signal based on a speech input;

based on the audio signal being input, obtaining an output value of a first speech recognition model that outputs a character string at a first level and a first encoded audio signal of the audio signal representing user acoustic information, wherein the first speech recognition model comprises a first encoder and a first decoder, wherein the character string at the first level is obtained by the first decoder based on the first encoded audio signal from the first encoder;

obtaining an output value of a second speech recognition model that outputs a character string at a second level corresponding to the audio signal based on the output value of the first speech recognition model including the character string at the first level and the first encoded audio signal, wherein the second speech recognition model includes a second encoder and a second decoder, wherein the character string at the second level is obtained by the second decoder based on the character string at the first level and a second encoded audio signal, and wherein the second encoded audio signal is obtained by the second encoder based on the first encoded audio signal; and

recognizing the speech from the output value of the second speech recognition model.

2. The method of claim 1 , wherein the character string at the second level comprises sub-sets of a set including, at least one character within the character string at the first level.

3. The method of claim 1 , wherein the character string at the second level comprises sub-strings that are more similar to a semantically-completed word than sub-strings within the character string at the first level.

4. The method of claim 1 , wherein the obtaining of the audio signal comprises:

splitting the audio signal into frames; and

obtaining a feature value of each of the frames of the audio signal.

5. The method of claim 1 , wherein the obtaining of the output value of the second speech recognition model comprises:

applying an attention an output value of the second encoder and an output value of the first decoder; and

obtaining the output value of the second speech recognition model from an output value of the second encoder to which the attention has been applied and an output value of the first decoder to which the attention has been applied.

6. The method of claim 5 , wherein the first decoder comprises a plurality of stacked LSTM layers and an attention layer, wherein the attention layer is configured to apply an attention to an output value of the first encoder based on an output value of the first decoder at a previous time, and

an output value of the first decoder comprises a sequence of context vectors generated by weighted summing the output value of the first encoder based on the attention.

7. The method of claim 5 , wherein based on training of the first speech recognition model for outputting the character string at the first level being completed, the second speech recognition model is trained to output the character string at the second level, based on the output value of the first speech recognition model.

8. The method of claim 1 , wherein each of the first encoder and the second encoder comprises a plurality of stacked long short-term memory (LSTM) layers, and

an output value of the first encoder comprises a sequence of hidden layer vectors respectively output by LSTM layers selected from the plurality of stacked LSTM layers included in the first encoder, and

an output value of the second encoder comprises a sequence of hidden layer vectors output by LSTM layers selected from the plurality of stacked LSTM layers included in the second encoder.

9. An electronic device configured to recognize speech, the electronic device comprising:

a memory storing a program comprising one or more instructions; and

a processor configured to execute the one or more instructions to control the electronic device to:

obtain an audio signal based on a speech input;

based on the audio signal being input, obtain an output value of a first speech recognition model configured to output a character string at a first level and first encoded audio signal of the audio signal representing user acoustic information, wherein the first speech recognition model comprises a first encoder and a first decoder, wherein the character string at the first level is obtained by the first decoder based on the first encoded audio signal from the first encoder;

obtain an output value of a second speech recognition model configured to output a character string at a second level corresponding to the audio signal based on the output value of the first speech recognition model including the character string at the first level and the first encoded audio signal, wherein the second speech recognition model includes a second encoder and a second decoder, wherein the character string at the second level is obtained by the second decoder based on the character string at the first level and a second encoded audio signal, wherein the second encoded audio signal is obtained by the second encoder based on the first encoded audio signal; and

recognize the speech from the output value of the second speech recognition model.

10. The electronic device of claim 9 , wherein the character string at the second level comprises sub-sets of a set including, as an element, at least one character within the character string at the first level.

11. The electronic device of claim 9 , wherein the processor is further configured to execute the one or more instructions to control the electronic device to: split the audio signal into frames and obtain a feature value of each of the frames of the audio signal.

12. The electronic device of claim 9 , wherein the processor is further configured to execute the one or more instructions to control the electronic device to: apply an attention to an output value of the second encoder and an output value of the first decoder, and obtain the output value of the second speech recognition model from an output value of the second encoder to which the attention has been applied and an output value of the first decoder to which the attention has been applied.

13. The electronic device of claim 9 , wherein each of the first encoder and the second encoder comprises a plurality of stacked long short-term memory (LSTM) layers, an output value of the first encoder comprises a sequence of hidden layer vectors respectively output by LSTM layers selected from the plurality of stacked LSTM layers included in the first encoder, and an output value of the second encoder comprises a sequence of hidden layer vectors output by LSTM layers selected from the plurality of stacked LSTM layers included in the second encoder.

14. A non-transitory computer-readable recording medium having recorded thereon a computer program, which, when executed by a computer, performs a method comprising:

obtaining an audio signal based on a speech input;

based on the audio signal being input to an electronic device, obtaining an output value of a first speech recognition model that outputs a character string at a first level and first encoded audio signal of the audio signal representing user acoustic information, wherein the first speech recognition model comprises a first encoder and a first decoder, wherein the character string at the first level is obtained by the first decoder based on the first encoded audio signal from the first encoder;

obtaining an output value of a second speech recognition model that outputs a character string at a second level corresponding to the audio signal based on the output value of the first speech recognition model including the character string at the first level and the first encoded audio signal, wherein the second speech recognition model includes a second encoder and a second decoder, wherein the character string at the second level is obtained by the second decoder based on the character string at the first level and a second encoded audio signal, wherein the second encoded audio signal is obtained by the second encoder based on the first encoded audio signal; and

recognizing the speech from the output value of the second speech recognition model.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jan 23, 2020
From: GOWDA, DHANANJAYA N.; KIM, KWANGYOUN; GARG, ABHINAV; KIM, CHANWOO
To: SAMSUNG ELECTRONICS CO., LTD.
Reel/Frame 051596/0643 →
Priority Claims (1)
KR 10-2019-0159359 · Dec 3, 2019 · national
Continuity (3)
Provisional Application 62848698 · May 16, 2019
Provisional Application 62795736 · Jan 23, 2019
Related Publication 20200234713A1 · Jul 23, 2020
Cited By (4)
US 12,266,373 US 12,505,839 US 12,579,975 US 12,597,413