IP Library › Granted Patent US 11,367,438
Granted Patent B2
US 11,367,438 · App. 16/490,020 · Granted Jun 21, 2022

Artificial intelligence apparatus for recognizing speech of user and method for the same

Inventors: Jaehong Kim (Seoul, KR); Hyoeun Kim (Seoul, KR); Hangil Jeong (Seoul, KR); Heeyeon Choi (Seoul, KR)
Assignee: LG ELECTRONICS INC.
G10L15/22G06N3/08G10L15/26G10L25/30G10L2015/225
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,367,438
App. No.
16/490,020
Granted
Jun 21, 2022
Kind
B2
Abstract

An embodiment of the present invention provides an artificial intelligence (AI) apparatus for recognizing a speech of a user, the artificial intelligence apparatus includes a memory to store a speech recognition model and a processor to obtain a speech signal for a user speech, to convert the speech signal into a text using the speech recognition model, to measure a confidence level for the conversion, to perform a control operation corresponding to the converted text if the measured confidence level is greater than or equal to a reference value, and to provide feedback for the conversion if the measured confidence level is less than the reference value.

Claims (35)

1. An artificial intelligence (AI) apparatus for recognizing a speech of a user, the artificial intelligence apparatus comprising:

a memory configured to store a speech recognition model; and

a processor configured to:

obtain a speech signal for a user speech,

convert the speech signal into a text using the speech recognition model,

measure a confidence level for the conversion of the converted text,

based on the measured confidence level being equal to or greater than a reference value, perform a control operation corresponding to the converted text, and

based on the measured confidence level being less than the reference value,

obtain information on a cause of lowering the measured confidence level,

extract a data feature set including a plurality of first features from the speech signal,

determine at least one abnormal feature among the extracted data feature set by determining an abnormal degree for each plurality of first features based on an extent that each feature deviates from a normal range using an abnormal feature determination model, wherein an abnormal feature is determined as the cause of lowering the confidence level, wherein the abnormal feature determination model includes information on threshold ranges for the plurality of first features and rank information on a rank of the corresponding threshold range to each of the plurality of first features, wherein the rank of the corresponding threshold range is determined to be higher as an influence on the confidence level is higher,

determine a rank between abnormal features based on the determined abnormal degree for each plurality of first features,

generate feedback based on the obtained information, the determined abnormal feature, and the rank between the abnormal features, and

provide the feedback for the conversion of the converted text.

2. The AI apparatus of claim 1 , wherein the processor is configured to generate the feedback including a notification of information on the abnormal feature as the cause of lowering the confidence level.

3. The AI apparatus of claim 1 , wherein the processor is configured to generate the feedback including a suggestion of a manner of changing the abnormal feature to a normal feature to enhance the confidence level.

4. The AI apparatus of claim 1 , wherein the data feature set includes at least one of a single speech source state, a speech level, a noise level, a signal to noise ratio (SNR), a speech speed, a word number, a word length, a clipping existence state, or a clipping ratio.

5. The AI apparatus of claim 1 , wherein the processor is configured to:

extract a recognition feature set including a plurality of second features from the speech signal, and

determine the confidence level by using a confidence level measurement model and the plurality of second features,

wherein the confidence level measurement model is a model to output a confidence level for the corresponding speech signal when the plurality of second features are input.

6. The AI apparatus of claim 5 , wherein the confidence level measurement model is an artificial neural network learned based on a machine learning algorithm or a deep learning algorithm and is learned to reduce a difference between a value, which is output when the recognition feature set extracted from a training speech signal is input, and a confidence level, which is previously provided, corresponding to the training speech signal.

7. A method for recognizing a user speech, the method comprising:

obtaining a speech signal for the user speech;

converting the speech signal into a text using a speech recognition model;

measuring a confidence level for the conversion of the converted text;

based on the measured confidence level being equal to or greater than a reference value, performing a control operation corresponding to the converted text; and

based on the measured confidence level being less than the reference value,

obtaining information on a cause of lowering the measured confidence level,

extracting a data feature set including a plurality of first features from the speech signal,

determining at least one abnormal feature among the extracted data feature set by determining an abnormal degree for each plurality of first features based on an extent that each feature deviates from a normal range using an abnormal feature determination model, wherein an abnormal feature is determined as the cause of lowering the confidence level, wherein the abnormal feature determination model includes information on threshold ranges for the plurality of first features and rank information on a rank of the corresponding threshold range to each of the plurality of first features, wherein the rank of the corresponding threshold range is determined to be higher as an influence on the confidence level is higher,

determining a rank between abnormal features based on the determined abnormal degree for each plurality of first features,

generating feedback based on the obtained information, the determined abnormal feature, and the rank between the abnormal features, and

providing the feedback for the conversion of the converted text.

8. A non-transitory recording medium having recorded thereon a program for performing a method for recognizing a user speech, the method comprising: obtaining a speech signal for the user speech; converting the speech signal into a text using a speech recognition model; measuring a confidence level for the conversion of the converted text; based on the measured confidence level being equal to or greater than a reference value, performing a control operation corresponding to the converted text; and based on the measured confidence level being less than the reference value, obtaining information on a cause of lowering the measured confidence level, extracting a data feature set including a plurality of first features from the speech signal, determining at least one abnormal feature among the extracted data feature set by determining an abnormal degree for each plurality of first features based on an extent that each feature deviates from a normal range using an abnormal feature determination model, wherein an abnormal feature is determined as the cause of lowering the confidence level, wherein the abnormal feature determination model includes information on threshold ranges for the plurality of first features and rank information on a rank of the corresponding threshold range to each of the plurality of first features, wherein the rank of the corresponding threshold range is determined to be higher as an influence on the confidence level is higher, determining a rank between abnormal features based on the determined abnormal degree for each plurality of first features, generating feedback based on the obtained information, the determined abnormal feature, and the rank between the abnormal features, and providing the feedback for the conversion of the converted text.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Aug 29, 2019
From: KIM, JAEHONG; KIM, HYOEUN; JEONG, HANGIL; CHOI, HEEYEON
To: LG ELECTRONICS INC.
Reel/Frame 050217/0381 →
Continuity (1)
Related Publication 20210407503A1 · Dec 30, 2021