IP Library › Granted Patent US 11,393,470
Granted Patent B2
US 11,393,470 · App. 16/739,811 · Granted Jul 19, 2022

Method and apparatus for providing speech recognition service

Inventor: Da Hae Kim (Anyang-si, KR)
Assignee: LG ELECTRONICS INC.
G10L15/22G06F3/167G06F40/295G06N20/00G10L15/16G10L15/30G10L2015/223
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,393,470
App. No.
16/739,811
Granted
Jul 19, 2022
Kind
B2
Abstract

Disclosed are a method for providing a speech recognition service and a speech recognition apparatus, which may perform speech recognition by executing an artificial intelligence (AI) algorithm and/or a machine learning algorithm, which are mounted therein, so that a speech recognition apparatus and a server may communicate with each other in a 5G communication environment. The method and the speech recognition apparatus provide a response based on a user's intention analysis with respect to the ambiguous utterance of the user.

Claims (71)

1. A method for providing a speech recognition service, comprising:

receiving a speech input of a user;

obtaining a plurality of candidate actions extracted from the speech input;

deciding relevance between the speech input and each candidate action of the plurality of candidate actions based on current context information of the user; and

deciding a final action of the plurality of candidate actions based on the relevance,

wherein the deciding the relevance comprises:

deciding a weight of each candidate action for each type of each context information by analyzing accumulated context information with respect to the user; and

calculating the relevance by combining weights for each candidate action,

wherein the deciding the weight comprises:

deciding a frequency of having performed each candidate action by analyzing the accumulated context information with respect to the user; and

deciding the weight of each candidate action with respect to each type of each context information based on the frequency.

2. The method of claim 1 ,

wherein the obtaining the plurality of candidate actions comprises:

transmitting the speech input to an external server; and

receiving the plurality of candidate actions extracted by performing a natural language processing for a text representing the speech input by the external server.

3. The method of claim 1 ,

wherein each candidate action of the plurality of candidate actions comprises a named entity that is a target of the action and domain information having classified a category of the action.

4. The method of claim 3 ,

wherein each candidate action of the plurality of candidate actions are distinguished from each other by a combination of the named entity and the domain information.

5. The method of claim 1 ,

wherein the current context information comprises at least one of information on a time of the speech input uttered by the user, a place of the speech input uttered by the user, whether the user is moving, a moving speed of the user, or a device being used by the user.

6. The method of claim 1 ,

wherein the accumulated context information with respect to the user comprises at least one of information on a time received by a command, a place received by the command, a frequency performed by an action, whether the user was moving when the command was received, a moving speed of the user when the command is received, or a device used in receiving the command, for each action performed by the user's command.

7. The method of claim 1 ,

wherein a first type of context information is time information of the speech input uttered by the user,

wherein the deciding the frequency comprises deciding the frequency by counting a number of times of each candidate action performed by the user for each predetermined time unit, and

wherein the deciding the weight of each candidate action with respect to each type of each context information based on the frequency comprises:

normalizing the number of times that has performed each candidate action with respect to an entire time period; and

deciding a result value of the normalization corresponding to the time information of having uttered the speech input from the current context information as a weight for the first type of each candidate action.

8. The method of claim 1 ,

wherein a second type of context information is place information of the speech input uttered by the user,

wherein the deciding the frequency comprises deciding the frequency by counting a number of times of each candidate action performed by the user in at least one place where the user stays for a predetermined time or more, and

wherein the deciding the weight of each candidate action with respect to each type of each context information based on the frequency comprises:

normalizing the number of times that has performed each candidate action with respect to all places recorded in the accumulated context information; and

deciding a result value of the normalization corresponding to a place of having uttered the speech input from the current context information as a weight for the second type of each candidate action.

9. The method of claim 1 ,

wherein a third type of context information is frequency information of the action performed by the user,

wherein the deciding the frequency comprises deciding the frequency for each action performed by the user by counting a number of times performed for each action performed by the user for a predetermined time period, and

wherein the deciding the weight of each candidate action with respect to each type of each context information based on the frequency comprises:

normalizing the frequency for each action; and

deciding a result value of the normalization corresponding to each candidate action as a weight for the third type of each candidate action.

10. The method of claim 1 ,

wherein the deciding the relevance is repeatedly performed for each predetermined time period.

11. The method of claim 1 ,

wherein the deciding the final action comprises deciding a candidate action having a maximum relevance as the final action.

12. The method of claim 11 ,

wherein, when the candidate action having the maximum relevance is in plural, the deciding the final action comprises deciding the final action among the plurality of candidate actions having the maximum relevance according to a predetermined priority.

13. A speech recognition apparatus, comprising:

a microphone configured to receive a speech input of a user; and

a processor configured to decide one of a plurality of candidate actions extracted from the speech input as a final action,

wherein the processor is configured to perform:

an operation of deciding relevance between the speech input and each candidate action of the plurality of candidate actions based on current context information of the user, and

an operation of deciding the final action of the plurality of candidate actions based on the relevance,

wherein the operation of deciding the relevance comprises:

an operation of deciding a weight of each candidate action for each type of each context information by analyzing accumulated context information with respect to the user; and

an operation of calculating the relevance by combining weights for each candidate action, and

wherein the operation of deciding the weight comprises:

an operation of deciding a frequency of having performed each candidate action by analyzing the accumulated context information with respect to the user; and

an operation of deciding the weight of each candidate action with respect to each type of each context information based on the frequency.

14. The speech recognition apparatus of claim 13 ,

wherein each candidate action of the plurality of candidate actions are distinguished from each other by a combination of a named entity that is a target of the action and domain information having classified a category of the action.

15. The speech recognition apparatus of claim 13 ,

wherein the operation of calculating the relevance comprises an operation of summing the weights for each candidate action.

16. The speech recognition apparatus of claim 13 ,

wherein the current context information comprises at least one of information on a time of the speech input uttered by the user, a place of the speech input uttered by the user, whether the user is moving, a moving speed of the user, or a device being used by the user.

17. The speech recognition apparatus of claim 13 ,

wherein the accumulated context information with respect to the user comprises at least one of information on a time received by a command, a place received by the command, a frequency performed by an action, whether the user was moving when the command was received, a moving speed of the user when the command is received, or a device used in receiving the command, for each action performed by the user's command.

18. The speech recognition apparatus of claim 13 ,

wherein the operation of deciding the final action comprises an operation of deciding a candidate action having a maximum relevance as the final action.

19. The speech recognition apparatus of claim 18 ,

wherein, when the candidate action having the maximum relevance is in plural, the operation of deciding the final action comprises an operation of deciding the final action among the plurality of candidate actions having the maximum relevance according to a predetermined priority.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jan 13, 2020
From: KIM, DA HAE
To: LG ELECTRONICS INC.
Reel/Frame 051496/0572 →
Priority Claims (1)
WO PCT/KR2019/011068 · Aug 29, 2019 · international
Continuity (1)
Related Publication 20210065704A1 · Mar 4, 2021
Cited By (3)
US 12,216,690 US 12,718,811 US 12,731,583