IP Library › Granted Patent US 11,398,238
Granted Patent B2
US 11,398,238 · App. 16/487,415 · Granted Jul 26, 2022

Speech recognition method in edge computing device

Inventors: Sungjin Kim (Seoul, KR); Dongho Kim (Seoul, KR); Jingyeong Kim (Seoul, KR); Taehyun Kim (Seoul, KR)
Assignee: LG ELECTRONICS INC.
G10L15/30G10L15/063G10L15/183G10L15/34
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,398,238
App. No.
16/487,415
Granted
Jul 26, 2022
Kind
B2
Abstract

Disclosed herein is a speech recognition method in a distributed network environment. A method of performing a speech recognition operation in an edge computing device includes receiving a natural language understanding (NLU) model from the cloud server, storing the received NLU model, receiving voice data spoken by a user from the client device, performing a natural language processing operation on the received voice data using the NLU model, performing speech recognition according to the natural language processing operation, and transmitting a result of the speech recognition to the client device. At least one of the edge computing device, a voice recognition device, and a server may be associated with an artificial intelligence module, a drone (an unmanned aerial vehicle (UAV)), a robot, an augmented reality (AR) device, a virtual reality (VR) device, a device related to a 5G service, and the like.

Claims (71)

1. A speech recognition method in an edge computing device in a distributed network system including a client device, the edge computing device, and a cloud server, the speech recognition method comprising:

receiving voice data spoken by a user from the client device;

based on the voice data being received from the client device, performing an auto speech recognition (ASR) operation;

determining whether a natural language understanding (NLU) model exists on the edge computing device, wherein the NLU model is configured to apply NLU understanding on generated text based on the ASR operation;

based on a determination that the NLU model exists on the edge computing device, performing a natural language processing operation on the received voice data using the NLU model;

based on a determination that the NLU model does not exist on the edge computing device, transmitting the received voice data to the cloud server and receiving, from the cloud server, a result of the natural language processing operation on the received voice data using a previously generated NLU model; and

transmitting a result of the natural language processing operation to the client device.

2. The speech recognition method of claim 1 , wherein

the NLU model is a user-personalized NLU model generated by analyzing a voice command pattern of the user.

3. The speech recognition method of claim 2 , wherein

the NLU model is generated by the cloud server and then received from the cloud server as having a size in a compressed state on the basis of the voice command pattern of the user.

4. The speech recognition method of claim 1 , wherein

the edge computing device comprises an artificial intelligence (AI) module, and

the ASR operation and the natural language processing operation are performed through the AI module.

5. The speech recognition method of claim 1 , further comprising:

performing an initial access procedure with the client device by periodically transmitting a synchronization signal block (SSB);

performing a random access procedure with the client device;

transmitting an uplink (UL) grant to the client device for scheduling transmission of the voice data; and

receiving the voice data from the client device on the basis of the UL grant.

6. The speech recognition method of claim 5 , wherein

the performing of a random access procedure comprises:

receiving a physical random access channel (PRACH) preamble from the client device; and

transmitting a response to the PRACH preamble to the client device.

7. The speech recognition method of claim 6 , further comprising:

performing a downlink beam management (DL BM) procedure using the SSB.

8. The speech recognition method of claim 7 , wherein

the performing of a DL BM procedure comprises:

transmitting a CSI-ResourceConfig IE including a CSI-SSB-ResourceSetList to a user equipment (UE);

transmitting a signal on SSB resources to the client device; and

receiving a best SSBRI and a corresponding RSRP from the client device.

9. The speech recognition method of claim 5 , further comprising:

transmitting configuration information of a reference signal related to beam failure detection to the client device; and

receiving a PRACH preamble requesting beam failure recovery from the client device.

10. A method for generating a natural language understanding (NLU) model in a cloud server in a distributed network system including a client device, an edge computing device, and the cloud server, the method comprising:

receiving voice data spoken by a user from the client device;

defining a rule for generating an NLU model by analyzing a voice pattern of the user from the voice data;

training the NLU model using a plurality of pieces of voice data collected in the cloud server on the basis of the rule, as input data;

generating the NLU model according to a training result;

compressing the NLU model to a predetermined size according to the voice pattern of the user;

transmitting the compressed NLU model to the edge computing device;

based on second voice data being received from the client device, performing an auto speech recognition (ASR) operation;

determining whether the compressed NLU model exists on the edge computing device, wherein the NLU model is configured to apply NLU understanding on generated text based on the ASR operation;

based on a determination that the compressed NLU model exists on the edge computing device, performing a natural language processing operation on the received second voice data using the compressed NLU model; and

based on a determination that the compressed NLU model does not exist on the edge computing device, transmitting the received second voice data to the cloud server and receiving, from the cloud server, a result of the natural language processing operation on the received second voice data using a previously generated NLU model.

11. The method of claim 10 , wherein

the training of the NLU model is performed through automated machine learning AutoML, and

the AutoML comprises:

performing a preprocessing process on the plurality of pieces of voice data collected in the cloud server; and

generating the input data for training the NLU model by classifying the plurality of pieces of voice data collected in the cloud server as a result of the preprocessing according to a predetermined criterion.

12. The method of claim 11 , wherein

the predetermined criterion comprises attribute information of each of the plurality of pieces of voice data collected in the cloud server, and the attribute information includes at least one of age, sex, and area.

13. A speech recognition system comprising:

a cloud server configured to receive first voice data spoken by a user from a client device and to generate a natural language understanding (NLU) model based on a rule defined on the basis of a voice pattern of the user analyzed from the first voice data; and

an edge computing device configured to:

receive second voice data spoken by the user from the client device;

based on the second voice data being received from the client device, performing an auto speech recognition (ASR) operation;

determining whether the NLU model exists on the edge computing device, wherein the NLU model is configured to apply NLU understanding on generated text based on the ASR operation; and

based on a determination that the NLU model exists on the edge computing device, perform a natural language processing operation on the received second voice data using the NLU model; and

based on a determination that the NLU model does not exist on the edge computing device, transmit the received second voice data to the cloud server and receiving, from the cloud server, a result of the natural language processing operation on the received second voice data using a previously generated NLU model.

14. A non-transitory computer-readable medium storing a computer-executable component configured to be executed in one or more processors of a computing device,

wherein the computer-executable component is configured to:

receive voice data spoken by a user from a client device,

define a rule for generating a natural language understanding (NLU) model by analyzing a voice pattern of the user from the voice data,

train the NLU using a plurality of pieces of voice data collected in a cloud server on the basis of the rule as input data,

generate the NLU model according to a training result,

compress the NLU model to a predetermined size according to the voice pattern of the user,

transmit the compressed NLU model to an edge computing device,

based on second voice data being received from the client device, perform an auto speech recognition (ASR) operation;

determine whether the compressed NLU model exists on the edge computing device, wherein the NLU model is configured to apply NLU understanding on generated text based on the ASR operation;

based on a determination that the compressed NLU model exists on the edge computing device, perform a natural language processing operation on the received second voice data using the compressed NLU model; and

based on a determination that the compressed NLU model does not exist on the edge computing device, transmit the received second voice data to the cloud server and receive, from the cloud server, a result of the natural language processing operation on the received second voice data using a previously generated NLU model.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Aug 21, 2019
From: KIM, DONGHO; KIM, TAEHYUN; KIM, SUNGJIN; KIM, JINGYEONG
To: LG ELECTRONICS INC.
Reel/Frame 050117/0976 →
Continuity (1)
Related Publication 20210358502A1 · Nov 18, 2021
Cited By (1)
US 12,499,873