IP Library › Granted Patent US 11,475,891
Granted Patent B2
US 11,475,891 · App. 17/077,943 · Granted Oct 18, 2022

Low delay voice processing system

Inventors: Soonpil Jang (Seoul, KR); Seongjae Jeong (Seoul, KR); Wonkyum Kim (Seoul, KR); Jonghoon Chae (Seoul, KR)
Assignee: LG ELECTRONICS INC.
G10L15/22G10L15/1822G10L15/26H04R29/004G10L2015/088G10L2015/223
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,475,891
App. No.
17/077,943
Granted
Oct 18, 2022
Kind
B2
Abstract

Disclosed is a speech processing method. The speech processing method controls activation timing of a microphone based on a response pattern of the microphone from a user in order to implement a natural conversation. The speech processing device and the NLP system of the present disclosure may be associated with an artificial intelligence module, a drone (or unmanned aerial vehicle (UAV)), a robot, an augmented reality (AR) device, a virtual reality (VR) device, a device related to 5G service, etc.

Claims (43)

1. A speech processing method comprising:

generating information on an activation timing of a microphone, wherein the activation timing corresponds to a time period when a first utterance is received;

tagging utterances with the generated information in response to receiving the first utterance;

activating the microphone by generating a signal in response to an event based at least in part on detecting that the event satisfies the activation timing while the first utterance is provided,

receiving a second utterance through the activated microphone; and

generating a response to the received second utterance based at least in part on extracting at least one of an intent or a subject of the second utterance, wherein generating the response further comprises:

splitting the second utterance into an N number of sub-sections, wherein N is a first number; and

extracting at least one of the intent or the subject from a first extraction section comprising a first sub-section to an M-th sub-section among the N sub-sections, wherein M is a second number smaller than N, wherein the response is generated based on the extracted intent or the extracted subject.

2. The speech processing method of claim 1 , further comprising providing the first utterance by transmitting, to an external terminal comprising a speaker or a display, audio or text related to the first utterance.

3. The speech processing method of claim 2 , wherein the microphone is activated based at least in part on controlling a transceiver to transmit the generated signal to the external terminal.

4. The speech processing method of claim 1 , wherein the first utterance comprises a first informing utterance comprising guide information for multiple services stored in a memory and a second informing utterance generated based at least in part on generating the first utterance.

5. The speech processing method of claim 1 , wherein generating the information on the activation timing further comprises:

analyzing an embedding vector related to the first utterance; and

applying the analyzed embedding vector to a neural network model and determining the activation timing corresponding to the first utterance based on an output of the neural network model.

6. The speech processing method of claim 5 , wherein the neural network model is pre-trained based on log data related to a timing at which the response to the first utterance is received in association with a total output time of the first utterance.

7. The speech processing method of claim 1 , further comprising providing the generated response by transmitting, to an external terminal comprising a speaker or a display, audio or text related to the provided generated response.

8. The speech processing method of claim 1 ,

wherein based on a determination that a reliability of the extracted intent is less than a preset reference value, generating the response further comprises:

extracting at least one of the intent or the subject from a second extraction section comprising the first sub-section to a P-th sub-section among the N sub-sections, wherein P is a third number that is less than or equal to N and greater than M; and

terminating a natural language understanding process based on a determination that a reliability of the extracted intent from the second extraction section is equal to or greater than the preset reference value, wherein the response to the second utterance is generated based on the extracted intent or the extracted subject from the second extraction section.

9. The speech processing method of claim 1 ,

wherein based on a determination that a reliability of the extracted intent is less than a preset reference value, generating the response further comprises extracting at least one of the intent or the subject from a third extraction section including the first sub-section to a K-th sub-section among the N sub-sections, wherein K is a number that is less than or equal to N and greater than M, and

wherein K is increased by 1 based at least in part on a determination that a reliability of the extracted intent from the third extraction section is greater than the preset reference value.

10. The speech processing method of claim 1 , further comprising controlling a transceiver to transmit the generated response to the second utterance to an external terminal comprising a speaker or a display.

11. A speech processing device comprising:

a microphone for receiving a speech signal;

a memory storing multiple first utterances; and

a processor configured to:

generate information on an activation timing of a microphone, wherein the activation timing corresponds to a time period when a first utterance is received,

tag utterances with the generated information in response to receiving the first utterance,

activate the microphone by generating a signal in response to an event when the event based at least in part on detecting that the event satisfies the activation timing while the first utterance is provided,

receive a second utterance through the activated microphone, and

generate a response to the received second utterance based at least in part on receiving at least one of intent or subject of the second utterance, wherein generating the response further comprises:

split the second utterance into an N number of sub-sections, wherein N is a first number; and

extract at least one of the intent or the subject from a first extraction section comprising a first sub-section to an M-th sub-section among the N sub-sections, wherein M is a second number smaller than N, wherein the response is generated based on the extracted intent or the extracted subject.

12. The speech processing device of claim 11 , wherein the processor is further configured to provide the generated response by transmitting, to an external terminal comprising a speaker or a display, audio, or text related to the provided generated response.

13. The speech processing device of claim 11 , wherein based on a determination that a reliability of the extracted intent is less than a preset reference value, generating the response further comprises:

extracting at least one of the intent or the subject from a second extraction section comprising the first sub-section to a P-th sub-section among the N sub-sections, wherein P is a third number that is less than or equal to N and greater than M; and

terminating a natural language understanding process based on a determination that a reliability of the extracted intent from the second extraction section is equal to or greater than the preset reference value, wherein the response to the second utterance is generated based on the extracted intent or the extracted subject from the second extraction section.

14. The speech processing device of claim 11 , wherein based on a determination that a reliability of the extracted intent is less than a preset reference value, generating the response further comprises extracting at least one of the intent or the subject from a third extraction section including the first sub-section to a K-th sub-sections among the N sub-sections, wherein K is a number that is less than or equal to N and greater than M, and

wherein K is increased by 1 based at least in part on a determination that a reliability of the extracted intent from the third extraction section is greater than the preset reference value.

15. The speech processing device of claim 11 , wherein the processor is further configured to control a transceiver to transmit the generated response to the second utterance to an external terminal comprising a speaker or a display.

16. The speech processing device of claim 13 , wherein the first utterance comprises a first informing utterance comprising guide information for multiple services stored in a memory and a second informing utterance generated based at least in part on generating the first utterance.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Oct 22, 2020
From: JANG, SOONPIL; JEONG, SEONGJAE; KIM, WONKYUM; CHAE, JONGHOON
To: LG ELECTRONICS INC
Reel/Frame 054144/0526 →
Priority Claims (1)
KR 10-2020-0012957 · Feb 4, 2020 · national
Continuity (1)
Related Publication 20210241762A1 · Aug 5, 2021
Cited By (1)
US 12,499,889