IP Library Granted Patent US 12,670,911
Granted Patent B2
US 12,670,911 · App. 18/081,382 · Granted Jun 30, 2026

Dialogue system and dialogue processing method

Inventors: Sungwang Kim (Seoul, KR); Jaemin Moon (Yongin, KR); Minjae Park (Seongnam, KR)
Assignees: Hyundai Motor Company; Kia Corporation
G10L15/26G06F40/20G08B3/1008G08B5/222
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,670,911
App. No.
18/081,382
Granted
Jun 30, 2026
Kind
B2
Abstract

Provided is a dialogue system including: a speech recognizer module configured to convert a speech of a user into a plurality of candidate texts, and prioritize the plurality of candidate texts; a understanding module configured to determine a first action corresponding to a first candidate text with a highest priority among the plurality of candidate texts; and a controller configured to attempt to perform the determined first action, and when the first action is not performable, reprioritize other candidate texts of the plurality of candidate texts.

Claims (36)

1 . A dialogue system, comprising:

a speech recognizer module comprising at least one processor configured to execute a speech-to-text engine and convert a speech of a user into a plurality of candidate texts, and prioritize the plurality of candidate texts;

an understanding module in communication with the speech recognizer module, the understanding module comprising the at least one processor configured to execute a natural language understanding engine and determine a first action corresponding to a first candidate text having a highest priority from among the plurality of candidate texts;

a controller in communication with the understanding module and the speech recognizer module, the controller comprising the at least one processor configured to attempt to perform the first action as determined by the understanding module and generate a first action signal for performing the determined first action, and

a communicator configured to transmit the generated first action signal to an external server or a vehicle,

wherein if a failure signal is received from the external server or the vehicle indicating that an operation corresponding to the generated first action signal is not performable, the controller is further configured to reprioritize other candidate texts from among the plurality of candidate texts based on whether a completeness of a sentence of the speech of the user and whether an action corresponding with each of other candidate texts from among the plurality of candidate texts is performable.

2 . The dialogue system of claim 1 , wherein the understanding module is further configured to determine a second action corresponding to a second candidate text having a highest priority from among the reprioritized candidate texts.

3 . The dialogue system of claim 2 , wherein the controller is further configured to attempt to perform the determined second action, and if the second action is performable, generate a guide signal for providing the user with information about the second action.

4 . The dialogue system of claim 3 , wherein the controller is further configured to generate a visual guide signal for visually providing the information about the second action if the user is looking at a display.

5 . The dialogue system of claim 4 , wherein the controller is further configured to: generate a visual guide signal for displaying an incorrect word which is misrecognized in the first candidate text and a corrected word which is correctly recognized in the second candidate text, and

if a speech including the corrected word is input from the user, transmit a second action signal for performing the second action to the external server or the vehicle through the communicator.

6 . The dialogue system of claim 3 , wherein the controller is further configured to generate an audible guide signal for audibly providing the information about the second action if the user is not looking at a display.

7 . The dialogue system of claim 3 , wherein the controller is further configured to generate a visual guide signal for visually providing the information about the second action and an audible guide signal for audibly providing the information about the second action, and the communicator is further configured to transmit the visual guide signal and the audible guide signal to the vehicle.

8 . The dialogue system of claim 1 , wherein the controller is further configured to reprioritize the other candidate texts based on at least one of: an utterance frequency of the user or an utterance frequency of entire users.

9 . A dialogue processing method, comprising:

converting, by at least one processor, a speech of a user into a plurality of candidate texts;

prioritizing, by the at least one processor, the plurality of candidate texts;

determining, by the at least one processor, a first action corresponding to a first candidate text having a highest priority from among the plurality of candidate texts;

attempting, by the at least one processor, to perform the determined first action;

generating, by the at least one processor, a first action signal for performing the determined first action;

transmitting, by a communicator, the generated first action signal to an external server or a vehicle; and

reprioritizing, by the at least one processor, other candidate texts of the plurality of candidate texts if the first action is not performable,

wherein, if a failure signal is received from the external server or the vehicle, the reprioritizing step further comprises reprioritizing the other candidate texts based on whether a completeness of a sentence of the speech of the user and whether an action corresponding with each of other candidate texts from among the plurality of candidate texts is performable, and

wherein the failure signal indicates that an operation corresponding to the generated first action signal is not performable.

10 . The dialogue processing method of claim 9 , further comprising:

determining a second action corresponding to a second candidate text with a highest priority from among the reprioritized candidate texts.

11 . The dialogue processing method of claim 10 , further comprising:

attempting to perform the determined second action; and

generating a guide signal for providing the user with information about the second action if the second action is performable.

12 . The dialogue processing method of claim 11 , wherein the generating of the guide signal step further comprises generating a visual guide signal for visually providing the information about the second action if the user is looking at a display.

13 . The dialogue processing method of claim 12 , wherein the generating of the guide signal step further comprises generating a visual guide signal for displaying a word which is misrecognized in the first candidate text and corrected in the second candidate text, and

when a speech including the corrected word is input from the user, transmitting a second action signal for performing the second action to the external server or the vehicle.

14 . The dialogue processing method of claim 11 , wherein the generating of the guide signal further comprises generating an audible guide signal for audibly providing the information about the second action if the user is not looking at a display.

15 . The dialogue processing method of claim 11 , wherein the generating of the guide signal step further comprises generating a visual guide signal for visually providing the information about the second action and an audible guide signal for audibly providing the information about the second action, and

transmitting the visual guide signal and the audible guide signal to the vehicle.

16 . The dialogue processing method of claim 9 , wherein the reprioritizing step further comprises reprioritizing the other candidate texts based on at least one of: an utterance frequency of the user or an utterance frequency of entire users.