IP Library › Granted Patent US 10,319,379
Granted Patent B2
US 10,319,379 · App. 15/704,691 · Granted Jun 11, 2019

Methods and systems for voice dialogue with tags in a position of text for determining an intention of a user utterance

Inventors: Atsushi Ikeno (Kyoto, JP); Yusuke Jinguji (Hiroo-gun, JP); Toshifumi Nishijima (Kasugai, JP); Fuminori Kataoka (Nisshin, JP); Hiromi Tonegawa (Okazaki, JP); Norihide Umeyama (Nisshin, JP)
Assignee: TOYOTA JIDOSHA KABUSHIKI KAISHA
G10L15/22G10L13/08G10L15/1815G10L15/30G10L2015/223
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 10,319,379
App. No.
15/704,691
Granted
Jun 11, 2019
Kind
B2
Abstract

A voice dialog system includes: a voice input unit which acquires a user utterance, an intention understanding unit that interprets an intention of utterance of a voice acquired by the voice input unit, a dialog text creator that creates a text of a system utterance, and a voice output unit that outputs the system utterance as voice data. When creating a text of a system utterance, the dialog text creator creates the text by inserting a tag in a position in the system utterance. The intention understanding unit interprets an utterance intention of a user in accordance with whether a timing at which the user utterance is made is before or after an output of a system utterance at a position corresponding to the tag from the voice output unit.

Claims (33)

1. A voice dialogue system, comprising:

a voice input unit configured to acquire a user utterance;

an intention understanding unit configured to interpret an intention of utterance of a voice acquired by the voice input unit;

a dialogue text creator configured to create a text of a system utterance; and

a voice output unit configured to output the system utterance as voice data, wherein

the dialogue text creator is further configured to create the text of a system utterance by inserting a tag in a position in the system utterance, and

the intention understanding unit is interpret an utterance intention of a user in accordance with whether a timing at which the user utterance is made is before or after an output of a system utterance at a position corresponding to the tag from the voice output unit,

wherein the dialogue text creator generates the system utterance as a combination of a connective portion and a content portion, and inserts the tag between the connective portion and the content portion, and the connective portion comprises one of an interjection, a gambit, or a repetition of a part of a user utterance previously acquired.

2. The voice dialogue system according to claim 1 , wherein the intention understanding unit is further configured to:

interpret that the user utterance is a response to the system utterance in response to the user utterance being made after the output of the system utterance at the position corresponding to the tag from the voice output unit, and

interpret that the user utterance is not a response to the system utterance in response to the user utterance being input before the output of the system utterance at the position corresponding to the tag from the voice output unit.

3. The voice dialogue system according to claim 1 , wherein the intention understanding unit is further configured to:

calculate a first period of time, which is a period of time from the output of the system utterance from the voice output unit until the output of all texts preceding the tag from the voice output unit;

acquire a second period of time, which is a period of time from the output of the system utterance from the voice output unit until the start of input of the user utterance; and

compare the first period of time and the second period of time with each other to determine whether the timing at which the user utterance is made is before or after the output of a system utterance at the position corresponding to the tag from the voice output unit.

4. The voice dialogue system according to claim 1 , wherein the voice output unit does not output as a voice the tag in the text of the system utterance.

5. A method of understanding an utterance intention, the method comprising:

a voice input step of acquiring a user utterance;

an intention understanding step of interpreting an intention of utterance of a voice acquired in the voice input step;

a dialogue text creation step of creating a text of a system utterance; and

a voice output step of outputting the system utterance as voice data, wherein

in the dialogue text creation step, creating the text of the system utterance by inserting a tag in a position in the system utterance, and

in the intention understanding step, an utterance intention of a user is interpreted in accordance with whether a timing at which the user utterance is made is before or after an output of a system utterance at a position corresponding to the tag,

wherein the dialogue text creation step comprises generating the system utterance as a combination of a connective portion and a content portion, and inserting the tag between the connective portion and the content portion, and the connective portion comprises one of an interjection, a gambit, or a repetition of a part of a user utterance previously acquired.

6. The voice dialogue method according to claim 5 , wherein the intention understanding step comprises:

interpreting that, the user utterance is a response to the system utterance in response to the user utterance being made after the output of the system utterance at the position corresponding to the tag, and

interpreting that, the user utterance is not a response to the system utterance in response to the user utterance being input before the output of the system utterance at the position corresponding to the tag.

7. The voice dialogue method according to claim 5 , wherein the intention understanding step further comprises:

calculating a first period of time, which is a period of time from the output of the system utterance until the output of all texts preceding the tag;

acquiring a second period of time, which is a period of time from the output of the system utterance until the start of input of the user utterance; and

comparing the first period of time and the second period of time with each other to determine whether the timing at which the user utterance is made is before or after the output of a system utterance at the position corresponding to the tag from the voice output unit.

8. The voice dialogue method according to claim 5 , wherein, in the voice output step, the tag in the text of the system utterance is not output as a voice.

9. A computer-readable medium non-transitorily storing a program for causing a computer to execute the respective steps of the method according to claim 5 .

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Sep 14, 2017
From: IKENO, ATSUSHI; JINGUJI, YUSUKE; NISHIJIMA, TOSHIFUMI; KATAOKA, FUMINORI; TONEGAWA, HIROMI; UMEYAMA, NORIHIDE
To: TOYOTA JIDOSHA KABUSHIKI KAISHA
Reel/Frame 043592/0897 →
Priority Claims (1)
JP 2016-189406 · Sep 28, 2016 · national
Continuity (1)
Related Publication 20180090144A1 · Mar 29, 2018
Cited By (1)
US 12,340,803