IP Library › Granted Patent US 11,450,316
Granted Patent B2
US 11,450,316 · App. 16/559,816 · Granted Sep 20, 2022

Agent device, agent presenting method, and storage medium

Inventors: Toshikatsu Kuramochi (Tokyo, JP); Wataru Endo (Wako, JP); Ryosuke Tanaka (Wako, JP)
Assignee: HONDA MOTOR CO., LTD.
G10L15/22B60K35/00B60K2370/148B60K2370/149B60K2370/1575G10L2015/223G10L2015/227
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,450,316
App. No.
16/559,816
Granted
Sep 20, 2022
Kind
B2
Abstract

An agent device includes a microphone which collects audio in a vehicle cabin, a speaker which outputs audio to the vehicle cabin, an interpreter which interprets the meaning of the audio collected by the microphone, a display provided in the vehicle cabin, and an agent controller which displays an agent image in a form of speaking to an occupant in a region of the display and causes the speaker to output audio by which the agent image speaks to at least one occupant, and the agent controller changes the face direction of the agent image to an direction different from an direction of the occupant who is a conversation target in a case that an utterance with respect to the face direction is interpreted by the interpreter after the agent image is displayed on the display.

Claims (53)

1. An agent device comprising:

a microphone which collects audio in a vehicle cabin;

a plurality of speakers which output audio to the vehicle cabin;

an interpreter which interprets the meaning of audio collected by the microphone;

a display provided in the vehicle cabin; and

an agent controller which displays an agent image in a form of speaking to an occupant in a region of the display in a state in which a face direction is recognizable and causes the plurality of speakers to output audio,

wherein the agent controller changes the face direction of the agent image to a direction different from a direction of the occupant who is a conversation target in a case that an utterance with respect to the face direction is interpreted by the interpreter after the agent image is displayed on the display,

a sound image can be located through a combination of outputs of the plurality of speakers, and

the agent controller displays the agent image in an area near the conversation target among one or more displays present in the vicinity of a plurality of occupants and controls the plurality of speakers such that a sound image is located at the display position of the agent image.

2. The agent device according to claim 1 , wherein the agent controller preferentially selects the occupant who is not a driver as the conversation target.

3. The agent device according to claim 2 , wherein the occupant preferentially selected as the conversation target is an occupant sitting on a passenger seat in the vehicle cabin.

4. The agent device according to claim 1 , wherein, in a case that the interpreter further performs the interpretation with respect to the face direction of the agent image after the face direction of the agent image has been changed, the agent controller sets the face direction as non-directional.

5. The agent device according to claim 1 , wherein the agent controller changes the face direction in a case that the interpreter interprets that input of a name of the agent image has been repeatedly received.

6. The agent device according to claim 1 , wherein the agent controller changes the face direction in a case that an increase rate of the sound pressure of the audio received through the microphone is equal to or higher than a predetermined rate.

7. The agent device according to claim 1 , wherein the agent controller changes the face direction in a case that it is determined that an utterance interpreted by the interpreter has been interrupted halfway.

8. The agent device according to claim 1 , further comprising a camera which captures an image of the occupants,

wherein the interpreter further interprets an image collected by the camera,

in a case that speaking of an occupant to the agent image has been recognized, the agent controller displays the agent image in a state in which the agent image faces in at least a direction in which any occupant is present before the agent image replies to speaking of the occupant,

the interpreter interprets images collected by the camera before and after the agent image is displayed in a state in which the agent image faces in at least a direction in which any occupant is present, and

the agent controller changes the face direction if it is determined that a facial expression of the occupant has changed in a case that the agent image is displayed in a state in which the agent image faces in at least a direction in which any occupant is present.

9. The agent device according to claim 8 , wherein the agent controller changes the face direction in a case that a negative facial expression change of the occupant selected as a conversation target has been detected.

10. The agent device according to claim 1 , wherein the agent controller displays the agent image at one end of the display and displays the face direction of the agent image such that the face direction faces in a direction of the other end of the display.

11. An agent presenting method using a computer, comprising:

collecting audio in a vehicle cabin;

outputting audio to the vehicle cabin;

interpreting the meaning of the collected audio;

displaying an agent image in a form speaking to an occupant in a state in which a face direction is recognizable and outputting the audio;

changing the face direction of the agent image to a direction different from a direction of the occupant who is a conversation target in a case that an utterance with respect to the face direction is interpreted after the agent image is displayed;

locating a sound image through a combination of outputs of a plurality of speakers; and

displaying the agent image in an area near the conversation target among one or more displays present in the vicinity of a plurality of occupants and controlling the plurality of speakers such that a sound image is located at the display position of the agent image.

12. An agent device comprising:

a microphone which collects audio in a vehicle cabin;

a speaker which outputs audio to the vehicle cabin;

an interpreter which interprets the meaning of audio collected by the microphone;

a display provided in the vehicle cabin;

a camera which captures an image of occupants; and

an agent controller which displays an agent image in a form of speaking to an occupant in a region of the display in a state in which a face direction is recognizable and causes the speaker to output audio,

wherein the agent controller changes the face direction of the agent image to a direction different from a direction of the occupant who is a conversation target in a case that an utterance with respect to the face direction is interpreted by the interpreter after the agent image is displayed on the display,

the interpreter further interprets an image collected by the camera,

in a case that speaking of an occupant to the agent image has been recognized, the agent controller displays the agent image in a state in which the agent image faces in at least a direction in which any occupant is present before the agent image replies to speaking of the occupant,

the interpreter interprets images collected by the camera before and after the agent image is displayed in a state in which the agent image faces in at least a direction in which any occupant is present, and

the agent controller changes the face direction if it is determined that a facial expression of the occupant has changed in a case that the agent image is displayed in a state in which the agent image faces in at least a direction in which any occupant is present.

13. An agent presenting method using a computer, comprising:

collecting audio in a vehicle cabin;

outputting audio to the vehicle cabin;

interpreting the meaning of the collected audio;

displaying an agent image in a form speaking to an occupant in a state in which a face direction is recognizable and outputting the audio;

changing the face direction of the agent image to a direction different from a direction of the occupant who is a conversation target in a case that an utterance with respect to the face direction is interpreted after the agent image is displayed;

capturing an image of occupants by a camera,

further interpreting an image collected by the camera,

in a case that speaking of an occupant to the agent image has been recognized, displaying the agent image in a state in which the agent image faces in at least a direction in which any occupant is present before the agent image replies to speaking of the occupant,

interpreting images collected by the camera before and after the agent image is displayed in a state in which the agent image faces in at least a direction in which any occupant is present, and

changing the face direction if it is determined that a facial expression of the occupant has changed in a case that the agent image is displayed in a state in which the agent image faces in at least a direction in which any occupant is present.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Sep 4, 2019
From: KURAMOCHI, TOSHIKATSU; ENDO, WATARU; TANAKA, RYOSUKE
To: HONDA MOTOR CO., LTD.
Reel/Frame 050258/0813 →
Priority Claims (1)
JP JP2018-189708 · Oct 5, 2018 · national
Continuity (1)
Related Publication 20200111489A1 · Apr 9, 2020