IP Library Granted Patent US 9,286,029
Granted Patent B2
US 9,286,029 · App. 13/934,396 · Granted Mar 15, 2016

System and method for multimodal human-vehicle interaction and belief tracking

Inventors: Antoine Raux (Cupertino, CA); Ian Lane (Sunnyvale, CA); Rakesh Gupta (Mountain View, CA)
Assignees: Honda Motor Co., Ltd.; Ian Lane
G06F3/167G06F3/01G06F3/017
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 9,286,029
App. No.
13/934,396
Granted
Mar 15, 2016
Kind
B2
Abstract

A method and system for multimodal human-vehicle interaction including receiving input from an occupant in a vehicle via more than one mode and performing multimodal recognition of the input. The method also includes augmenting at least one recognition hypothesis based on at least one visual point of interest and determining a belief state of the occupant's intent based on the recognition hypothesis. The method further includes selecting an action to take based on the determined belief state.

Claims (27)

1. A method for multimodal human-vehicle interaction, comprising:

receiving input from an occupant in a vehicle via more than one mode, wherein the input includes a speech input and a gesture input;

performing multimodal recognition of the input to determine a reference to a point of interest based on the speech input and to extract a visual point of interest based on the gesture input and the reference to the point of interest in the speech input;

augmenting at least one recognition hypothesis based on the visual point of interest;

determining a belief state of the occupant's intent, wherein the belief state is determined based on joint probability distribution tables of probabilistic ontology trees and the probabilistic ontology trees are based on the recognition hypothesis; and

selecting an action to take based on the determined belief state.

2. The method of claim 1 , wherein performing multimodal recognition includes speech recognition of the speech input and gesture recognition of the gesture input.

3. The method of claim 2 , wherein the speech recognition includes determining whether the point of interest is present in a current dialog.

4. The method of claim 1 , wherein extracting the visual point of interest is based on the speech input, the gesture input, and a location of the vehicle.

5. The method of claim 4 , wherein the gesture input includes at least an eye gaze of the occupant.

6. A method for multimodal human-vehicle interaction, comprising:

receiving an input from an occupant of a vehicle including a first input and a second input, wherein the first and second inputs represent different modalities;

performing multimodal recognition of the first and second inputs to determine a reference of a point of interest based on the first input and to extract a visual point of interest based on the second input and the reference to the point of interest in the first input;

modifying a recognition hypothesis of the first input with the second input;

determining a belief state of the occupant's intent, wherein the belief state is determined based on joint probability distribution tables of probabilistic ontology trees and the probabilistic ontology trees are based on the recognition hypothesis; and

selecting an action to take based on the determined belief state.

7. The method of claim 6 , wherein extracting the visual point of interest is based on the first input, the second input, and a location of the vehicle.

8. The method of claim 7 , wherein the recognition hypothesis is modified based on the visual point of interest.

9. The method of claim 7 , including determining whether the reference to the visual point of interest is within a current dialog.

10. A system for multimodal human-vehicle interaction, comprising:

a plurality of sensors for sensing interaction data from a vehicle occupant, wherein the interaction data includes a speech input and a gesture input;

a multimodal recognition module for performing multimodal recognition of the interaction data;

a point of interest identification module for determining a reference to a point of interest based on the speech input and to extract a visual point of interest based on the gesture input and the reference to the point of interest in the speech input, wherein the multimodal recognition module augments a recognition hypothesis based on the visual point of interest,

a belief tracking module for determining a belief state of the occupant's intent, wherein the belief state is determined based on joint probability distribution tables of probabilistic ontology trees and the probabilistic ontology trees are based on the recognition hypothesis; and

a dialog management and action module for selecting an action to take based on the determined belief state.

11. The system of claim 10 , including a point of interest history database wherein the point of interest identification module utilizes the database to determine whether the reference to the point of interest is within a current dialog.

12. The system of claim 10 , wherein extracting the visual point of interest from the gesture input is based on the speech input, the gesture input, and a location of the vehicle.

Assignments (2)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Dec 5, 2018
From: LANE, IAN
To: CAPIO, INC.
Reel/Frame 047684/0177 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Oct 4, 2013
From: RAUX, ANTOINE; LANE, IAN; GUPTA, RAKESH
To: HONDA MOTOR CO., LTD.; LANE, IAN
Reel/Frame 031348/0043 →
Continuity (2)
Provisional Application 61831783 · Jun 6, 2013
Related Publication 20140361973A1 · Dec 11, 2014