IP Library Granted Patent US 8,886,530
Granted Patent B2
US 8,886,530 · App. 13/529,585 · Granted Nov 11, 2014

Displaying text and direction of an utterance combined with an image of a sound source

Inventor: Kazuhiro Nakadai (Wako, JP)
Assignee: Honda Motor Co., Ltd.
G10L21/06G01L2021/02166
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 8,886,530
App. No.
13/529,585
Granted
Nov 11, 2014
Kind
B2
Abstract

An information processing device includes a display data creating unit configured to create display data including characters representing the content of an utterance based on a sound and a symbol surrounding the characters and indicating a first direction, and an image combining unit configured to determine the position of the display data based on a display position of an image representing a sound source of the utterance, and to combine the display data and the image of the sound source so that an orientation in which the sound is radiated is matched with the first direction.

Claims (30)

1. An information processing device comprising:

a display data creating unit configured to create display data including characters representing contents of an utterance based on a sound and a symbol surrounding the characters and indicating a first direction;

an image acquiring unit configured to acquire an image representing the sound source of the utterance;

a data input unit configured to input a viewpoint which is a position where the image is observed; and

an image combining unit configured to determine the position of the display data based on a display position of the image representing the sound source, and to combine the display data and the image of the sound source so that an orientation in which the sound is radiated is matched with the first direction, wherein

the image combining unit is configured to perform a viewpoint change based on the viewpoint input from the data input unit on the display data created by the display data creating unit, and to combine the display data, of which the viewpoint is changed, with the image acquired by the image acquiring unit, and

the display data creating unit is configured to determine the size of the characters representing the contents of the utterance based on a distance from the viewpoint to the position of the sound source.

2. The information processing device according to claim 1 , further comprising a position detecting unit configured to detect its own position,

wherein the data input unit is configured to input the position detected by the position detecting unit as the viewpoint.

3. The information processing device according to claim 1 , further comprising an emotion estimating unit configured to estimate an emotion of a speaker producing the sound of the utterance,

wherein the display data creating unit is configured to change the display form of the symbol based on the emotion estimated by the emotion estimating unit.

4. The information processing device according to claim 1 , wherein the display data creating unit is configured to determine the time at which the symbol is displayed based on the number of characters included in the display data.

5. An information processing system comprising:

a sound source position estimating unit configured to estimate the position of a sound source;

a orientation estimating unit configured to estimate an orientation in which the sound source radiates a sound wave;

a sound recognizing unit configured to recognize contents of an utterance from the sound source;

a display data creating unit configured to create display data including characters representing the contents of the utterance recognized by the sound recognizing unit and a symbol surrounding the characters and indicating a first direction;

an image acquiring unit configured to acquire an image representing the sound source of the utterance;

a data input unit configured to input a viewpoint which is a position where the image is observed; and

an image combining unit configured to determine the position of the display data based on a display position of the image representing the sound source of the utterance, and to combine the display data and the image of the sound source so that an orientation in which the sound is radiated is matched with the first direction, wherein

the image combining unit is configured to perform a viewpoint change based on the viewpoint input from the data input unit on the display data created by the display data creating unit, and to combine the display data, of which the viewpoint is changed, with the image acquired by the image acquiring unit, and

the display data creating unit is configured to determine the size of the characters representing the contents of the utterance based on a distance from the viewpoint to the position of the sound source.

6. The information processing system according to claim 5 , further comprising an imaging unit configured to capture an image representing the sound source of the utterance.

7. An information processing method in an information processing device, comprising the steps of:

creating display data including characters representing contents of an utterance based on a sound and a symbol surrounding the characters and indicating a first direction;

acquiring an image representing the sound source of the utterance;

inputting a viewpoint which is a position where the image is observed; and

determining the position of the display data based on a display position of the image representing the sound source of the utterance and combining the display data and the image of the sound source so that an orientation in which the sound is radiated is matched with the first direction, wherein,

in the step of combining the display data and the image of the sound source, a viewpoint change is performed based on the viewpoint on the display data, and the display data, of which the viewpoint is changed, are combined with the image representing the sound source, and

in the step of creating display data, the size of the characters representing the contents of the utterance is determined based on a distance from the viewpoint to the position of the sound source.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jul 24, 2012
From: NAKADAI, KAZUHIRO
To: HONDA MOTOR CO., LTD.
Reel/Frame 028621/0560 →
Continuity (2)
Provisional Application 61500653 · Feb 24, 2011
Related Publication 20120330659A1 · Dec 27, 2012