IP Library Granted Patent US 9,899,028
Granted Patent B2
US 9,899,028 · App. 14/826,527 · Granted Feb 20, 2018

Information processing device, information processing system, information processing method, and information processing program

Inventors: Kazuhiro Nakadai (Wako, JP); Takeshi Mizumoto (Wako, JP); Keisuke Nakamura (Wako, JP); Masayuki Takigahira (Tokyo, JP)
Assignee: HONDA MOTOR CO., LTD.
G10L15/32
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 9,899,028
App. No.
14/826,527
Granted
Feb 20, 2018
Kind
B2
Abstract

An information processing device includes a first information processing unit, a communication unit, and a control unit. The first information processing unit performs predetermined information processing on input data to generate first processing result data. The communication unit is capable of receiving second processing result data generated by a second information processing unit capable of executing the same kind of information processing as the information processing on the input data under a condition with higher versatility. The control unit selects either the first processing result data or the second processing result data according to the use environment of the device.

Claims (48)

1. An information processing device comprising:

a first speech recognizer which performs speech recognition on an input speech signal using first speech recognition data to generate first text data;

a transceiver which is capable of receiving, from a second speech recognizer which performs speech recognition using second speech recognition data with higher versatility than the first speech recognition data to generate second text data, the second text data; and

a controller that determines whether or not to stop an operation of the first speech recognizer based on a communication state with the second speech recognizer; and

a first preprocessor which performs pre-processing on the speech signal to generate a first acoustic feature quantity, wherein

the first preprocessor includes hierarchical processors of L (where L is a prescribed integer equal to or greater than 1) hierarchies,

an m-th (where m is an integer equal to or greater than 1 and equal to or less than L) hierarchical processor performs an m-th hierarchical processing on m-th hierarchical data to generate (m+1)th hierarchical data, first hierarchical data is the speech signal, and (L+1)th hierarchical data is the first acoustic feature quantity, and

the controller determines to which hierarchy of the hierarchical processors is operated in response to the communication state when stopping, an operation of the first speech recognizer, wherein

a first hierarchical processor is a sound source localizer which calculates a sound source direction of each sound source from a speech signal of a plurality of channels,

a second hierarchical processor is a sound source separator which separates the speech signal of the plurality of channels into sound source-specific speech signals of each of the sound sources, and

a third hierarchical processor is a feature quantity calculator which calculates acoustic feature quantities from the sound source-specific speed signals.

2. An information processing system comprising:

a first information processing device; and a second information processing device, wherein

the first information processing device includes:

a first speech recognizer which performs speech recognition on an input speech signal using first speech recognition data to generate first text data;

a transceiver which is capable of receiving second text data from the second information processing device; and

a controller that determines whether or not to stop an operation of the first speech recognizer based on a communication state with the second information processing device, and

a preprocessor which performs pre-processing on the speech signal to generate a first acoustic feature quantity, and wherein

the preprocessor includes hierarchical processors of L (where L is a prescribed integer equal to or greater than 1) hierarchies,

an m-th (where m is an integer equal to or greater than 1 and equal to or less than L) hierarchical processor performs an m-th hierarchical processing on m-th hierarchical data to generate (m+1)th hierarchical data, first hierarchical data is the speech signal, and (L+1)th hierarchical data is the first acoustic feature quantity, and

the controller determines to which hierarchy of the hierarchical processors is operated in response to the communication state when stopping an operation of the first speech recognizer, and wherein

the second information processing device includes:

a second speech recognizer which performs speech recognition on the speech signal using second speech recognition data with higher versatility than the first speech recognition data to generate second text data, wherein

L is 3,

a first hierarchical processor is a sound source localizer which calculates a sound source direction of each sound source from a speech signal of a plurality of channels,

a second hierarchical processor is a sound source separator which separates the speech signal of the plurality of channels into sound source-specific speed′ signals of each of the sound sources, and

a third hierarchical processor is a feature quantity calculator which calculates acoustic feature quantities from the sound source-specific speech signals.

3. An information processing method in an information processing device that comprises a transceiver which is capable of receiving, from a speech recognizer which performs speech recognition using second speech recognition data with higher versatility than first speech recognition data to generate second text data, the second text data, the method comprising:

a speech recognition process of performing speech recognition on an input speech signal using the first speech recognition data to generate first text data;

a control process of determining whether or not to stop the speech recognition process based on a communication state with the speech recognizer; and

a preprocessing process of performing pre-processing on the speech signal to generate a first acoustic feature quantity, wherein

the preprocessing process includes hierarchical processing processes of L (where L is a prescribed integer equal to or greater than 1) hierarchies,

an m-th (where m is an integer equal to or greater than 1 and equal to or less than L) hierarchical processing process performs an m-th hierarchical processing on m-th hierarchical data to generate (m+1)th hierarchical data, first hierarchical data is the speech signal, and (L+1)th hierarchical data is the first acoustic feature quantity, and

the control process determines to which hierarchy of the hierarchical processing processes is executed in response to the communication state when stopping the speech recognition process based on the communication state, wherein

L is 3,

a first hierarchical processor is a sound source localizer which calculates a sound source direction of each sound source from a speech signal of a plurality of channels,

a second hierarchical processor is a sound source separator which separates the speech signal of the plurality of channels into sound source-specific speech signals of each of the sound sources, and

a third hierarchical processor is a feature quantity calculator which calculates acoustic feature quantities from the sound source-specific speech signals.

4. A non-transitory computer readable, medium storing an information processing program that causes a computer of an information processing device that comprises a transceiver which is capable of receiving, from a speech recognizer which performs speech recognition using second speech recognition data with higher versatility than first speech recognition data to generate second text data, the second text data to execute:

a speech recognition sequence of performing speech recognition on an input speech signal using the first speech recognition data to generate first text data;

a control sequence of determining whether or not to stop the speech recognition sequence based on a communication state with the speech recognizer; and

a preprocessing sequence of performing pre-processing on the speech signal to generate a first acoustic feature quantity, wherein

the first preprocessing sequence includes hierarchical processing sequences of L (where L is a prescribed integer equal to or greater than 1) hierarchies,

an m-th (where m is an integer equal to or greater than 1 and equal to or less than L) hierarchical processing sequence performs an m-th hierarchical processing on m-th hierarchical data to generate (m+1)th hierarchical data, first hierarchical data is the speech signal, and (L+1)th hierarchical data is the first acoustic feature quantity, and

the control sequence determines to which hierarchy of the hierarchical processing sequences is executed in response to the communication state when stopping the speech recognition sequence, wherein

a first hierarchical processor is a sound source localizer which calculates a sound source direction of each sound source from a speech signal of a plurality of channels,

a second hierarchical processor is a sound source separator which separates the speech signal of the plurality of channels into sound source-specific speech signals of each of the sound sources, and

a third hierarchical processor is a feature quantity calculator which calculates acoustic feature quantities from the sound source-specific speed signals.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Aug 14, 2015
From: NAKADAI, KAZUHIRO; MIZUMOTO, TAKESHI; NAKAMURA, KEISUKE; TAKIGAHIRA, MASAYUKI
To: HONDA MOTOR CO., LTD.
Reel/Frame 036330/0351 →
Priority Claims (2)
JP 2014-168632 · Aug 21, 2014 · national
JP 2015-082359 · Apr 14, 2015 · national
Continuity (1)
Related Publication 20160055850A1 · Feb 25, 2016