IP Library › Granted Patent US 9,208,782
Granted Patent B2
US 9,208,782 · App. 14/265,612 · Granted Dec 8, 2015

Speech processing device, speech processing method, and speech processing program

Inventors: Kazuhiro Nakadai (Wako, JP); Keisuke Nakamura (Wako, JP); Randy Gomez (Wako, JP)
Assignee: HONDA MOTOR CO., LTD.
G10L15/20G10L2021/02082
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 9,208,782
App. No.
14/265,612
Granted
Dec 8, 2015
Kind
B2
Abstract

A speech processing device includes a reverberation characteristic selection unit configured to correlate correction data indicating a contribution of a reverberation component based on a corresponding reverberation characteristic with an adaptive acoustic model which is trained using reverbed speech to which a reverberation based on the corresponding reverberation characteristic is added for each of reverberation characteristics, to calculate likelihoods based on the adaptive acoustic models for a recorded speech, and to select correction data corresponding to the adaptive acoustic model having the calculated highest likelihood, and a dereverberation unit configured to remove the reverberation component from the speech based on the correction data.

Claims (21)

1. A speech processing device comprising:

a reverberation characteristic selection unit configured to correlate correction data indicating a contribution of a reverberation component based on a corresponding reverberation characteristic with an adaptive acoustic model which is trained using reverbed speech to which a reverberation based on the corresponding reverberation characteristic is added for each of reverberation characteristics, to calculate likelihoods based on the adaptive acoustic models for a recorded speech, and to select correction data corresponding to the adaptive acoustic model having the calculated highest likelihood;

a dereverberation unit configured to remove the reverberation component from the speech based on the correction data,

wherein the reverberation characteristics differ in the contribution of a component which is inversely proportional to a distance between a sound collection unit configured to record speech from a sound source and the sound source, and

wherein the reverberation characteristic selection unit correlates distance data indicating the distances corresponding to the reverberation characteristics with the correction data and the adaptive acoustic models and selects the distance data corresponding to the adaptive acoustic model having the calculated highest likelihood;

an acoustic model prediction unit configured to predict an acoustic model corresponding to the distance indicated by the distance data selected by the reverberation characteristic selection unit from a first acoustic model trained using reverbed speech to which a reverberation corresponding to the reverberation characteristic based on a predetermined distance is added and a second acoustic model trained using speech in an environment in which reverberations are negligible; and

a speech recognition unit configured to perform a speech recognizing process on the speech using the acoustic model predicted by the acoustic model prediction unit.

2. A speech processing method comprising:

a reverberation characteristic selecting step of calculating a likelihood for a recorded speech based on an adaptive acoustic model trained using reverbed speech to which a reverberation based on a corresponding reverberation characteristic is added for each of reverberation characteristics and selecting correction data corresponding to the adaptive acoustic model having the calculated highest likelihood from a storage unit in which the adaptive acoustic model and the correction data are stored in correlation for each of the reverberation characteristics;

a dereverbing step of removing a reverberation component from the speech based on the correction data,

wherein the reverberation characteristic selecting step differs in the contribution of a component which is inversely proportional to a distance between a sound collection unit configured to record speech from a sound source and the sound source, and

wherein the reverberation characteristic selecting step correlates distance data indicating the distances corresponding to the reverberation characteristics with the correction data and the adaptive acoustic models and selects the distance data corresponding to the adaptive acoustic model having the calculated highest likelihood;

an acoustic model prediction step of predicting an acoustic model corresponding to the distance indicated by the distance data selected by the reverberation characteristic selection unit from a first acoustic model trained using reverbed speech to which a reverberation corresponding to the reverberation characteristic based on a predetermined distance is added and a second acoustic model trained using speech in an environment in which reverberations are negligible; and

a speech recognition step of performing a speech recognizing process on the speech using the acoustic model predicted by the acoustic model prediction step.

3. A non-transitory computer-readable storage medium comprising a speech processing program causing a computer of a speech processing device to execute:

a reverberation characteristic selecting process of calculating a likelihood for a recorded speech based on an adaptive acoustic model trained using reverbed speech to which a reverberation based on a corresponding reverberation characteristic is added for each of reverberation characteristics and selecting correction data corresponding to the adaptive acoustic model having the calculated highest likelihood from a storage unit in which the adaptive acoustic model and the correction data are stored in correlation for each of the reverberation characteristics;

a dereverbing process of removing a reverberation component from the speech based on the correction data,

wherein the reverberation characteristics differ in the contribution of a component which is inversely proportional to a distance between a sound collection unit configured to record speech from a sound source and the sound source, and

wherein the reverberation characteristic selecting process correlates distance data indicating the distances corresponding to the reverberation characteristics with the correction data and the adaptive acoustic models and selects the distance data corresponding to the adaptive acoustic model having the calculated highest likelihood;

an acoustic model prediction process of predicting an acoustic model corresponding to the distance indicated by the distance data selected by the reverberation characteristic selecting process from a first acoustic model trained using reverbed speech to which a reverberation corresponding to the reverberation characteristic based on a predetermined distance is added and a second acoustic model trained using speech in an environment in which reverberations are negligible; and

a speech recognition process of performing a speech recognizing process on the speech using the acoustic model predicted by the acoustic model prediction process.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Apr 30, 2014
From: NAKADAI, KAZUHIRO; NAKAMURA, KEISUKE; GOMEZ, RANDY
To: HONDA MOTOR CO., LTD.
Reel/Frame 032787/0697 →
Priority Claims (1)
JP 2013-143079 · Jul 8, 2013 · national
Continuity (1)
Related Publication 20150012268A1 · Jan 8, 2015