IP Library › Granted Patent US 10,950,227
Granted Patent B2
US 10,950,227 · App. 15/904,738 · Granted Mar 16, 2021

Sound processing apparatus, speech recognition apparatus, sound processing method, speech recognition method, storage medium

Inventor: Takehiko Kagoshima (Kanagawa, JP)
Assignee: Kabushiki Kaisha Toshiba
G10L15/20G01S5/18G10L15/00G10L15/02G10L15/063G10L15/22G10L25/51G10L25/30
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 10,950,227
App. No.
15/904,738
Granted
Mar 16, 2021
Kind
B2
Abstract

According to one embodiment, a sound processing apparatus extracts a feature of first speech uttered outside an objective area from first speech obtained at positions different from each other in a space of the objective area and a place outside the objective area. The apparatus creates, by learning, a determination model configured to determine whether an utterance position of second speech in the space is outside the objective area based at least in part on the feature uttered outside the objective area. The apparatus eliminates a portion of the second speech uttered outside the objective area from the second speech obtained by a second microphone based at least in part on the feature and the model. The apparatus detects and outputs remaining speech from the second speech.

Claims (72)

1. A sound processing apparatus comprising:

a plurality of first microphones that are arranged at positions different from each other in a space comprising an objective area and a place outside the objective area, and obtains a first sound comprising first speech uttered at the place outside the objective area;

a second microphone that is arranged in the space, and obtains second speech;

a feature extractor that extracts a first feature from the first sound obtained by the plurality of first microphones, and extracts a second feature from the second speech obtained by the second microphone, the first feature comprising a speech feature corresponding to the first speech obtained by the plurality of first microphones, and a background sound corresponding to sound in the first sound other than the first speech;

a speech feature extractor configured to extract the speech feature from the first feature, and store the speech feature in a storage;

a determination model creator that creates a determination model configured to determine whether an utterance position of a portion of the second speech is in the place outside the objective area by learning based at least in part on the speech feature stored in the storage;

a speech detector that eliminates the portion of the second speech from the second speech obtained by the second microphone when it is determined that the portion of the second speech is uttered at the place outside the objective area based at least in part on the speech feature of the first speech and the determination model, wherein the speech detector detects and outputs remaining speech from the second speech; and

a switch configured to switch between an initial setting mode and an operation mode in response to instructions sent from a controller, wherein the first speech is obtained in the initial setting mode, the second speech is obtained in the operation mode, and the controller compares a feature of the portion of the second speech and the determination model in the operation mode.

2. The sound processing apparatus of claim 1 , wherein the determination model creator carries out learning of the determination model based at least in part on the first feature of the first speech uttered at the place outside the objective area in the space and a third feature of a third speech uttered at the place outside of the objective area, the third speech obtained by the first plurality of microphones.

3. The sound processing apparatus of claim 1 , wherein the plurality of first microphones comprises the second microphone.

4. The sound processing apparatus of claim 1 , further comprising a noise elimination device that eliminates noise from the second speech obtained by the second microphone, wherein the speech detector eliminates the portion of the second speech uttered outside of the objective area from the second speech after the noise has already been eliminated by the noise elimination device based at least in part on the first feature of the first speech and the determination model, wherein the speech detector detects and outputs the remaining speech from the second speech.

5. The sound processing apparatus of claim 1 , further comprising a recognition device configured to recognize contents of the remaining speech detected by the speech detector.

6. A speech recognition apparatus comprising:

a plurality of first microphones that are arranged at positions different from each other in a space comprising an objective area and a place outside the objective area, and obtains a first sound comprising first speech uttered at the place outside the objective area;

a second microphone that is arranged in the space, and obtains second speech;

a feature extractor that extracts a first feature from the first sound obtained by the plurality of first microphones, and extracts a second feature from the second speech obtained by the second microphone, the first feature comprising a speech feature corresponding to the first speech obtained by the plurality of first microphones, and a background sound corresponding to sound in the first sound other than the first speech;

a speech feature extractor configured to extract the speech feature from the first feature, and store the speech feature in a storage;

a determination model creator that creates a determination model configured to determine whether an utterance position of a portion of the second speech is outside the objective area by learning based at least in part on the speech feature stored in the storage;

a speech detector that eliminates the portion of the second speech from the second speech obtained by the second microphone when it is determined that the portion of the second speech is uttered at the place outside the objective area based at least in part on the speech feature of the first speech and the determination model, wherein the speech detector detects and outputs remaining speech from the second speech;

a recognition device that recognizes contents of the remaining speech detected by the speech detector; and

a switch configured to switch between an initial setting mode and an operation mode in response to instructions sent from a controller, wherein the first speech is obtained in the initial setting mode, the second speech is obtained in the operation mode, and the controller compares a feature of the portion of the second speech and the determination model in the operation mode.

7. A method for sound processing, the method comprising:

arranging a plurality of first microphones at positions different from each other in a space comprising an objective are and a place outside the objective area;

obtaining a first sound comprising first speech uttered at the place outside the objective area by the plurality of first microphones;

arranging a second microphone in the space;

obtaining second speech by the second microphone;

extracting a first feature from the first sound obtained by the plurality of first microphones, the first feature comprising a speech feature corresponding to the first speech obtained by the plurality of first microphones, and a background sound corresponding to sound in the first sound other than the first speech;

extracting a second feature from the second speech;

extracting the speech feature from the first feature;

storing the speech feature in a storage;

creating a determination model configured to determine whether an utterance position of a portion of the second speech is outside the objective area by learning based at least in part on the speech feature stored in the storage; and

eliminating the portion of the second speech from the second speech obtained by the second microphone when it is determined that the portion of the second speech is uttered at the place outside the objective area based at least in part on the speech feature of the first speech and the determination model;

detecting remaining speech from the second speech;

outputting the remaining speech;

switching between an initial setting mode and an operation mode in response to instructions sent from a controller, wherein the first speech is obtained in the initial setting mode, the second speech is obtained in the operation mode; and

comparing a feature of the portion of the second speech and the determination model in the operation mode.

8. A method of speech recognition, the method comprising:

arranging a plurality of first microphones at positions different from each other in a space comprising an objective area and a place outside an objective area;

obtaining a first sound comprising first speech uttered at the place outside the objective area by the plurality of first microphones;

arranging a second microphone in the space;

obtaining second speech by the second microphone;

extracting a first feature from first sound obtained by the plurality of first microphones, the first feature comprising a speech feature corresponding to the first speech obtained by the plurality of first microphones, and a background sound corresponding to sound in the first sound other than the first speech;

extracting a second feature from the second speech;

extracting the speech feature from the first feature;

storing the speech feature in a storage;

creating a determination model configured to determine whether an utterance position of a portion of the second speech is outside the objective area by learning based at least in part on the speech feature stored in the storage;

eliminating the portion of the second speech from the second speech obtained by the second microphone when it is determined that the portion of the second speech is uttered at the place outside the objective area based at least in part on the speech feature of the first speech and the determination model;

detecting remaining speech from the second speech;

recognizing contents of the remaining speech;

switching between an initial setting mode and an operation mode in response to instructions sent from a controller, wherein the first speech is obtained in the initial setting mode, the second speech is obtained in the operation mode; and

comparing a feature of the portion of the second speech and the determination model in the operation mode.

9. A non-transitory computer-readable storage comprising a computer program that is executable by a computer used in a sound processing program, the computer program comprising instructions for causing the computer to execute functions of:

extracting a first feature from a first sound comprising first speech uttered at a place outside an objective area, and obtained by a plurality of first microphones that are arranged at positions different from each other in a space comprising the objective area and the place outside the objective area, the first feature comprising a speech feature corresponding to the first speech obtained by the plurality of first microphones, and a background sound corresponding to sound in the first sound other than the first speech;

extracting the speech feature from the first feature;

storing the speech feature in a storage;

creating a determination model configured to determine whether an utterance position of a portion of second speech obtained by a second microphone that is arranged in the space is outside the objective area by learning based at least in part on the speech feature stored in the storage; and

eliminating the portion of the second speech from the second speech when it is determined that the portion of the second speech is uttered at the place outside the objective area based at least in part on the speech feature of the first speech and the determination model;

detecting remaining speech from the second speech;

outputting the remaining speech;

switching between an initial setting mode and an operation mode in response to instructions sent from a controller, wherein the first speech is obtained in the initial setting mode, the second speech is obtained in the operation mode; and

compares a feature of the portion of the second speech and the determination model in the operation mode.

10. A non-transitory computer-readable storage comprising a computer program that is executable by a computer used in a speech recognition program, the computer program comprising instructions for causing the computer to execute functions of:

extracting a first feature from a first sound comprising first speech uttered at a place outside an objective area, and obtained by a plurality of first microphones that are arranged at positions different from each other in a space comprising the objective area the place outside the objective area, the first feature comprising a speech feature corresponding to the first speech obtained by the plurality of first microphones, and a background sound corresponding to sound in the first sound other than the first speech;

extracting the speech feature from the first feature;

storing the speech feature in a storage;

creating a determination model configured to determine whether an utterance position of a portion of second speech obtained by a second microphone that is arranged in the space is outside the objective area by learning based at least in part on the speech feature amount of the first speech uttered outside the objective area stored in the storage; and

eliminating the portions of the second speech from the second speech when it is determined that the portion of the second speech is uttered at the place outside the objective area based at least in part on the speech feature of the first speech and the determination model;

detecting remaining speech from the second speech;

outputting the remaining speech;

recognizing contents of the remaining speech;

switching between an initial setting mode and an operation mode in response to instructions sent from a controller, wherein the first speech is obtained in the initial setting mode, the second speech is obtained in the operation mode; and

compares a feature of the portion of the second speech and the determination model in the operation mode.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Feb 26, 2018
From: KAGOSHIMA, TAKEHIKO
To: KABUSHIKI KAISHA TOSHIBA
Reel/Frame 045041/0140 →
Priority Claims (1)
JP 2017-177022 · Sep 14, 2017 · national
Continuity (1)
Related Publication 20190080689A1 · Mar 14, 2019