IP Library Granted Patent US 9,460,714
Granted Patent B2
US 9,460,714 · App. 14/485,202 · Granted Oct 4, 2016

Speech processing apparatus and method

View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 9,460,714
App. No.
14/485,202
Granted
Oct 4, 2016
Kind
B2
Abstract

In a speech processing apparatus, an acquisition unit is configured to acquire a speech. A separation unit is configured to separate the speech into a plurality of sections in accordance with a prescribed rule. A calculation unit is configured to calculate a degree of similarity in each combination of the sections. An estimation unit is configured to estimate, with respect to the each section, a direction of arrival of the speech. A correction unit is configured to group the sections whose directions of arrival are mutually similar into a same group and correct the degree of similarity with respect to the combination of the sections in the same group. A clustering unit is configured to cluster the sections by using the corrected degree of similarity.

Claims (28)

1. A speech processing apparatus comprising:

an acquisition processor configured to acquire a speech, wherein a microphone array including a plurality of microphones to acquire the speech and to provide the acquired speech to the acquisition processor;

a separation processor configured to separate the speech into a plurality of sections in accordance with a prescribed rule;

a calculation processor configured to calculate a degree of similarity in each combination of the sections;

an estimation processor configured to estimate, with respect to the each section, a direction of arrival of the speech;

a correction processor configured to group the sections whose directions of arrival are mutually similar into a same group and correct the degree of similarity with respect to the combination of the sections in the same group; and

a clustering processor configured to cluster the sections by using the corrected degree of similarity.

2. The apparatus according to claim 1 , wherein the calculation processor includes:

a feature calculation processor configured to calculate acoustic features of each section, and

a similarity calculation processor configured to calculate the degree of similarity in each combination of the sections by using the calculated acoustic features.

3. The apparatus according to claim 2 ,

wherein the correction processor corrects the calculated degree of similarity with respect to the combination of the sections in the same group to a higher degree when the calculated degree of similarity is higher than a prescribed threshold.

4. The apparatus according to claim 3 ,

wherein the correction processor corrects the calculated degree of similarity to a higher degree by multiplying the calculated degree of similarity by N(N is a real number whose value is more than 1), or by raising the calculated degree of similarity to an M-th power(M is a real number whose value is more than 1).

5. A speech processing method comprising:

acquiring a speech, wherein the acquiring the speech acquires the speech through a microphone array;

separating the speech into a plurality of sections in accordance with a prescribed rule;

calculating a degree of similarity in each combination of the sections;

estimating, with respect to the each section, a direction of arrival of the speech;

grouping the sections whose directions of arrival are mutually similar into a same group and correcting the degree of similarity with respect to the combination of the sections in the same group; and

clustering the sections by using the corrected degree of similarity.

6. A non-transitory computer readable medium including computer executable instructions, wherein the instructions, when executed by a processor, cause the processor to perform a method comprising:

acquiring a speech, wherein the acquiring the speech acquires the speech through a microphone array;

separating the speech into a plurality of sections in accordance with a prescribed rule;

calculating a degree of similarity in each combination of the sections;

estimating, with respect to the each section, a direction of arrival of the speech;

grouping the sections whose directions of arrival are mutually similar into a same group and correcting the degree of similarity with respect to the combination of the sections in the same group; and

clustering the sections by using the corrected degree of similarity.

Assignments (4)
CORRECTIVE ASSIGNMENT TO CORRECT THE RECEIVING PARTY'S ADDRESS PREVIOUSLY RECORDED ON REEL 048547 FRAME 0187. ASSIGNOR(S) HEREBY CONFIRMS THE ASSIGNMENT OF ASSIGNORS INTEREST. Recorded May 6, 2020
From: KABUSHIKI KAISHA TOSHIBA
To: TOSHIBA DIGITAL SOLUTIONS CORPORATION
Reel/Frame 052595/0307 →
CORRECTIVE ASSIGNMENT TO CORRECT THE ADD SECOND RECEIVING PARTY PREVIOUSLY RECORDED AT REEL: 48547 FRAME: 187. ASSIGNOR(S) HEREBY CONFIRMS THE ASSIGNMENT. Recorded Aug 13, 2019
From: KABUSHIKI KAISHA TOSHIBA
To: KABUSHIKI KAISHA TOSHIBA; TOSHIBA DIGITAL SOLUTIONS CORPORATION
Reel/Frame 050041/0054 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Mar 8, 2019
From: KABUSHIKI KAISHA TOSHIBA
To: TOSHIBA DIGITAL SOLUTIONS CORPORATION
Reel/Frame 048547/0187 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Sep 12, 2014
From: DING, NING; KIDA, YUSUKE; HIROHATA, MAKOTO
To: KABUSHIKI KAISHA TOSHIBA
Reel/Frame 033733/0586 →