IP Library Granted Patent US 9,626,958
Granted Patent B2
US 9,626,958 · App. 15/167,683 · Granted Apr 18, 2017

Speech retrieval method, speech retrieval apparatus, and program for speech retrieval apparatus

Inventors: Gakuto Kurata (Tokyo, JP); Tohru Nagano (Tokyo, JP); Masafumi Nishimura (Kanagawa, JP)
Assignee: SINOEAST CONCEPT LIMITED
G10L15/02G10L15/04G10L15/08G10L15/187G10L25/51G10L2015/025G10L2015/027G10L2015/088
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 9,626,958
App. No.
15/167,683
Granted
Apr 18, 2017
Kind
B2
Abstract

A method for speech retrieval includes acquiring a keyword designated by a character string, and a phoneme string or a syllable string, detecting one or more coinciding segments by comparing a character string that is a recognition result of word speech recognition with words as recognition units performed for speech data to be retrieved and the character string of the keyword, calculating an evaluation value of each of the one or more segments by using the phoneme string or the syllable string of the keyword to evaluate a phoneme string or a syllable string that is recognized in each of the detected one or more segments and that is a recognition result of phoneme speech recognition with phonemes or syllables as recognition units performed for the speech data, and outputting a segment in which the calculated evaluation value exceeds a predetermined threshold.

Claims (16)

1. A method for speech retrieval, comprising:

detecting one or more coinciding segments for speech data by comparing a character string of a recognition result and a character string of a keyword, the keyword being designated by the character string and a phoneme string or a syllable string;

calculating an evaluation value of each of the one or more coinciding segments using the phoneme string or the syllable string of the keyword to evaluate a phoneme string or a syllable string recognized in each of the one or more coinciding segments and that is a recognition result of phoneme speech recognition, wherein the phoneme string or the syllable string associated with each of the coinciding segments is a phoneme string or a syllable string associated with a segment in which a start and an end of the segment is expanded by a predetermined time; and

outputting a segment in which the calculated evaluation value exceeds a predetermined threshold.

2. The method according to claim 1 , wherein the recognition result of word speech recognition includes words as recognition units performed for the speech data.

3. The method according to claim 1 , wherein the recognition result of phoneme speech recognition includes phonemes or syllables as recognition units performed for the speech data.

4. The method according to claim 1 , wherein calculating comprises comparing a phoneme string or a syllable string that is an N-best recognition result of phoneme speech recognition with phonemes or syllables as recognition units performed for speech data associated with each of the detected one or more coinciding segments and the phoneme string of the keyword to set a rank of the coinciding N-best recognition result as the evaluation value.

5. The method according to claim 1 , wherein calculating comprises setting, as the evaluation value, an edit distance between a phoneme string or a syllable string that is a 1-best recognition result of phoneme speech recognition with phonemes or syllables as recognition units performed for speech data associated with each of the detected one or more coinciding segments and the phoneme string or the syllable string of the keyword.

6. The method according to claim 5 , wherein the edit distance is a distance matched by matching based on dynamic programming.

7. The method according to claim 1 , further comprising performing word speech recognition of the speech data to be retrieved, with words as recognition units.

8. The method according to claim 1 , further comprising performing phoneme speech recognition of the speech data associated with each of the detected one or more coinciding segments, with phonemes or syllables as recognition units.

9. The method according to claim 1 , further comprising performing phoneme speech recognition of the speech data to be retrieved, with phonemes or syllables as recognition units.

10. The method of claim 1 , wherein the calculating the evaluation value of each of the one or more coinciding segments further includes using the character string of the keyword to evaluate the character string in each of the detected one or more coinciding segments.

11. The method of claim 1 , further comprising adjusting the predetermined threshold to alter at least one of a precision value and a recall value of the output segment, the precision value being positively correlated with the predetermined threshold and the recall value being negatively correlated with the predetermined threshold.

12. The method of claim 11 , wherein the precision value is a ratio of retrieval results satisfying a retrieval request to all documents satisfying the retrieval request.

13. The method of claim 11 , wherein the recall value is a ratio of retrieval results satisfying a retrieval request to all retrieval results.

Assignments (2)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jan 17, 2017
From: INTERNATIONAL BUSINESS MACHINES CORPORATION
To: SINOEAST CONCEPT LIMITED
Reel/Frame 041388/0557 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded May 27, 2016
From: KURATA, GAKUTO; NAGANO, TOHRU; NISHIMURA, MASAFUMI
To: INTERNATIONAL BUSINESS MACHINES CORPORATION
Reel/Frame 038740/0598 →
Priority Claims (1)
JP 2014-087325 · Apr 21, 2014 · national
Continuity (3)
Continuation 14745912 · Jun 22, 2015
Continuation 14692105 · Apr 21, 2015
Related Publication 20160275940A1 · Sep 22, 2016