IP Library › Granted Patent US 11,557,288
Granted Patent B2
US 11,557,288 · App. 16/845,849 · Granted Jan 17, 2023

Hindrance speech portion detection using time stamps

Inventors: Nobuyasu Itoh (Kanagawa, JP); Gakuto Kurata (Tokyo, JP); Masayuki Suzuki (Tokyo, JP)
Assignee: INTERNATIONAL BUSINESS MACHINES CORPORATION
G10L15/197
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,557,288
App. No.
16/845,849
Granted
Jan 17, 2023
Kind
B2
Abstract

A computer-implemented method of detecting a portion of audio data to be removed is provided. The method includes obtaining a recognition result of audio data. The recognition result includes recognized text data and time stamps. The method also includes extracting one or more candidate phrases from the recognition result using n-gram counts. The method further includes, for each candidate phrase, making pairs of same phrases with different time stamps and clustering the pairs of the same phrase by using differences in time stamps. The method includes further determining a portion of the audio data to be removed using results of the clustering.

Claims (39)

1. A computer-implemented method for detecting a portion of audio data to be removed, the method comprising:

obtaining a recognition result of audio data, the recognition result including recognized text data and time stamps;

extracting one or more candidate phrases from the recognition result based on a comparison between a number of times an n-gram appears within the recognition result and a threshold;

for each candidate phrase, making a plurality of pairs of same phrases with different time stamps;

for each candidate phrase, clustering the plurality of pairs of the same phrases by using a difference in time stamps for each pair of the same phrases; and

determining a portion of the audio data to be removed using results of the clustering.

2. The method according to claim 1 , further comprising preparing a training data for training a model by removing the portion of the audio data from the audio data.

3. The method according to claim 1 , wherein the clustering the plurality of pairs of the same phrases includes clustering the plurality of pairs of the same phrases while allowing a difference in time stamps less than a predetermined threshold for each pair.

4. The method according to claim 3 , wherein the predetermined threshold is determined depending on the portion of the audio data to be removed.

5. The method according to claim 4 , wherein the portion of the audio data to be removed is at least one selected from the group consisting of recorded speech data, repeated speech data, and speech data spoken by a same speaker.

6. The method according to claim 3 , wherein the making the plurality of pairs of the same phrases includes, for each pair of the same phrases, obtaining a sum of differences in time stamps between corresponding words in the pair of the same phrases as the difference in time stamps for each pair of the same phrases.

7. The method according to claim 6 wherein the differences in the time stamps includes a difference in a duration time length of a word.

8. The method according to claim 6 , wherein the differences in the time stamps includes a difference in silent time length between two adjacent words.

9. A computer system for detecting a portion of audio data to be removed, by executing program instructions, the computer system comprising:

a memory tangibly storing the program instructions; and

a processor in communications with the memory, wherein the processor is configured to:

obtain a recognition result of audio data, the recognition result including recognized text data and time stamps;

extract one or more candidate phrases from the recognition result based on a comparison between a number of times an n-gram appears within the recognition result and a threshold;

make, for each candidate phrase, a plurality of pairs of same phrases with different time stamps;

cluster, for each candidate phrase, the plurality of pairs of the same phrases by using differences in time stamps; and

determine a portion of the audio data to be removed using results of the clustering.

10. The computer system of claim 9 , wherein the plurality of pairs of the same phrases is clustered while allowing a difference in time stamps less than a predetermined threshold for each pair.

11. The computer system of claim 10 , wherein the predetermined threshold is determined depending on the portion of the audio data to be removed.

12. The computer system of claim 9 , wherein the processor is configured to:

prepare a training data for training a model by removing the portion of the audio data from the audio data.

13. The computer system of claim 9 , wherein the processor is configured to:

obtain, for each pair of the same phrases, a sum of differences in time stamps between corresponding words in the pair of the same phrases as the difference in time stamps.

14. The computer system of claim 13 , wherein the differences in the time stamps includes a difference in a duration time length of a word.

15. The computer system of claim 13 , wherein the differences in the time stamps includes a difference in silent time length between two adjacent words.

16. A computer program product for detecting a portion of audio data to be removed, the computer program product comprising a computer readable storage medium having program instructions embodied therewith, the program instructions executable by a computer to cause the computer to perform a computer-implemented method comprising:

obtaining a recognition result of audio data, the recognition result including recognized text data and time stamps;

extracting one or more candidate phrases from the recognition result based on a comparison between a number of times an n-gram appears within the recognition result and a threshold;

for each candidate phrase, making a plurality of pairs of same phrases with different time stamps;

for each candidate phrase, clustering the plurality of pairs of the same phrases by using differences in time stamps; and

determining a portion of the audio data to be removed using results of the clustering.

17. The computer program product of claim 16 , wherein the clustering the plurality of pairs of the same phrases includes clustering the plurality of pairs of the same phrases while allowing a difference in time stamps less than a predetermined threshold for each pair.

18. The computer program product of claim 17 , wherein the predetermined threshold is determined depending on the portion of the audio data to be removed.

19. The computer program product of claim 16 , wherein the computer-implemented method further comprises preparing a training data for training a model by removing the portion of the audio data from the audio data.

20. The computer program product of claim 16 , wherein the making the plurality of pairs of the same phrases includes, for each pair of the same phrases, obtaining a sum of differences in time stamps between corresponding words in the pair of the same phrases as the difference in time stamps.

Assignments (2)
CORRECTIVE ASSIGNMENT TO CORRECT THE THIRD ASSIGNOR'S NAME PREVIOUSLY RECORDED AT REEL: 052367 FRAME: 0740. ASSIGNOR(S) HEREBY CONFIRMS THE ASSIGNMENT. Recorded Apr 13, 2020
From: ITOH, NOBUYASU; KURATA, GAKUTO; SUZUKI, MASAYUKI
To: INTERNATIONAL BUSINESS MACHINES CORPORATION
Reel/Frame 052382/0705 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Apr 10, 2020
From: ITOH, NOBUYASU; KURATA, GAKUTO
To: INTERNATIONAL BUSINESS MACHINES CORPORATION
Reel/Frame 052367/0740 →
Continuity (1)
Related Publication 20210319787A1 · Oct 14, 2021