IP Library Granted Patent US 7,725,318
Granted Patent B2
US 7,725,318 · App. 11/195,144 · Granted May 25, 2010

System and method for improving the accuracy of audio searching

View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 7,725,318
App. No.
11/195,144
Granted
May 25, 2010
Kind
B2
Abstract

A system and method for improving the accuracy of audio searching using multiple models to process an audio file or stream to obtain search tracks. The search tracks are processed to locate at least one search term and generate multiple search results. The number of search results is equivalent to the number of models used to process the audio stream. The search results are combined to generate a unified search result. The multiple models may represent different languages, dialects and accents.

Claims (43)

1. A method for improving the searching of an audio stream with improved accuracy, the method comprising:

gathering the audio stream carrying voice of an unknown speaker, by a call recording system;

determining a plurality of acoustic models;

indexing said audio stream using said plurality of acoustic models to generate a plurality of phonetic search tracks, at least one of the plurality of phonetic search tracks comprising a first sequence of phonemes;

collecting at least one keyword;

processing said plurality of phonetic search tracks and said at least one keyword to obtain a plurality of search results by matching a pattern of phonemes in the at least one keyword with a pattern of phonemes in each of said plurality of phonetic search tracks, such that each of said plurality of search results corresponds to one of said plurality of acoustic models, and each of said plurality of search results indicates whether the at least one keyword was found in one of said plurality of search tracks, wherein each of said plurality of search results includes at least one hit indicating detection of the at least one keyword within one of said plurality of phonetic search tracks, the at least one hit having a time offset; and

combining said plurality of search results into a unified search result˜said combining comprising:

grouping at least two hits having time offsets which differ in at most a predetermined threshold into a cluster; and

determining a single hit from the cluster as the unified search result, the single hit indicating that the at least one keyword appears in the audio stream; and

wherein each of said plurality of acoustic models represents a language or dialect.

2. The method according to claim 1 , wherein each of said plurality of acoustic models is a phonetic acoustic model.

3. The method according to claim 1 , wherein the at least one hit comprises a confidence score.

4. The method according to claim 3 , wherein the step of combining said search results further comprises determining a resultant confidence score for cluster.

5. The method according to claim 4 , wherein the step of determining a resultant confidence score includes computing a simple average.

6. The method according to claim 4 , wherein the step of determining a resultant confidence score includes computing a weighted average.

7. The method according to claim 4 , wherein the step of determining a resultant confidence score includes computing a maximal confidence.

8. The method according to claim 4 , wherein the step of determining a resultant confidence score includes computing a confidence value with a non-linear rule.

9. The method according to claim 1 wherein each of the plurality of acoustic models represents phonemes.

10. A method for searching an audio streams with improved accuracy, the method comprising:

gathering the audio stream carrying voice of an unknown speaker, by a call recording system;

determining a plurality of acoustic models;

reducing said plurality of acoustic models using a language determining module;

indexing said audio stream using said plurality of acoustic models to generate a plurality of phonetic search tracks, at least one of the plurality of phonetic search tracks comprising a first sequence of phonemes;

collecting at least one keyword;

processing said plurality of phonetic search tracks and said at least one keyword to obtain a plurality of search results by matching a pattern of phonemes in the at least one keyword with a pattern of phonemes in each of said plurality of phonetic search tracks, such that each of said plurality of search results corresponds to one of said plurality of acoustic models, and wherein each of said plurality of search results indicates whether the at least one keyword was found in one of said plurality of phonetic search tracks, wherein each of said plurality of search results includes at least one hit indicating detection of the at least one keyword within one of said plurality of phonetic search tracks, the at least one hit having a time offset, and;

combining said plurality of search results into a unified search result, said combining comprising:

grouping hits having time offsets which differ in at most a predetermined threshold into a cluster; and

determining a single hit from the cluster, the single hit indicating that the at least one keyword appears in the audio stream; and

wherein each of said plurality of acoustic models represents a language or dialect.

11. The method according to claim 10 further comprising:

training for estimating the plurality of acoustic models, wherein at least one of the plurality of models is in a target language; and

testing for determining a probabilistic score that a speech utterance signal is in the target language.

12. The method according to claim 11 , wherein said training further comprises:

inputting a plurality of speech utterance signals corresponding to a plurality of target languages;

processing said plurality of speech utterance signals to extract feature vectors from said plurality of signals; and

estimating the plurality of acoustic models from said feature vectors.

13. The method according to claim 11 , wherein testing further comprises the steps of:

inputting the speech utterance signal;

processing the speech utterance signal to extract feature vectors from said signal; and

applying a pattern matching technique to the speech utterance signal to calculate a probabilistic score.

14. The method according to claim 13 , wherein said pattern matching technique is performed by an algorithm.

15. The method according to claim 13 , wherein said probabilistic score represents the likelihood that said speech utterance was spoken in the target language.

16. The method according to claim 10 wherein each of the plurality of acoustic models represents phonemes.

Assignments (3)
SECURITY INTEREST Recorded Feb 26, 2026
From: NICE LTD; NICE SYSTEMS INC.; NICE SYSTEMS TECHNOLOGIES INC.; INCONTACT, INC.
To: JPMORGAN CHASE BANK, N.A., AS ADMINISTRATIVE AGENT
Reel/Frame 074986/0208 →
PATENT SECURITY AGREEMENT Recorded Dec 6, 2016
From: NICE LTD.; NICE SYSTEMS INC.; AC2 SOLUTIONS, INC.; ACTIMIZE LIMITED; INCONTACT, INC.; NEXIDIA, INC.; NICE SYSTEMS TECHNOLOGIES, INC.
To: JPMORGAN CHASE BANK, N.A., AS ADMINISTRATIVE AGENT
Reel/Frame 040821/0818 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Dec 19, 2005
From: GAVALDA, MARSAL; WASSERBLAT, MOSHE
To: NICE SYSTEMS INC.
Reel/Frame 017368/0604 →