IP Library Granted Patent US 11,605,373
Granted Patent B2
US 11,605,373 · App. 17/875,964 · Granted Mar 14, 2023

System and method for combining phonetic and automatic speech recognition search

Inventors: William Mark Finlay (Tucker, GA); Robert William Morris (Decatur, GA); Peter S. Cardillo (Atlanta, GA); Maria Kunin (San Rafael, CA)
Assignee: Nice Ltd.
G10L15/02G10L15/08G10L2015/025
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,605,373
App. No.
17/875,964
Granted
Mar 14, 2023
Kind
B2
Abstract

A text search query including one or more words may be received. An ASR index created for an audio recording may be searched over using the query to produce ASR search results including words, each word associated with a confidence score. For each of the words in the ASR search results associated with a confidence score below a threshold (and in some cases having one or more preceding words in the ASR index and one or more subsequent words in the ASR index), a phonetic representation of the audio recording may be searched for the word having the confidence score below the threshold, where it occurs in the audio recording, possibly after the one or more preceding words and in the audio recording before the one or more subsequent words, to produce phonetic search results. Search results may be returned include ASR and phonetic results.

Claims (56)

1. A method for searching an audio recording using text, the method comprising:

accepting a text search query;

converting the text search query to a phonetic representation of the text search query;

searching over an automatic speech recognition (ASR) index created for an audio file using the text search query to produce ASR search results wherein the ASR index comprises textual representations of words, each textual representation associated with a confidence score;

for one or more words in the ASR search results associated with a confidence score below a threshold, searching over a portion of the phonetic representation of the audio file using the phonetic representation of the text search query to produce phonetic search results, wherein the phonetic representation is extended by a window comprised of at least one word before and at least one word after the one or more words; and

returning as search results ASR search results and phonetic search results wherein the searching further comprises:

processing each word in the ASR index sequentially;

starting a new time segment to be used with phonetic processing when an ASR word is encountered with confidence below a threshold, wherein the new time segment start is based on at least one of: start time of the current low confidence ASR word, start or end time of the previous high confidence ASR word; start or end time of the N previous confidence words; and fixed time window added to or subtracted from any of the aforementioned values;

continuing processing each word in the ASR and incrementing the phonetic portion until an ASR index word is found with a confidence score equal to or above the threshold or until the end of the ASR index is reached;

setting the end of the of the time sequence and adding a master list of media segments requiring a phonetic index; and

continuing to process each word until the entire ASR index has been processed.

2. The method of claim 1 , wherein the phonetic representation represents portions of the audio file corresponding to low confidence scores.

3. The method of claim 1 wherein the phonetic representation and ASR index are comprised 1, within a composite index, and wherein the phonetic representation represents only portions of the audio file comprising words associated with a confidence score below the threshold.

4. The method of claim 1 wherein the phonetic representation and ASR index are comprised within a composite index, and wherein the phonetic representation represents portions of the audio file comprising words associated with a confidence score below the threshold and an overlap portion including words associated with a confidence score not below the threshold.

5. The method of claim 1 , wherein the confidence score indicates the confidence that the word accurately represents the corresponding word in the audio recording.

6. The method of claim 3 , wherein the search results comprise a location in the audio recording corresponding to the text search query.

7. A method for searching an audio recording using text, the method comprising:

accepting a text search query comprising a plurality of words;

searching over an automatic speech recognition (ASR) index created for an audio recording using the text search query to produce ASR search results, the ASR search results comprising words, each word associated with a confidence score;

for one or more words comprised in the ASR search results associated with a confidence score below a threshold and having one or more preceding words in the ASR index and one or more subsequent words in the ASR index, searching over a portion of the phonetic representation of the audio recording for the word associated with a confidence score below the threshold where it occurs in the audio recording after the one or more preceding words and in the audio recording before the one or more subsequent words, to produce phonetic search results,

wherein the phonetic representation is extended by a window comprised of at least one word before and at least one word after the word position; and

returning as search results ASR search results and phonetic search results; the searching further comprising:

processing each word in the ASR index sequentially;

starting a new time segment to be used with phonetic processing when an ASR word is encountered with confidence below a threshold, wherein the new time segment start is based on at least one of: start time of the current low confidence ASR word, start or end time of the previous high confidence ASR word; start or end time of the N previous confidence words; and fixed time window added to or subtracted from any of the aforementioned values;

continuing processing each word in the ASR and incrementing the phonetic portion until an ASR index word is found with a confidence score equal to or above the threshold or until the end of the ASR index is reached;

setting the end of the of the time sequence and adding a master list of media segments requiring a phonetic index; and

continuing to process each word until the entire ASR index has been processed.

8. The method of claim 7 , comprising searching over a phonetic representation of the audio file before the end of a preceding word and after the beginning of a subsequent word, to produce phonetic search results.

9. The method of claim 7 wherein the phonemic representation and ASR index are comprised within a composite index, and wherein the phonetic representation represents only portions of the audio file comprising words associated with a confidence score below the threshold.

10. The method of claim 7 wherein the phonetic representation and ASR index are comprised within a composite index, and wherein the phonetic representation represents portions of the audio file comprising words associated with a confidence score below the threshold and an overlap portion including words associated with a confidence score not below the threshold.

11. The method of claim 7 , wherein the confidence score indicates the confidence that the word accurately represents the corresponding word in the audio recording.

12. The method of claim 7 , wherein the search results comprise a location in the audio recording corresponding to the text search query.

13. The method of claim 7 , wherein searching over the ASR index comprises:

converting the text search query to a phoneme representation; and

using the phoneme representation to access a phoneme sequence lookup table, to return an index to the ASR index.

14. A system for searching an audio recording using text, the system comprising:

a memory; and

a controller configured to:

accept a text search query comprising a plurality of words;

search over an automatic speech recognition (ASR) index created for an audio recording using the text search query to produce ASR search results, the ASR search results comprising words, each word associated with a confidence score;

for one or more words comprised in the ASR search results associated with a confidence score below a threshold and having one or more preceding words in the ASR index and one or more subsequent words in the ASR index, search over a portion of the phonetic representation of the audio recording for the word associated with a confidence score below the threshold where it occurs in the audio recording after the one or more preceding words and in the audio recording before the one or more subsequent words, to produce phonetic search results, wherein the phonetic representation is extended by a window comprised of at least one word before and at least one word after the position; and

returning as search results ASR search results and phonetic search results;

the searching further comprising:

processing each word in the ASR index sequentially;

starting a new time segment to be used with phonetic processing when an ASR word is encountered with confidence below a threshold, wherein the new time segment start is based on at least one of: start time of the current low confidence ASR word, start or end time of the previous high confidence ASR word; start or end time of the N previous confidence words; and fixed time window added to or subtracted from any of the aforementioned values;

continuing processing each word in the ASR and incrementing the phonetic portion until an ASR index word is found with a confidence score equal to or above the threshold or until the end of the ASR index is reached;

setting the end of the of the time sequence and adding a master list of media segments requiring a phonetic index; and

continuing to process each word until the entire ASR index has been processed.

15. The system of claim 14 , wherein the controller is configured to search over a phonetic representation of the audio file before the end of a preceding word and after the beginning of a subsequent word, to produce phonetic search results.

16. The system of claim 14 wherein the phonetic representation and ASR index are comprised within a composite index, and wherein the phonetic representation represents only portions of the audio file comprising words associated with a confidence score below the threshold.

17. The system of claim 14 wherein the phonetic representation and ASR index are comprised within a composite index, and wherein the phonetic representation represents portions of the audio file comprising words associated with a confidence score below the threshold and an overlap portion including words associated with a confidence score not below the threshold.

18. The system of claim 14 , wherein the confidence score indicates the confidence that the word accurately represents the corresponding word in the audio recording.

19. The system of claim 14 , wherein the search results comprise a location in the audio recording corresponding to the text search query.

20. The system of claim 14 , wherein searching over the ASR index comprises:

converting the text search query to a phoneme representation; and

using the phoneme representation to access a phoneme sequence lookup table, to return an index to the ASR index.

Assignments (2)
SECURITY INTEREST Recorded Feb 26, 2026
From: NICE LTD; NICE SYSTEMS INC.; NICE SYSTEMS TECHNOLOGIES INC.; INCONTACT, INC.
To: JPMORGAN CHASE BANK, N.A., AS ADMINISTRATIVE AGENT
Reel/Frame 074986/0208 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Nov 21, 2022
From: FINLAY, WILLIAM MARK; MORRIS, ROBERT WILLIAM; CARDILLO, PETER S.; KUNIN, MARIA MICHAELA
To: NICE LTD.
Reel/Frame 061841/0891 →