IP Library Granted Patent US 9,965,552
Granted Patent B2
US 9,965,552 · App. 15/055,975 · Granted May 8, 2018

System and method of lattice-based search for spoken utterance retrieval

View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 9,965,552
App. No.
15/055,975
Granted
May 8, 2018
Kind
B2
Abstract

A system and method are disclosed for retrieving audio segments from a spoken document. The spoken document preferably is one having moderate word error rates such as telephone calls or teleconferences. The method comprises converting speech associated with a spoken document into a lattice representation and indexing the lattice representation of speech. These steps are performed typically off-line. Upon receiving a query from a user, the method further comprises searching the indexed lattice representation of speech and returning retrieved audio segments from the spoken document that match the user query.

Claims (37)

1. A method comprising:

receiving data corresponding to a text query from a user;

retrieving a spoken document associated with the text query;

searching a word index of the spoken document associated with the text query using the text query, to yield first search results;

searching a sub-word index of the spoken document associated with the text query using the text query, to yield second search results; and

returning, via a computer network and according to the first search results and the second search results, audio segments from the spoken document associated with the text query which correspond to the text query.

2. The method of claim 1 , further comprising combining the first search results and the second search results to yield combined results.

3. The method of claim 2 , wherein returning the audio segments from the spoken document associated with the text query is further based on the combined results.

4. The method of claim 1 , wherein searching the word index of the spoken document associated with the text query using the text query further comprises searching the word index for an in-vocabulary portion of the text query.

5. The method of claim 1 , wherein searching the sub-word index of the spoken document associated with the text query using the text query further comprises searching the sub-word index for an out-of-vocabulary portion of the text query.

6. The method of claim 1 , wherein the text query comprises at least one text query word.

7. The method of claim 1 , wherein the text query is converted to text via automatic speech recognition.

8. A system comprising:

a processor; and

a computer-readable storage medium having instructions stored which, when executed by the processor, cause the processor to perform operations comprising:

receiving data corresponding to a text query from a user;

retrieving a spoken document associated with the text query;

searching a word index of the spoken document associated with the text query using the text query, to yield first search results;

searching a sub-word index of the spoken document associated with the text query using the text query, to yield second search results; and

returning, via a computer network and according to the first search results and the second search results, audio segments from the spoken document associated with the text query which correspond to the text query.

9. The system of claim 8 , the computer-readable storage medium having additional instructions stored which result in operations comprising combining the first search results and the second search results to yield combined results.

10. The system of claim 9 , wherein returning the audio segments from the spoken document associated with the text query is further based on the combined results.

11. The system of claim 8 , wherein searching the word index of the spoken document associated with the text query using the text query further comprises searching the word index for an in-vocabulary portion of the text query.

12. The system of claim 11 , wherein searching the sub-word index of the spoken document associated with the text query using the text query further comprises searching the sub-word index for an out-of-vocabulary portion of the text query.

13. The system of claim 12 , wherein the text query comprises at least one text query word.

14. The system of claim 8 , wherein the text query from the user is converted to text via automatic speech recognition.

15. A computer-readable storage device having instructions stored which, when executed by a processor, cause the processor to perform operations comprising:

receiving data corresponding to a text query from a user;

retrieving, based on the text query, a spoken document;

searching a word index of the spoken document associated with the text query using the text query, to yield first search results;

searching a sub-word index of the spoken document associated with the text query using the text query, to yield second search results; and

returning, via a computer network and according to the first search results and the second search results, audio segments from the spoken document associated with the text query which correspond to the text query.

16. The computer-readable storage device of claim 15 , having additional instructions stored which result in operations comprising combining the first search results and the second search results to yield combined results.

17. The computer-readable storage device of claim 16 , wherein returning the audio segments from the spoken document associated with the text query is further based on the combined results.

18. The computer-readable storage device of claim 15 , wherein searching the word index of the spoken document associated with the text query using the text query further comprises searching the word index for an in-vocabulary portion of the text query.

19. The computer-readable storage device of claim 15 , wherein searching the sub-word index of the spoken document associated with the text query using the text query further comprises searching the sub-word index for an out-of-vocabulary portion of the text query.

20. The computer-readable storage device of claim 15 , wherein the text query comprises at least one text query word.

Assignments (5)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Nov 9, 2023
From: NUANCE COMMUNICATIONS, INC.
To: MICROSOFT TECHNOLOGY LICENSING, LLC
Reel/Frame 065533/0389 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jan 26, 2017
From: AT&T INTELLECTUAL PROPERTY II, L.P.
To: NUANCE COMMUNICATIONS, INC.
Reel/Frame 041512/0608 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Aug 23, 2016
From: SARACLAR, MURAT; SPROAT, RICHARD WILLIAM
To: AT&T CORP.
Reel/Frame 039509/0458 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Aug 23, 2016
From: AT&T CORP.
To: AT&T PROPERTIES, LLC
Reel/Frame 039511/0604 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Aug 23, 2016
From: AT&T PROPERTIES, LLC
To: AT&T INTELLECTUAL PROPERTY II, L.P.
Reel/Frame 039511/0643 →