IP Library Granted Patent US 9,355,638
Granted Patent B2
US 9,355,638 · App. 14/737,708 · Granted May 31, 2016

System and method for improving speech recognition accuracy using textual context

View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 9,355,638
App. No.
14/737,708
Granted
May 31, 2016
Kind
B2
Abstract

Disclosed herein are systems, methods, and computer-readable storage media for improving speech recognition accuracy using textual context. The method includes retrieving a recorded utterance, capturing text from a device display associated with the spoken dialog and viewed by one party to the recorded utterance, and identifying words in the captured text that are relevant to the recorded utterance. The method further includes adding the identified words to a dynamic language model, and recognizing the recorded utterance using the dynamic language model. The recorded utterance can be a spoken dialog. A time stamp can be assigned to each identified word. The method can include adding identified words to and/or removing identified words from the dynamic language model based on their respective time stamps. A screen scraper can capture text from the device display associated with the recorded utterance. The device display can contain customer service data.

Claims (31)

1. A method comprising:

identifying, via a processor, words in optically-recognized text that are relevant to a recorded utterance based on references within the optically-recognized text to external data, to yield identified words;

adding the identified words to a dynamic language model to generate a modified dynamic language model; and

recognizing the recorded utterance using the modified dynamic language model, wherein the modified dynamic language model comprises a speech recognition model.

2. The method of claim 1 , wherein the external data comprises information about one of a product or a service.

3. The method of claim 2 , wherein a user has purchased the one of the product or the service, and wherein the optically-recognized text is on a display of an agent interacting with the user.

4. The method of claim 1 , wherein the external data comprises a link to a social media address.

5. The method of claim 1 , wherein the recorded utterance is a portion of a spoken dialog.

6. The method of claim 1 , wherein each identified word in the identified words is assigned a respective time stamp, and the method further comprising adding an identified word of the identified words to the modified dynamic language model based on the respective time stamp assigned to the identified word.

7. The method of claim 1 , wherein each identified word in the identified words is assigned a respective time stamp, and the method further comprising removing an identified word of the identified words from the modified dynamic language model based on the respective time stamp assigned to the identified word.

8. A system comprising:

a processor; and

a computer-readable storage medium having instructions stored which, when executed by the processor, cause the processor to perform operations comprising:

identifying words in optically-recognized text that are relevant to a recorded utterance based on references within the optically-recognized text to external data, to yield identified words;

adding the identified words to a dynamic language model to generate a modified dynamic language model; and

recognizing the recorded utterance using the modified dynamic language model, wherein the modified dynamic language model comprises a speech recognition model.

9. The system of claim 8 , wherein the external data comprises information about one of a product or a service.

10. The system of claim 9 , wherein a user has purchased the one of the product or the service, and wherein the optically-recognized text is on a display of an agent interacting with the user.

11. The system of claim 8 , wherein the external data comprises a link to a social media address.

12. The system of claim 8 , wherein the recorded utterance is a portion of a spoken dialog.

13. The system of claim 8 , wherein each identified word in the identified words is assigned a respective time stamp, and the computer-readable storage medium having instructions stored which, when executed by the processor, cause the processor to perform operations comprising adding an identified word of the identified words to the modified dynamic language model based on the respective time stamp assigned to the identified word.

14. The system of claim 8 , wherein each identified word in the identified words is assigned a respective time stamp, and the computer-readable storage medium having instructions stored which, when executed by the processor, cause the processor to perform operations comprising removing an identified word of the identified words from the modified dynamic language model based on the respective time stamp assigned to the identified word.

15. A computer-readable storage device having instructions stored which, when executed by a computing device, cause the computing device to perform operations comprising:

identifying words in optically-recognized text that are relevant to a recorded utterance based on references within the optically-recognized text to external data, to yield identified words;

adding the identified words to a dynamic language model to generate a modified dynamic language model; and

recognizing the recorded utterance using the modified dynamic language model, wherein the modified dynamic language model comprises a speech recognition model.

16. The computer-readable storage device of claim 15 , wherein the external data comprises information about one of a product or a service.

17. The computer-readable storage device of claim 16 , wherein a user has purchased the one of the product or the service, and wherein the optically-recognized text is on a display of an agent interacting with the user.

18. The computer-readable storage device of claim 15 , wherein the external data comprises a link to a social media address.

19. The computer-readable storage device of claim 15 , wherein the recorded utterance is a portion of a spoken dialog.

20. The computer-readable storage device of claim 15 , wherein each identified word in the identified words is assigned a respective time stamp, and the computer-readable storage device having additional instructions stored which, when executed by the computing device, cause the computing device to perform operations comprising adding an identified word of the identified words to the modified dynamic language model based on the respective time stamp assigned to the identified word.

Assignments (3)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Nov 9, 2023
From: NUANCE COMMUNICATIONS, INC.
To: MICROSOFT TECHNOLOGY LICENSING, LLC
Reel/Frame 065532/0152 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jan 26, 2017
From: AT&T INTELLECTUAL PROPERTY I, L.P.
To: NUANCE COMMUNICATIONS, INC.
Reel/Frame 041504/0952 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Dec 18, 2015
From: MELAMED, DAN; BANGALORE, SRINIVAS; JOHNSTON, MICHAEL
To: AT&T INTELLECTUAL PROPERTY I, L.P.
Reel/Frame 037333/0248 →