IP Library Granted Patent US 9,536,519
Granted Patent B2
US 9,536,519 · App. 14/926,544 · Granted Jan 3, 2017

Method and apparatus to generate a speech recognition library

View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 9,536,519
App. No.
14/926,544
Granted
Jan 3, 2017
Kind
B2
Abstract

Methods and apparatus to generate a speech recognition library for use by a speech recognition system are disclosed. An example method comprises identifying a plurality of video segments having closed caption data corresponding to a phrase, the plurality of video segments associated with respective ones of a plurality of audio data segments, computing a plurality of difference metrics between a baseline audio data segment associated with the phrase and respective ones of the plurality of audio data segments, selecting a set of the plurality of audio data segments based on the plurality of difference metrics, identifying a first one of the audio data segments in the set as a representative audio data segment, determining a first phonetic transcription of the representative audio data segment, and adding the first phonetic transcription to a speech recognition library when the first phonetic transcription differs from a second phonetic transcription associated with the phrase in the speech recognition library.

Claims (49)

1. A device, comprising:

a processing system including a processor; and

a memory that stores executable instructions that, when executed by the processing system, facilitate performance of operations, comprising:

obtaining video media content, wherein the video media content comprises images, audio content, and closed captioning of text from the audio content;

detecting an occurrence of a textual phrase in the closed captioning data of the video media content as a detected occurrence;

obtaining an audio segment from the audio content corresponding to the textual phrase as a selected audio segment;

computing a phonetic transcription for the selected audio segment as a computed transcription;

selecting, from a speech recognition library, a plurality of identified phonetic transcriptions associated with the textual phrase, wherein the speech recognition library comprises audio pronunciation data for the textual phrase and identified phonetic transcriptions of the textual phrase;

comparing the computed transcription with the plurality of identified phonetic transcriptions from the speech recognition library;

determining if the computed transcription differs from the plurality of identified phonetic transcriptions from the speech recognition library; and

responsive to determining that the computed transcription differs from the plurality of identified phonetic transcriptions, adding the computed transcription and the textual phrase to the audio pronunciation data in a group of the plurality of identified phonetic transcriptions in the speech recognition library.

2. The device of claim 1 , wherein the speech recognition library arranges phrase data for the textual phrase to include the textual phrase, the computed transcription, and audio data of the detected occurrence.

3. The device of claim 1 , wherein the operations further comprise:

detecting a second occurrence of the textual phrase in the closed captioning as a detected second occurrence;

obtaining a second audio segment corresponding to the textual phrase as a second selected audio segment; and

computing a second phonetic transcription for the second selected audio segment as a second computed transcription.

4. The device of claim 3 , wherein the operations further comprise:

comparing the second computed transcription with the plurality of identified phonetic transcriptions from the speech recognition library;

determining if the second computed transcription differs from the identified phonetic transcriptions from the speech recognition library; and

responsive to determining that the second computed transcription differs from the plurality of identified phonetic transcriptions, adding the second computed transcription and the textual phrase to the audio pronunciation data in the speech recognition library.

5. The device of claim 3 , wherein the operations comprise adding the second audio segment to the speech recognition library in relational association with the textual phrase.

6. The device of claim 3 , wherein the operations comprise comparing the second audio segment to a second baseline audio pronunciation associated with the textual phrase from the speech recognition library.

7. A non-transitory, machine-readable storage medium, comprising executable instructions that, when executed by a processing system including a processor, facilitate performance of operations, comprising:

obtaining video media content, wherein video media content comprises images, audio content, and closed captioning of text from the audio content;

obtaining a textual phrase in the closed captioning data as a detected occurrence, wherein the textual phrase is associated with an audio data segment from the audio content of the video media content;

computing a phonetic transcription for the detected occurrence as a computed transcription;

selecting, from a speech recognition library, identified phonetic transcriptions associated with the textual phrase, wherein the speech recognition library comprises audio pronunciation data for the textual phrase and identified phonetic transcriptions of the textual phrase;

calculating a difference metric between the computed transcription and the identified phonetic transcriptions associated with the textual phrase from the speech recognition library; and

responsive to the difference metric exceeding a threshold between the identified phonetic transcriptions and the computed transcription, adding the computed transcription and the textual phrase to a group including the identified phonetic transcriptions and the computed transcription associated with the textual phrase in the audio pronunciation data in the speech recognition library.

8. The non-transitory, machine-readable storage medium described in claim 7 , wherein the operations comprise receiving the textual phrase from an electronic program guide.

9. The non-transitory, machine-readable storage medium described in claim 7 , wherein the operations comprise adding the audio data segment to the speech recognition library.

10. The non-transitory, machine-readable storage medium described in claim 7 , wherein the operations further comprise:

detecting a second occurrence of the textual phrase in the closed captioning as a detected second occurrence;

obtaining a second audio segment corresponding to the textual phrase as a selected audio segment; and

computing a second phonetic transcription for the second occurrence as a second computed transcription.

11. The non-transitory, machine-readable storage medium described in claim 10 , wherein the operations further comprise calculating a difference metric between the second computed transcription and the identified phonetic transcriptions associated with the textual phrase from the speech recognition library.

12. The non-transitory, machine-readable storage medium described in claim 11 , wherein the operations further comprise, responsive to the difference metric exceeding the threshold, adding the second computed transcription and the textual phrase to the audio pronunciation data in the speech recognition library.

13. The non-transitory, machine-readable storage medium described in claim 11 , wherein the operations comprise adding the second computed transcription to the speech recognition library in relational association with the textual phrase.

14. The non-transitory, machine-readable storage medium described in claim 11 , wherein the difference metric comprises one of a mean-square error, a difference in formants, or a linear predictive coding coefficient difference.

15. A method, comprising:

obtaining, by a processing system including a processor, a textual phrase responsive to detecting a difference in pronunciation between a phonetic transcription from a video media source and a baseline phonetic transcription associated with the textual phrase from a speech recognition library, wherein the speech recognition library comprises the baseline phonetic transcription and collected phonetic pronunciation transcriptions of the textual phrase, and wherein the video media source comprises image data, audio data, and closed captioning data; and

responsive to detecting, by the processing system, the difference in the pronunciation:

storing, by the processing system, a phonetic transcription of the textual phrase in the speech recognition library as one of the phonetic pronunciation transcriptions to populate the speech recognition library; and

adding, by the processing system, an audio segment from the video media source corresponding to the collected phonetic pronunciation transcriptions associated with the textual phrase in the speech recognition library to populate the collected phonetic pronunciation transcriptions associated with the textual phrase.

16. The method of claim 15 , wherein the textual phrase comprises one of a proper name, a title, or a location.

17. The method of claim 15 , wherein the textual phrase is received from an electronic program guide.

18. The method of claim 15 , further comprising detecting, by the processing system, occurrences of the textual phrase in the closed captioning of the video media source.

19. The method of claim 18 , further comprising selecting, by the processing system, the audio segment from the video media source where the textual phrase is detected in the closed captioning data.

20. The method of claim 15 , wherein the textual phrase comprises a single word, and wherein the difference in pronunciation is of a single syllable of the single word.

Assignments (4)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Nov 13, 2023
From: NUANCE COMMUNICATIONS, INC.
To: MICROSOFT TECHNOLOGY LICENSING, LLC
Reel/Frame 065552/0934 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jan 26, 2017
From: AT&T INTELLECTUAL PROPERTY I, L.P.
To: NUANCE COMMUNICATIONS, INC.
Reel/Frame 041504/0952 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Nov 16, 2015
From: CHANG, HISAO M.
To: AT&T KNOWLEDGE VENTURES, LP
Reel/Frame 037045/0959 →
CHANGE OF NAME Recorded Nov 16, 2015
From: AT&T KNOWLEDGE VENTURES, L.P.
To: AT&T INTELLECTUAL PROPERTY I, LP
Reel/Frame 037112/0613 →