IP Library Granted Patent US 11,651,775
Granted Patent B2
US 11,651,775 · App. 17/384,970 · Granted May 16, 2023

Word correction using automatic speech recognition (ASR) incremental response

Inventor: Jeffry Copps Robert Jose (Tamil Nadu, IN)
Assignee: ROVI GUIDES, INC.
G10L15/32G10L15/1815G10L15/22G10L2015/088G10L2015/223
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,651,775
App. No.
17/384,970
Granted
May 16, 2023
Kind
B2
Abstract

An exemplary automatic speech recognition (ASR) system may receive an audio input including a segment of speech. The segment of speech may be independently processed by general ASR and domain-specific ASR to generate multiple ASR results. A selection between the multiple ASR results may be performed based on respective confidence levels for the general ASR and domain-specific ASR. As incremental ASR is performed, a composite result may be generated based on general ASR and domain-specific ASR.

Claims (82)

1. A method comprising:

identifying, based on a sequence aware general speech recognition application, a first plurality of candidate words corresponding to a speech-based audio input;

assigning, by the sequence aware general speech recognition application, a first respective plurality of confidence values to each of the first plurality of candidate words;

identifying, based on a sequence unaware general speech recognition application, a second plurality of candidate words corresponding to the speech-based audio input;

assigning, by the sequence unaware general speech recognition application, a second respective plurality of confidence values to each of the second plurality of candidate words;

comparing each of the first respective plurality of confidence values to each of the second respective plurality of confidence values;

selecting, based on the comparing, at least one of the first plurality of candidate words and at least one of the second plurality of candidate words; and

generating a composite plurality of words representing the speech-based audio input based on the selecting.

2. The method of claim 1 , further comprising:

identifying, from the speech-based audio input, a first speech segment and a second speech segment;

identifying, based on the sequence aware general speech recognition application, a first set of candidate words corresponding to the first segment;

assigning, by the sequence aware general speech recognition application, a first respective set of confidence values to each of the first set of candidate words;

for each of the first set of candidate words corresponding to the first segment:

identifying, based on domain-specific speech recognition, a second set of candidate words corresponding to the first segment;

assigning, by the domain-specific speech recognition, a second respective set of confidence values;

comparing each of the first respective set of confidence values to the second set of confidence values;

selecting, based on the comparing, at least one of the first set of candidate words or at least one of the second set of candidate words; and

updating a first segment of the composite plurality of words with the selected at least one of the first set of candidate words or the selected at least one of the second set of candidate words.

3. The method of claim 2 , wherein the selecting, based on the comparing, comprises selecting at least one of the first set of candidate words and at least one of the second set of candidate words.

4. The method of claim 2 , further comprising:

identifying, based on the sequence aware general speech recognition application, a third set of candidate words corresponding to the second segment;

assigning, by the sequence aware general speech recognition application, a third respective set of confidence values to each of the third set of candidate words;

for each of the third set of candidate words corresponding to the second segment:

identifying, based on domain-specific speech recognition, a fourth set of candidate words corresponding to the second segment;

assigning, by the domain-specific speech recognition, a fourth respective set of confidence values;

comparing each of the third respective set of confidence values to the fourth set of confidence values;

selecting, based on the comparing, at least one of the third set of candidate words or at least one of the fourth set of candidate words; and

updating a second segment of the composite plurality of words with the selected at least one of the third set of candidate words or the selected at least one of the fourth set of candidate words.

5. The method of claim 4 , wherein the selecting, based on the comparing, comprises selecting at least one of the third set of candidate words and at least one of the fourth set of candidate words.

6. The method of claim 1 , wherein each of the first respective plurality of confidence values and each of the second respective plurality of confidence values are based on a likelihood of a match of each of the first plurality of candidate words and each of the second plurality of candidate words, respectively, to the speech-based audio input.

7. The method of claim 2 , wherein the domain-specific specific speech recognition comprises a media guidance application.

8. The method of claim 7 , wherein the domain-specific speech recognition further comprises retrieving a plurality of metadata types for each of a plurality of media assets.

9. The method of claim 8 , the plurality of metadata types comprises at least one of title, genre, and character.

10. The method of claim 1 , further comprising:

for each of the second plurality of candidate words corresponding to the speech-based audio input:

identifying, based on domain-specific speech recognition, an alternative set of candidate words corresponding to the second plurality of candidate words;

assigning, by the domain-specific speech recognition, an alternative respective set of confidence values;

comparing each of the second respective plurality of confidence values to the alternative respective set of confidence values;

selecting, based on the comparing, at least one of the second plurality of candidate words or at least one of the alternative set of candidate words; and

updating the composite plurality of words with the selected at least one of the second plurality of candidate words or the selected at least one of the alternative set of candidate words.

11. A system comprising:

at least one input interface; and

control circuitry control circuitry configured to:

identify, based on a sequence aware general speech recognition application, a first plurality of candidate words corresponding to a speech-based audio input;

assign, by the sequence aware general speech recognition application, a first respective plurality of confidence values to each of the first plurality of candidate words;

identify, based on a sequence unaware general speech recognition application, a second plurality of candidate words corresponding to the speech-based audio input;

assign, by the sequence unaware general speech recognition application, a second respective plurality of confidence values to each of the second plurality of candidate words;

compare each of the first respective plurality of confidence values to each of the second respective plurality of confidence values;

select, based on the comparing, at least one of the first plurality of candidate words and at least one of the second plurality of candidate words; and

generate a composite plurality of words representing the speech-based audio input based on the selecting.

12. The system of claim 11 , wherein the control circuitry is further configured to:

identify, from the speech-based audio input, a first speech segment and a second speech segment;

identify, based on the sequence aware general speech recognition application, a first set of candidate words corresponding to the first segment;

assign, by the sequence aware general speech recognition application, a first respective set of confidence values to each of the first set of candidate words;

for each of the first set of candidate words corresponding to the first segment:

identify, based on domain-specific speech recognition, a second set of candidate words corresponding to the first segment;

assign, by the domain-specific speech recognition, a second respective set of confidence values;

compare each of the first respective set of confidence values to the second set of confidence values;

select, based on the comparing, at least one of the first set of candidate words or at least one of the second set of candidate words; and

update a first segment of the composite plurality of words with the selected at least one of the first set of candidate words or the selected at least one of the second set of candidate words.

13. The system of claim 12 , wherein the control circuitry configured to select, based on the comparing, is further configured to select at least one of the first set of candidate words and at least one of the second set of candidate words.

14. The system of claim 12 , wherein the control circuitry is further configured to:

identify, based on the sequence aware general speech recognition application, a third set of candidate words corresponding to the second segment;

assign, by the sequence aware general speech recognition application, a third respective set of confidence values to each of the third set of candidate words;

for each of the third set of candidate words corresponding to the second segment:

identify, based on domain-specific speech recognition, a fourth set of candidate words corresponding to the second segment;

assign, by the domain-specific speech recognition, a fourth respective set of confidence values;

compare each of the third respective set of confidence values to the fourth set of confidence values;

select, based on the comparing, at least one of the third set of candidate words or at least one of the fourth set of candidate words; and

update a second segment of the composite plurality of words with the selected at least one of the third set of candidate words or the selected at least one of the fourth set of candidate words.

15. The system of claim 14 , wherein the control circuitry configured to select, based on the comparing, is further configured to select at least one of the third set of candidate words and at least one of the fourth set of candidate words.

16. The system of claim 11 , wherein the control circuitry configured to assign each of the first respective plurality of confidence values and each of the second respective plurality of confidence values is configured to assign respective confidence values based on a likelihood of a match of each of the first plurality of candidate words and each of the second plurality of candidate words, respectively, to the speech-based audio input.

17. The system of claim 12 , wherein the control circuitry configured to identify, based on domain-specific speech recognition, is further configured to identify, based on domain-specific speech recognition, wherein the domain-specific speech recognition comprises a media guidance application.

18. The system of claim 17 , wherein the control circuitry configured to identify based on domain-specific speech recognition, is further configured to retrieve a plurality of metadata types for each of a plurality of media assets.

19. The system of claim 18 , wherein the control circuitry is further configured to retrieve a plurality of metadata types comprising at least one of title, genre, and character.

20. The system of claim 11 , wherein the control circuitry is further configured to:

for each of the second plurality of candidate words corresponding to the speech-based audio input:

identify, based on domain-specific speech recognition, an alternative set of candidate words corresponding to the second plurality of candidate words;

assign, by the domain-specific speech recognition, an alternative respective set of confidence values;

compare each of the second respective plurality of confidence values to the alternative respective set of confidence values;

select, based on the comparing, at least one of the second plurality of candidate words or at least one of the alternative set of candidate words; and

update the composite plurality of words with the selected at least one of the second plurality of candidate words or the selected at least one of the alternative set of candidate words.

Assignments (3)
CHANGE OF NAME Recorded Oct 3, 2024
From: ROVI GUIDES, INC.
To: ADEIA GUIDES INC.
Reel/Frame 069106/0231 →
SECURITY INTEREST Recorded May 19, 2023
From: ADEIA GUIDES INC.; ADEIA MEDIA HOLDINGS LLC; ADEIA MEDIA SOLUTIONS INC.; ADEIA SEMICONDUCTOR BONDING TECHNOLOGIES INC.; ADEIA SEMICONDUCTOR SOLUTIONS LLC; ADEIA SEMICONDUCTOR TECHNOLOGIES LLC
To: BANK OF AMERICA, N.A., AS COLLATERAL AGENT
Reel/Frame 063707/0884 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jul 26, 2021
From: ROBERT JOSE, JEFFRY COPPS
To: ROVI GUIDES, INC.
Reel/Frame 056975/0441 →
Continuity (2)
Continuation 16408260 · May 9, 2019
Related Publication 20210350807A1 · Nov 11, 2021
Cited By (1)
US 12,573,405