IP Library Granted Patent US 11,107,475
Granted Patent B2
US 11,107,475 · App. 16/408,260 · Granted Aug 31, 2021

Word correction using automatic speech recognition (ASR) incremental response

Inventor: Jeffry Copps Robert Jose (Tamil Nadu, IN)
Assignee: Rovi Guides, Inc.
G10L15/32G10L15/1815G10L15/22G10L2015/088G10L2015/223
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,107,475
App. No.
16/408,260
Granted
Aug 31, 2021
Kind
B2
Abstract

An exemplary automatic speech recognition (ASR) system may receive an audio input including a segment of speech. The segment of speech may be independently processed by general ASR and domain-specific ASR to generate multiple ASR results. A selection between the multiple ASR results may be performed based on respective confidence levels for the general ASR and domain-specific ASR. As incremental ASR is performed, a composite result may be generated based on general ASR and domain-specific ASR.

Claims (67)

1. A method for identifying words from speech of a user, the method comprising:

receiving a segment of audio corresponding to the speech of the user;

identifying, by general speech recognition, a first plurality of candidate words from the segment;

determining, by the general speech recognition, a first plurality of confidence values, wherein each of the first plurality of confidence values is associated with one of the first plurality of candidate words;

identifying, by domain-specific speech recognition, a second plurality of candidate words from the segment;

determining, by the domain-specific speech recognition, a second plurality of confidence values, wherein each of the second plurality of confidence values is associated with one of the second plurality of candidate words;

comparing each of the first plurality of confidence values with one or more of the second plurality of confidence values;

selecting, based on the comparing, at least one of the first plurality of candidate words and at least one of the second plurality of candidate words; and

identifying a composite plurality of words for the segment based on the selected candidate words.

2. The method of claim 1 , wherein the segment of audio corresponds to a first incremental speech input, further comprising:

receiving a second incremental speech input;

identifying, by the general speech recognition, a first additional candidate word from the second incremental speech input;

determining, by the general speech recognition, a first additional confidence value associated with the first additional candidate word;

identifying, by the domain-specific speech recognition, a second additional candidate word from the second incremental speech input;

determining, by the domain-specific speech recognition, a second additional confidence value associated with the second additional candidate word;

comparing the first additional confidence value to the second additional confidence value;

selecting the first additional candidate word or the second additional candidate word based on the comparison; and

updating the composite plurality of words with the selected first additional candidate word or second additional candidate word.

3. The method of claim 2 , further comprising:

updating at least one of the first plurality of confidence values or one of the second plurality of confidence values based on the second incremental speech input; and

updating the composite plurality of words based on the updated confidence values.

4. The method of claim 2 , wherein the selection of the first additional candidate word or the second additional candidate word is based on the composite plurality of words for the segment prior to the updating.

5. The method of claim 1 , wherein the first plurality of confidence values are based on a likelihood of a match of each of the first plurality of candidate words by the general speech recognition.

6. The method of claim 1 , wherein the second plurality of confidence values are based on a likelihood of a match of each of the second plurality of candidate words by the domain-specific speech recognition.

7. The method of claim 1 , wherein a domain of the domain-specific speech recognition comprises a media guidance application.

8. The method of claim 7 , wherein the domain-specific speech recognition comprises a plurality of metadata types for each of a plurality of media assets.

9. The method of claim 8 , wherein the plurality of metadata types comprise title, genre, and character.

10. The method of claim 1 , wherein the general speech recognition is sequence aware, and wherein the method further comprises:

identifying, by sequence unaware speech recognition, a third plurality of candidate words from the segment;

determining, by the sequence unaware speech recognition, a third plurality of confidence values, wherein each of the third plurality of confidence values is associated with one of the third plurality of candidate words;

comparing each of the third plurality of confidence values with one or more of the first plurality of confidence values or with one or more of the second plurality of confidence values;

selecting, based on the comparing of the third plurality of confidence values, at least one of the first plurality of candidate words, at least one of the second plurality of candidate words, and at least one of the third plurality of candidate words; and

identifying a composite plurality of words for the segment based on the selected candidate words.

11. A system for identifying words from speech of a user, the system comprising:

control circuitry configured to:

receive a segment of audio corresponding to the speech of the user;

identify, by general speech recognition, a first plurality of candidate words from the segment;

determine, by the general speech recognition, a first plurality of confidence values, wherein each of the first plurality of confidence values is associated with one of the first plurality of candidate words;

identify, by domain-specific speech recognition, a second plurality of candidate words from the segment;

determine, by the domain-specific speech recognition, a second plurality of confidence values, wherein each of the second plurality of confidence values is associated with one of the second plurality of candidate words;

compare each of the first plurality of confidence values with one or more of the second plurality of confidence values;

select, based on the comparison, at least one of the first plurality of candidate words and at least one of the second plurality of candidate words; and

identify a composite plurality of words for the segment based on the selected candidate words.

12. The system of claim 11 , wherein the segment of audio corresponds to a first incremental speech input, and wherein the control circuitry is further configured to:

receive a second incremental speech input;

identify, by the general speech recognition, a first additional candidate word from the second incremental speech input;

determine, by the general speech recognition, a first additional confidence value associated with the first additional candidate word;

identify, by the domain-specific speech recognition, a second additional candidate word from the second incremental speech input;

determine, by the domain-specific speech recognition, a second additional confidence value associated with the second additional candidate word;

compare the first additional confidence value to the second additional confidence value;

select the first additional candidate word or the second additional candidate word based on the comparison; and

update the composite plurality of words with the selected first additional candidate word or second additional candidate word.

13. The system of claim 12 , wherein the control circuitry is further configured to:

update at least one of the first plurality of confidence values or one of the second plurality of confidence values based on the second incremental speech input; and

update the composite plurality of words based on the updated confidence values.

14. The system of claim 12 , wherein the selection of the first additional candidate word or the second additional candidate word is based on the composite plurality of words for the segment prior to the updating.

15. The system of claim 11 , wherein the first plurality of confidence values are based on a likelihood of a match of each of the first plurality of candidate words by the general speech recognition.

16. The system of claim 11 , wherein the second plurality of confidence values are based on a likelihood of a match of each of the second plurality of candidate words by the domain-specific speech recognition.

17. The system of claim 11 , wherein a domain of the domain-specific speech recognition comprises a media guidance application.

18. The system of claim 17 , wherein the domain-specific speech recognition comprises a plurality of metadata types for each of a plurality of media assets.

19. The system of claim 18 , wherein the plurality of metadata types comprise title, genre, and character.

20. The system of claim 11 , wherein the general speech recognition is sequence aware, and wherein the control circuitry is further configured to:

identify, by sequence unaware speech recognition, a third plurality of candidate words from the segment;

determine, by the sequence unaware speech recognition, a third plurality of confidence values, wherein each of the third plurality of confidence values is associated with one of the third plurality of candidate words;

compare each of the third plurality of confidence values with one or more of the first plurality of confidence values or with one or more of the second plurality of confidence values;

select, based on the comparison of the third plurality of confidence values, at least one of the first plurality of candidate words, at least one of the second plurality of candidate words, and at least one of the third plurality of candidate words; and

identify a composite plurality of words for the segment based on the selected candidate words.

Assignments (7)
CHANGE OF NAME Recorded Oct 3, 2024
From: ROVI GUIDES, INC.
To: ADEIA GUIDES INC.
Reel/Frame 069106/0231 →
RELEASE OF SECURITY INTEREST Recorded Jun 5, 2020
From: MORGAN STANLEY SENIOR FUNDING, INC.
To: ROVI SOLUTIONS CORPORATION; ROVI TECHNOLOGIES CORPORATION; ROVI GUIDES, INC.; TIVO SOLUTIONS, INC.; VEVEO, INC.
Reel/Frame 053481/0790 →
RELEASE OF SECURITY INTEREST Recorded Jun 5, 2020
From: HPS INVESTMENT PARTNERS, LLC
To: ROVI SOLUTIONS CORPORATION; ROVI TECHNOLOGIES CORPORATION; ROVI GUIDES, INC.; TIVO SOLUTIONS, INC.; VEVEO, INC.
Reel/Frame 053458/0749 →
SECURITY INTEREST Recorded Jun 1, 2020
From: ROVI SOLUTIONS CORPORATION; ROVI TECHNOLOGIES CORPORATION; ROVI GUIDES, INC.; TIVO SOLUTIONS INC.; VEVEO, INC.; INVENSAS CORPORATION; INVENSAS BONDING TECHNOLOGIES, INC.; TESSERA, INC.; TESSERA ADVANCED TECHNOLOGIES, INC.; DTS, INC.; PHORUS, INC.; IBIQUITY DIGITAL CORPORATION
To: BANK OF AMERICA, N.A.
Reel/Frame 053468/0001 →
PATENT SECURITY AGREEMENT Recorded Nov 25, 2019
From: ROVI SOLUTIONS CORPORATION; ROVI TECHNOLOGIES CORPORATION; ROVI GUIDES, INC.; TIVO SOLUTIONS, INC.; VEVEO, INC.
To: MORGAN STANLEY SENIOR FUNDING, INC., AS COLLATERAL AGENT
Reel/Frame 051110/0006 →
SECURITY INTEREST Recorded Nov 22, 2019
From: ROVI SOLUTIONS CORPORATION; ROVI TECHNOLOGIES CORPORATION; ROVI GUIDES, INC.; TIVO SOLUTIONS, INC.; VEVEO, INC.
To: HPS INVESTMENT PARTNERS, LLC, AS COLLATERAL AGENT
Reel/Frame 051143/0468 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jul 22, 2019
From: ROBERT JOSE, JEFFRY COPPS
To: ROVI GUIDES, INC.
Reel/Frame 049825/0106 →
Continuity (1)
Related Publication 20200357412A1 · Nov 12, 2020
Cited By (1)
US 12,573,405