IP Library Granted Patent US 12,573,405
Granted Patent B2
US 12,573,405 · App. 18/131,021 · Granted Mar 10, 2026

Word correction using automatic speech recognition (ASR) incremental response

Inventor: Jeffry Copps Robert Jose (Tamil Nadu, IN)
Assignee: Adeia Guides Inc.
G10L15/32G10L15/1815G10L15/22G10L2015/088G10L2015/223
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,573,405
App. No.
18/131,021
Granted
Mar 10, 2026
Kind
B2
Abstract

An exemplary automatic speech recognition (ASR) system may receive an audio input including a segment of speech. The segment of speech may be independently processed by general ASR and domain-specific ASR to generate multiple ASR results. A selection between the multiple ASR results may be performed based on respective confidence levels for the general ASR and domain-specific ASR. As incremental ASR is performed, a composite result may be generated based on general ASR and domain-specific ASR.

Claims (53)

1 . A method, comprising:

receiving, by a control circuitry, an incremental speech input of a user;

transcoding the incremental speech input into a format compatible with speech recognition processing by the control circuitry;

identifying, by a first speech recognition, a first additional candidate word based on the incremental speech input;

determining a first confidence value for the first additional candidate word;

identifying, by a second speech recognition, a second additional candidate word based on the incremental speech input;

determining a second confidence value for the second additional candidate word;

determining a speech type and simultaneously a speech pattern of the incremental speech input, wherein the speech pattern comprises at least one of (i) terse speech or (ii) choppy speech;

modifying a weight associated with the first confidence value or a weight associated with the second confidence value based on the determined speech type and the speech pattern of the incremental speech input;

selecting at least one candidate word from each of the first additional candidate word and the second additional candidate word based on the first confidence value and the associated weight of the first confidence value and the second confidence value and the associated weight of the second confidence value; and

providing, to the control circuitry, the selected at least one candidate word as an incremental output.

2 . The method of claim 1 , further comprising:

updating, based on the incremental output, at least one of the first speech recognition and the second speech recognition.

3 . The method of claim 1 , wherein the at least one candidate word is further selected based on shared additional candidate words that are included in both the first speech recognition and the second speech recognition.

4 . The method of claim 1 , wherein the first speech recognition is a general speech recognition and the second speech recognition is a domain-specific speech recognition.

5 . The method of claim 4 , wherein the domain-specific speech recognition comprises a plurality of metadata types for each of a plurality of media assets.

6 . The method of claim 1 , wherein the speech pattern comprises at least one of: terse, choppy, narrative, and conversational.

7 . The method of claim 1 , wherein the first confidence value and the second confidence value are based on an edit distance between the incremental speech input and the first additional candidate word and second additional candidate word, as well as edit distances associated with prior words in the incremental speech input.

8 . The method of claim 1 , wherein at least one of the first speech recognition and the second speech recognition is a sequence aware speech recognition.

9 . The method of claim 1 , further comprising:

processing the incremental output by a media guidance application.

10 . The method of claim 1 , wherein the first speech recognition is sequence aware, and wherein the method further comprises:

identifying, by sequence unaware speech recognition, a third additional candidate word from the incremental speech input;

determining, by the sequence unaware speech recognition, a third confidence value, for the third additional candidate word;

comparing the third confidence value with one or more of the first confidence value and the second confidence value; and

selecting, based on the comparing of the third confidence value, at least one of the first additional candidate word, the second additional candidate word, and the third additional candidate word as the incremental output.

11 . A system for identifying words from speech of a user, the system comprising:

control circuitry configured to:

receive an incremental speech input of a user;

transcode the incremental speech input into a format compatible with speech recognition processing by the control circuitry;

identify, by a first speech recognition, a first additional candidate word based on the incremental speech input;

determine a first confidence value for the first additional candidate word;

identify, by a second speech recognition, a second additional candidate word based on the incremental speech input;

determine a second confidence value for the second additional candidate word;

determine a speech type and simultaneously a speech pattern of the incremental speech input, wherein the speech pattern comprises at least one of (i) terse speech or (ii) choppy speech;

modify a weight associated with the first confidence value or a weight associated with the second confidence value based on the determined speech type and the speech pattern of the incremental speech input;

select at least one candidate word from each of the first additional candidate word and the second additional candidate word based on the first confidence value and the associated weight of the first confidence value and the second confidence value and the associated weight of the second confidence value; and

provide the selected at least one candidate word as an incremental output.

12 . The system of claim 11 , wherein the control circuitry is further configured to:

update, based on the incremental output, at least one of the first speech recognition and the second speech recognition.

13 . The system of claim 11 , wherein the at least one candidate word is further selected based on shared additional candidate words that are included in both the first speech recognition and the second speech recognition.

14 . The system of claim 11 , wherein the first speech recognition is a general speech recognition and the second speech recognition is a domain-specific speech recognition.

15 . The system of claim 14 , wherein the domain-specific speech recognition comprises a plurality of metadata types for each of a plurality of media assets.

16 . The system of claim 11 , wherein the speech pattern comprises at least one of: terse, choppy, narrative, and conversational.

17 . The system of claim 11 , wherein the first confidence value and the second confidence value are based on an edit distance between the incremental speech input and the first additional candidate word and second additional candidate word, as well as edit distances associated with prior words in the incremental speech input.

18 . The system of claim 11 , wherein at least one of the first speech recognition and the second speech recognition is a sequence aware speech recognition.

19 . The system of claim 11 , wherein the control circuitry is further configured to:

process the incremental output by a media guidance application.

20 . The system of claim 11 , wherein the first speech recognition is sequence aware, and wherein the control circuitry is further configured to:

identify, by sequence unaware speech recognition, a third additional candidate word from the incremental speech input;

determine, by the sequence unaware speech recognition, a third confidence value, for the third additional candidate word;

compare the third confidence value with one or more of the first confidence value and the second confidence value; and

select, based on the comparing of the third confidence value, at least one of the first additional candidate word, the second additional candidate word, and the third additional candidate word as the incremental output.

Assignments (2)
CHANGE OF NAME Recorded Oct 3, 2024
From: ROVI GUIDES, INC.
To: ADEIA GUIDES INC.
Reel/Frame 069106/0231 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Apr 5, 2023
From: ROBERT JOSE, JEFFRY COPPS
To: ROVI GUIDES, INC.
Reel/Frame 063228/0623 →
Continuity (3)
Continuation 17384970 · Jul 26, 2021
Continuation 16408260 · May 9, 2019
Related Publication 20230252997A1 · Aug 10, 2023
References Cited (40)
US 4805219A · Baker · 1989 [cited by examiner]
US 5566272A · Brems et al. · 1996 [cited by applicant]
US 5761687A · Hon · 1998 [cited by examiner]
US 6122613A · Baker · 2000 [cited by examiner]
US 8650031B1 · Mamou · 2014 [cited by examiner]
US 8775177B1 · Heigold · 2014 [cited by examiner]
US 8972253B2 · Deng et al. · 2015 [cited by applicant]
US 9336771B2 · Chelba · 2016 [cited by examiner]
US 11081104B1 · Su · 2021 [cited by examiner]
US 11107475B2 · Robert Jose · 2021 [cited by applicant]
US 11138205B1 · Mohajer · 2021 [cited by examiner]
US 11651775B2 · Robert Jose · 2023 [cited by applicant]
US 20040220813A1 · Weng et al. · 2004 [cited by applicant]
US 20040249634A1 · Degani · 2004 [cited by examiner]
US 20040260543A1 · Horowitz · 2004 [cited by examiner]
US 20080004877A1 · Tian · 2008 [cited by examiner]
US 20100305947A1 · Schwarz · 2010 [cited by examiner]
US 20110055256A1 · Phillips · 2011 [cited by examiner]
US 20110283189A1 · McCarty · 2011 [cited by examiner]
US 20120265819A1 · McGann · 2012 [cited by examiner]
US 20130073286A1 · Bastea-Forte · 2013 [cited by examiner]
US 20130289987A1 · Ganapathiraju · 2013 [cited by examiner]
US 20140191939A1 · Penn · 2014 [cited by examiner]
US 20140270114A1 · Kolbegger · 2014 [cited by examiner]
US 20150012271A1 · Peng · 2015 [cited by examiner]
US 20150120296A1 · Stern · 2015 [cited by examiner]
US 20150169284A1 · Quast · 2015 [cited by examiner]
US 20160140956A1 · Yu et al. · 2016 [cited by applicant]
US 20170140752A1 · Sugitani · 2017 [cited by examiner]
US 20180082677A1 · Yaghi · 2018 [cited by examiner]
US 20180096678A1 · Zhou · 2018 [cited by examiner]
US 20180211652A1 · Mun · 2018 [cited by examiner]
US 20190013008A1 · Kunitake · 2019 [cited by examiner]
US 20200160838A1 · Lee · 2020 [cited by examiner]
US 20200310749A1 · Miller · 2020 [cited by examiner]
US 20200357412A1 · Robert Jose · 2020 [cited by applicant]
US 20210350807A1 · Robert Jose · 2021 [cited by applicant]
US 20230252997A1 · Robert Jose · 2023 [cited by applicant]
EP 2838085A1 · 2015 [cited by applicant]
PCT International Search Report for International Application No. PCT/US2020/032148, dated Jul. 1, 2020 (13 pages). [cited by applicant]