IP Library Granted Patent US 11,721,322
Granted Patent B2
US 11,721,322 · App. 16/804,201 · Granted Aug 8, 2023

Automated word correction in speech recognition systems

Inventors: Ankur Anil Aher (Maharashtra, IN); Jeffry Copps Robert Jose (Tamil Nadu, IN)
Assignee: Rovi Guides, Inc.
G10L15/01G10L15/063G10L15/08G10L15/22G10L2015/0636G10L2015/086
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,721,322
App. No.
16/804,201
Filed
Feb 28, 2020
Granted
Aug 8, 2023
Kind
B2
Examiner
HAN, QI
Art Unit
2659
USPC
704/270
Abstract

Systems and methods for correcting recognition errors in speech recognition systems are disclosed herein. Natural conversational variations are identified to determine whether a query intends to correct a speech recognition error or whether the query is a new command. When the query intends to correct a speech recognition error, the system identifies a location of the error and performs the correction. The corrected query can be presented to the user or be acted upon as a command for the system.

Claims (55)

1. A method for correcting a speech recognition error, the method comprising:

detecting a first voice input;

generating a text string based on the first voice input;

detecting a second voice input, wherein the second voice input was input within a predetermined time threshold of the generating the text string based on the first voice input;

determining whether the second voice input is directed to correcting the text string generated based on the first voice input, by:

identifying a first plurality of sound properties corresponding to the first voice input;

identifying a second plurality of sound properties corresponding to the second voice input;

comparing each of the first plurality of sound properties of the first voice input to each corresponding one of the second plurality of sound properties of the second voice input; and

determining, based on the comparing, whether the second voice input is directed to correcting the text string previously generated based on the first voice input when a number of the first plurality of sound properties that matches the second plurality of sound properties is greater than a predetermined threshold and at least one of the first plurality of sound properties does not match the corresponding one of the second plurality of sound properties; and

in response to the determining that the second voice input is directed to correcting the text string, modifying the text string based on the second voice input.

2. The method of claim 1 , wherein determining whether the second voice input is directed to correcting the text string comprises:

determining that at least one of an acoustic envelope, an intensity, a pitch, a frequency, or an amplitude of the second voice input exceeds a threshold.

3. The method of claim 1 , wherein determining that at least the portion of the first plurality of sound properties matches the at least the portion of the second plurality of sound properties comprises determining that an acoustic envelope of the at least the portion of the second plurality of sound properties matches an acoustic envelope of the at least the portion of the first plurality of sound properties.

4. The method of claim 1 , wherein the text string previously generated based on the first voice input is a first text string, and wherein determining whether the second voice input is directed to correcting the first text string comprises:

converting the second voice input to a second text string; and

determining, based on a database of misrecognized words, that a first word in the first text string and a second word in the second text string are misrecognized for one another.

5. The method of claim 4 , wherein modifying the first text string based on the second voice input comprises replacing the first word in the first text string with the second word.

6. The method of claim 1 , further comprising:

identifying, based at least in part on the sound property of the second voice input, a portion of the text string to be modified,

wherein the modifying the text string comprises modifying only the identified portion of the text string.

7. The method of claim 1 , wherein the sound property of the voice input comprises at least one of an acoustic envelope, an intensity, a pitch, a frequency, or an amplitude.

8. The method of claim 1 , wherein the sound property corresponds to a correction expression.

9. A system for correcting a speech recognition error, the system comprising:

memory; and

control circuitry configured to:

detect a first voice input;

generate a text string based on the first voice input;

detect a second voice input, wherein the second voice input was input within a predetermined time threshold of the generating the text string based on the first voice input;

determine whether the second voice input is directed to correcting the text string previously generated based on the first voice input, by:

identifying a first plurality of sound properties corresponding to the first voice input;

identifying a second plurality of sound properties corresponding to the second voice input;

compare each of the first plurality of sound properties of the first voice input to each corresponding one of the second plurality of sound properties of the second voice input; and

determining, based on the comparing, whether the second voice input is directed to correcting the text string previously generated based on the first voice input when a number of the first plurality of sound properties that matches the second plurality of sound properties is greater than a predetermined threshold and at least one of the first plurality of sound properties does not match the corresponding one of the second plurality of sound properties; and

in response to the determining that the second voice input is directed to correcting the text string, modify the text string based on the second voice input.

10. The system of claim 9 , wherein the control circuitry is configured to determine whether the second voice input is directed to correcting the text string by:

determining that at least one of an acoustic envelope, an intensity, a pitch, a frequency, or an amplitude of the second voice input exceeds a threshold.

11. The system of claim 9 , wherein the control circuitry is configured to determine that at least the portion of the first plurality of sound properties matches the at least the portion of the second plurality of sound properties by determining that an acoustic envelope of the at least the portion of the second plurality of sound properties matches an acoustic envelope of the at least the portion of the first plurality of sound properties.

12. The system of claim 9 , wherein the text string previously generated based on the first voice input is a first text string, and wherein the control circuitry is configured to determine whether the second voice input is directed to correcting the first text string by:

converting the second voice input to a second text string; and

determining, based on a database of misrecognized words, that a first word in the first text string and a second word in the second text string are misrecognized for one another.

13. The system of claim 12 , wherein the control circuitry is configured to modify the first text string based on the second voice input by replacing the first word in the first text string with the second word.

14. The system of claim 9 , wherein the control circuitry is further configured to:

identify, based at least in part on the sound property of the second voice input, a portion of the text string to be modified,

wherein the modifying the text string comprises modifying only the identified portion of the text string.

15. The system of claim 9 , wherein the sound property of the voice input comprises at least one of an acoustic envelope, an intensity, a pitch, a frequency, or an amplitude.

16. The system of claim 9 , wherein the sound property corresponds to a correction expression.

17. A method for correcting a speech recognition error, the method comprising:

detecting a first voice input;

generating a text string based on the first voice input;

detecting a second voice input, wherein the second voice input was input within a predetermined time threshold of the generating the text string based on the first voice input;

identifying a first plurality of sound properties corresponding to the first voice input;

identifying a second plurality of sound properties corresponding to the second voice input;

comparing each of the first plurality of sound properties of the first voice input to each corresponding one of the second plurality of sound properties of the second voice input;

determining, based on the comparing, whether a number of the first plurality of sound properties that matches the second plurality of sound properties is greater than a predetermined threshold and at least one of the first plurality of sound properties does not match the corresponding one of the second plurality of sound properties; and

in response to the determining that the number of the first plurality of sound properties matches the second plurality of sound properties is greater than the predetermined threshold and at least one of the first plurality of sound properties does not match the corresponding one of the second plurality of sound properties, modifying the text string based on the second voice input.

Assignments (3)
CHANGE OF NAME Recorded Sep 24, 2024
From: ROVI GUIDES, INC.
To: ADEIA GUIDES INC.
Reel/Frame 069032/0008 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Feb 8, 2022
From: AHER, ANKUR ANIL; ROBERT JOSE, JEFFRY COPPS
To: ROVI GUIDES, INC.
Reel/Frame 058921/0213 →
SECURITY INTEREST Recorded Jun 1, 2020
From: ROVI SOLUTIONS CORPORATION; ROVI TECHNOLOGIES CORPORATION; ROVI GUIDES, INC.; TIVO SOLUTIONS INC.; VEVEO, INC.; INVENSAS CORPORATION; INVENSAS BONDING TECHNOLOGIES, INC.; TESSERA, INC.; TESSERA ADVANCED TECHNOLOGIES, INC.; DTS, INC.; PHORUS, INC.; IBIQUITY DIGITAL CORPORATION
To: BANK OF AMERICA, N.A.
Reel/Frame 053468/0001 →