IP Library Granted Patent US 12,125,471
Granted Patent B2
US 12,125,471 · App. 18/211,732 · Granted Oct 22, 2024

Automated word correction in speech recognition systems

Inventors: Ankur Anil Aher (Maharashtra, IN); Jeffry Copps Robert Jose (Tamil Nadu, IN)
Assignee: Rovi Guides, Inc.
G10L15/01G10L15/063G10L15/08G10L15/22G10L2015/0636G10L2015/086
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,125,471
App. No.
18/211,732
Granted
Oct 22, 2024
Kind
B2
Abstract

Systems and methods for correcting recognition errors in speech recognition systems are disclosed herein. Natural conversational variations are identified to determine whether a query intends to correct a speech recognition error or whether the query is a new command. When the query intends to correct a speech recognition error, the system identifies a location of the error and performs the correction. The corrected query can be presented to the user or be acted upon as a command for the system.

Claims (52)

1. A method for correcting a speech recognition error, the method comprising:

detecting a first voice input;

detecting a second voice input within a threshold time of detecting the first voice input; and

in response to detecting the second voice input within the threshold time of detecting the first voice input, determining whether the second voice input is directed to correcting the first voice input based on:

comparing a first plurality of sound properties of the first voice input to a second plurality of sound properties of the second voice input; and

determining, based on the comparing, that the second voice input is directed to correcting the first voice input when at least a threshold amount of the first plurality of sound properties match the second plurality of sound properties and at least one sound property of the first plurality of sound properties does not match the second plurality of sound properties.

2. The method of claim 1 , further comprising, determining, based on the comparing, that the second voice input is not directed to correcting the first voice input when at least the threshold amount of the first plurality of sound properties do not match the second plurality of sound properties.

3. The method of claim 1 , wherein the first plurality of sound properties comprises at least one of a first acoustic envelope, a first intensity, a first pitch, a first frequency, and a first amplitude, of the first voice input, and wherein the second plurality of sound properties comprises at least one of a second acoustic envelope, a second intensity, a second pitch, a second frequency, and a second amplitude, of the second voice input.

4. The method of claim 1 , wherein determining that the second voice input is directed to correcting the first voice input, further comprises:

identifying a first word in the first voice input;

identifying a second word in the second voice input; and

determining, that the first word and the second word are frequently misrecognized for one another.

5. The method of claim 1 , wherein in response to determining that the second voice input is directed to correcting the first voice input, further comprises modifying an interpretation of the first voice input based on the second voice input.

6. The method of claim 5 , further comprising

identifying a first interpretation and a second interpretation for a portion of the first voice input, wherein each interpretation is associated with a respective confidence value;

based on determining that the first interpretation is associated with a higher, respective confidence value than the second interpretation, selecting the first interpretation for the portion; and

wherein modifying an interpretation of the first voice input based on the second voice input comprises, replacing the first interpretation for the portion with the second interpretation for the portion.

7. The method of claim 5 , further comprising generating for output the modified interpretation.

8. The method of claim 5 , further comprising:

identifying an action corresponding to the modified interpretation; and

causing the action to be performed.

9. The method of claim 1 , further comprising:

detecting a third voice input subsequent to detecting the second voice input;

determining that a time of receipt of the third voice input is not within a threshold time of receipt of the second voice input; and

in response to determining that the time of receipt of the third voice input is not within the threshold time of receipt of the second voice input, determining that the third voice input is not directed to correcting the second voice input.

10. The method of claim 1 , wherein determining that the second voice input is directed to correcting the first voice input comprises detecting a corrective expression within the second voice input.

11. A system for correcting a speech recognition error, the system comprising control circuitry configured to:

detect a first voice input;

detect a second voice input within a threshold time of detecting the first voice input; and

in response to detecting the second voice input within the threshold time of detecting the first voice input, determine whether the second voice input is directed to correcting the first voice input based on:

comparing a first plurality of sound properties of the first voice input to a second plurality of sound properties of the second voice input; and

determining, based on the comparing, that the second voice input is directed to correcting the first voice input when at least a threshold amount of the first plurality of sound properties match the second plurality of sound properties and at least one sound property of the first plurality of sound properties does not match the second plurality of sound properties.

12. The system of claim 11 , wherein the control circuitry is further configured to, determine, based on the comparing, that the second voice input is not directed to correcting the first voice input when at least the threshold amount of the first plurality of sound properties do not match the second plurality of sound properties.

13. The system of claim 11 , wherein the first plurality of sound properties comprises at least one of a first acoustic envelope, a first intensity, a first pitch, a first frequency, and a first amplitude, of the first voice input, and wherein the second plurality of sound properties comprises at least one of a second acoustic envelope, a second intensity, a second pitch, a second frequency, and a second amplitude, of the second voice input.

14. The system of claim 11 , wherein the control circuitry is further configured, when determining that the second voice input is directed to correcting the first voice input, to:

identify a first word in the first voice input;

identify a second word in the second voice input; and

determine, that the first word and the second word are frequently misrecognized for one another.

15. The system of claim 11 , wherein the control circuitry is further configured, in response to determining that the second voice input is directed to correcting the first voice input, to modify an interpretation of the first voice input based on the second voice input.

16. The system of claim 15 , wherein the control circuitry is further configured to:

identify a first interpretation and a second interpretation for a portion of the first voice input, wherein each interpretation is associated with a respective confidence value;

based on determining that the first interpretation is associated with a higher, respective confidence value than the second interpretation, select the first interpretation for the portion; and

wherein modifying an interpretation of the first voice input based on the second voice input comprises, replace the first interpretation for the portion with the second interpretation for the portion.

17. The system of claim 15 , wherein the control circuitry is further configured to generate for output the modified interpretation.

18. The system of claim 15 , wherein the control circuitry is further configured to:

identify an action corresponding to the modified interpretation; and

cause the action to be performed.

19. The system of claim 11 , wherein the control circuitry is further configured to:

detect a third voice input subsequent to detecting the second voice input;

determine that a time of receipt of the third voice input is not within a threshold time of receipt of the second voice input; and

in response to determining that the time of receipt of the third voice input is not within the threshold time of receipt of the second voice input, determine that the third voice input is not directed to correcting the second voice input.

20. The system of claim 11 , wherein the control circuitry is further configured, when determining that the second voice input is directed to correcting the first voice input, to detect a corrective expression within the second voice input.

Assignments (2)
CHANGE OF NAME Recorded Sep 24, 2024
From: ROVI GUIDES, INC.
To: ADEIA GUIDES INC.
Reel/Frame 069032/0008 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jun 20, 2023
From: AHER, ANKUR ANIL; ROBERT JOSE, JEFFRY COPPS
To: ROVI GUIDES, INC.
Reel/Frame 063995/0673 →
Continuity (2)
Continuation 16804201 · Feb 28, 2020
Related Publication 20230410792A1 · Dec 21, 2023