IP Library Granted Patent US 10,410,635
Granted Patent B2
US 10,410,635 · App. 15/619,304 · Granted Sep 10, 2019

Dual mode speech recognition

Inventor: Bernard Mont-Reynaud (Sunnyvale, CA)
Assignee: SoundHound, Inc.
G10L15/32G10L15/02G10L15/063G10L15/1822G10L15/30G10L2015/0635
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 10,410,635
App. No.
15/619,304
Granted
Sep 10, 2019
Kind
B2
Abstract

A dual mode speech recognition system sends speech to two or more speech recognizers. If a first recognition result is received, whose recognition score exceeds a high threshold, the first result is selected without waiting for another result. If the score is below a low threshold, the first result is ignored. At intermediate values of recognition scores, a timeout duration is dynamically determined as a function of the recognition score. The timeout duration determines how long the system will wait for another result. Many functions of the recognition score are possible, but timeout durations generally decrease as scores increase. When receiving a second recognition score before the timeout occurs, a comparison based on recognition scores determines whether the first result or the second result is the basis for creating a response.

Claims (49)

1. A speech recognition method comprising:

sending speech to a first recognizer and a second recognizer;

receiving, from the first recognizer, a first result associated with a recognition score;

setting a value of a timeout duration as a function of a value of the recognition score, such that the value of the timeout duration is set, in dependence upon the value of the recognition score, from at least one of a maximum value, an intermediary value and a minimum value;

responsive to receiving no result from the second recognizer before the timeout duration expires, choosing the first result as a basis for creating a response; and

responsive to receiving a second result from the second recognizer, updating a speech recognition vocabulary of the first recognizer to include at least one of an updated vocabulary model, an updated language model and an updated acoustic model.

2. The method of claim 1 , wherein the first recognizer and the second recognizer are local.

3. The method of claim 1 , wherein the first recognizer and the second recognizer are remote.

4. The method of claim 1 , wherein the speech is a continuous audio stream.

5. The method of claim 1 , wherein the speech is a delimited spoken query.

6. The method of claim 1 , wherein the recognition score is based on a phonetic sequence score.

7. The method of claim 1 , wherein the recognition score is based on a transcription score.

8. The method of claim 1 , wherein the recognition score is based on a grammar parse score.

9. The method of claim 1 , wherein the recognition score is based on an interpretation score.

10. A non-transitory computer readable medium storing code that, when executed by one or more computer processors, causes the one or more computer processors to:

send speech to a first recognizer and a second recognizer;

receive, from the first recognizer, a first result associated with a recognition score;

set a value of a timeout duration as a function of the value of the recognition score, such that the value of the timeout duration is set, in dependence upon the value of the recognition score, from at least one of a maximum value, an intermediary value and a minimum value;

responsive to receiving no result from the second recognizer before the timeout duration expires, choose the first result as a basis for creating a response; and

responsive to receiving a second result from the second recognizer, updating a speech recognition vocabulary of the first recognizer to include at least one of an updated vocabulary model, an updated language model and an updated acoustic model.

11. A mobile device enabled to perform dual mode speech recognition, the device comprising:

a module for receiving speech from a user;

a module for sending speech to a first recognizer;

a module for sending speech to a second recognizer;

a module for receiving a recognition score corresponding to recognition by the first recognizer;

a module for detecting a timeout based on a timeout duration, a value of the timeout duration being selected as a function of the value of the recognition score, such that the value of the timeout duration is set in dependence upon the value of the recognition score, from at least one of a maximum value, an intermediary value and a minimum value; and

responsive to receiving a second result from the second recognizer, updating a speech recognition vocabulary of the first recognizer to include at least one of an updated vocabulary model, an updated language model and an updated acoustic model,

wherein the mobile device chooses a first result from the first recognizer if it does not receive a result from the second recognizer before the timeout occurs.

12. The mobile device of claim 11 wherein the first recognizer is local to the device and the second recognizer is remote from the mobile device.

13. The method of claim 1 , wherein the function is selected from a set of functions consisting of a linear function, a parabolic function and an s-shaped function.

14. The method of claim 1 , wherein the function is not a step function.

15. A speech recognition method comprising:

continuously sending speech to both (i) a first recognizer for first recognition of the speech and (ii) a second recognizer for second recognition of the speech;

receiving, from the first recognizer, a first speech recognition result and an associated first recognition score;

responsive to the first recognition score being above a threshold, choosing the first speech recognition result from the first recognizer as a basis for creating a response to the speech and responsive to the first recognition score being below the threshold, waiting a predetermined period to receive a second speech recognition result and an associated second recognition score from the second recognizer;

receiving, from the second recognizer, the second speech recognition result and the second recognition score; and

responsive to receiving the second speech recognition result and the second recognition score, choosing one of the first speech recognition result and the second speech recognition result, in dependence upon the first recognition score and the second recognition score,

wherein the first recognition of the speech by the first recognizer and the second recognition of the speech by the second recognizer are continuous, and

wherein the first recognition score and the second recognition score are recomputed and adjusted on a continuing basis by the first recognizer and the second recognizer as the speech continues to be sent to both the first recognizer and the second recognizer, such that the first recognition score and the second recognition score are continuously updated as new and continuous speech is recognized and until an end of the first recognition of the speech and the second recognition of the speech.

16. The method of claim 15 , further comprising:

responsive to the first recognition score being below a low threshold, ignoring the first speech recognition result; and

responsive to not receiving a second response before a timeout occurs, signaling an error.

17. The method of claim 15 , wherein the first recognizer and the second recognizer are local.

18. The method of claim 15 , wherein the first recognizer and the second recognizer are remote.

19. The method of claim 15 , wherein the speech is a delimited spoken query.

20. The method of claim 15 , wherein the first recognition score is based on a phonetic sequence score.

21. The method of claim 15 , wherein the first recognition score is based on a transcription score.

22. The method of claim 15 , wherein the first recognition score is based on a grammar parse score.

23. The method of claim 15 , wherein the first recognition score is based on an interpretation score.

Assignments (12)
TERMINATION AND RELEASE OF SECURITY INTEREST IN PATENTS Recorded Dec 3, 2024
From: MONROE CAPITAL MANAGEMENT ADVISORS, LLC, AS COLLATERAL AGENT
To: SOUNDHOUND, INC.
Reel/Frame 069480/0312 →
SECURITY INTEREST Recorded Aug 9, 2024
From: SOUNDHOUND, INC.
To: MONROE CAPITAL MANAGEMENT ADVISORS, LLC, AS COLLATERAL AGENT
Reel/Frame 068526/0413 →
RELEASE OF SECURITY INTEREST Recorded Jun 11, 2024
From: ACP POST OAK CREDIT II LLC, AS COLLATERAL AGENT
To: SOUNDHOUND, INC.; SOUNDHOUND AI IP, LLC
Reel/Frame 067698/0845 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jun 27, 2023
From: SOUNDHOUND AI IP HOLDING, LLC
To: SOUNDHOUND AI IP, LLC
Reel/Frame 064205/0676 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jun 23, 2023
From: SOUNDHOUND, INC.
To: SOUNDHOUND AI IP HOLDING, LLC
Reel/Frame 064083/0484 →
RELEASE OF SECURITY INTEREST Recorded Apr 21, 2023
From: FIRST-CITIZENS BANK & TRUST COMPANY, AS AGENT
To: SOUNDHOUND, INC.
Reel/Frame 063411/0396 →
RELEASE OF SECURITY INTEREST Recorded Apr 19, 2023
From: OCEAN II PLO LLC, AS ADMINISTRATIVE AGENT AND COLLATERAL AGENT
To: SOUNDHOUND, INC.
Reel/Frame 063380/0625 →
SECURITY INTEREST Recorded Apr 17, 2023
From: SOUNDHOUND, INC.; SOUNDHOUND AI IP, LLC
To: ACP POST OAK CREDIT II LLC
Reel/Frame 063349/0355 →
CORRECTIVE ASSIGNMENT TO CORRECT THE COVER SHEET PREVIOUSLY RECORDED AT REEL: 056627 FRAME: 0772. ASSIGNOR(S) HEREBY CONFIRMS THE SECURITY INTEREST. Recorded Apr 12, 2023
From: SOUNDHOUND, INC.
To: OCEAN II PLO LLC, AS ADMINISTRATIVE AGENT AND COLLATERAL AGENT
Reel/Frame 063336/0146 →
SECURITY INTEREST Recorded Jun 18, 2021
From: OCEAN II PLO LLC, AS ADMINISTRATIVE AGENT AND COLLATERAL AGENT
To: SOUNDHOUND, INC.
Reel/Frame 056627/0772 →
CORRECTIVE ASSIGNMENT TO CORRECT THE ASSIGNOR NAME PREVIOUSLY RECORDED AT REEL: 042760 FRAME: 0191. ASSIGNOR(S) HEREBY CONFIRMS THE ASSIGNMENT. Recorded Jun 29, 2017
From: MONT-REYNAUD, BERNARD
To: SOUNDHOUND, INC.
Reel/Frame 043060/0082 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jun 20, 2017
From: MONT-REYNAUD, BERNARD; MOHAJER, KAMYAR
To: SOUNDHOUND, INC.
Reel/Frame 042760/0191 →
Continuity (1)
Related Publication 20180358019A1 · Dec 13, 2018
Cited By (1)
US 12,603,088