IP Library Granted Patent US 9,691,390
Granted Patent B2
US 9,691,390 · App. 15/085,944 · Granted Jun 27, 2017

System and method for performing dual mode speech recognition

View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 9,691,390
App. No.
15/085,944
Granted
Jun 27, 2017
Kind
B2
Abstract

A system and method is presented for performing dual mode speech recognition, employing a local recognition module on a mobile device and a remote recognition engine on a server device. The system accepts a spoken query from a user, and both the local recognition module and the remote recognition engine perform speech recognition operations on the query, returning a transcription and confidence score, subject to a latency cutoff time. If both sources successfully transcribe the query, then the system accepts the result having the higher confidence score. If only one source succeeds, then that result is accepted. In either case, if the remote recognition engine does succeed in transcribing the query, then a client vocabulary is updated if the remote system result includes information not present in the client vocabulary.

Claims (47)

1. A method for performing dual mode speech recognition, comprising:

receiving at a device a query from a user;

sending the query to a first recognition system;

sending the query to a second recognition system;

receiving at least a first recognition result from either the first recognition system or the second recognition system;

producing a final result considering the first recognition result; and

setting a latency timer to a timeout value,

wherein the first recognition system maintains a first vocabulary and the second recognition system maintains a second vocabulary, and whereby the final result is produced at or before the time that the latency timer reaches the timeout value.

2. The method of claim 1 wherein the final result is produced at the time of receiving the first recognition result.

3. The method of claim 1 wherein the final result is produced at the time of receiving a second recognition result, the final result selected from either the first recognition result or the second recognition result.

4. The method of claim 1 further comprising:

receiving a first recognition score associated with the first recognition result; and

producing the final result based on the first recognition score.

5. The method of claim 4 further comprising:

receiving a second recognition result;

receiving a second recognition score associated with the second recognition result; and

basing the producing of the final result on the greater of the first recognition score and the second recognition score.

6. The method of claim 1 , wherein:

the first recognition system is a local recognition system local to the device that receives the query from the user;

the second recognition system is a remote recognition system; and

the first recognition system sends the query to the remote recognition system over a communications link.

7. The method of claim 1 , further comprising:

determining that the second vocabulary contains at least one word that is not contained in the first vocabulary.

8. The method of claim 1 , further comprising:

the first recognition system receiving vocabulary information from the second recognition system; and

updating the first vocabulary with the received vocabulary information.

9. The method of claim 1 , wherein one or more words from the first vocabulary are assigned at least one of:

a frequency value that indicates how often the word is used; and

a recency value that indicates when the word was last used.

10. The method of claim 9 , further comprising removing a word from the first vocabulary based at least on the frequency value or the recency value.

11. A client for dual mode speech recognition, the client comprising:

an interface enabled to receive a query from a user;

a communication module enabled to send the query to a server and receive a remote recognition result from a server;

a local recognition module enabled to create a local recognition result from the query;

a latency timer;

a control module enabled to receive a notification from the latency timer and to select between the local recognition result and the remote recognition result; and

a client vocabulary enabled to describe words or phrases available to the local recognition module.

12. The client of claim 11 further comprising a vocabulary update module enabled to update the client vocabulary.

13. The client of claim 12 , wherein one or more words from the client vocabulary are assigned at least one of a frequency value and a recency value, and the client vocabulary update module removes the one or more words from the client vocabulary based on the at least one of a frequency value and a recency value.

14. The client of claim 11 , wherein the control module is enabled to:

receive a recognition score from the server;

receive a recognition score from the local recognition module; and

choose a recognition result based on the recognition score from the server and the recognition score from the local recognition module.

15. A server for dual mode speech recognition, the server comprising:

a recognition engine enabled to create a recognition result from audio content;

a communication module enabled to receive a query from a client and send the recognition result to the client; and

a vocabulary download module enabled to respond to requests from the client to send updates to a client vocabulary.

Assignments (12)
TERMINATION AND RELEASE OF SECURITY INTEREST IN PATENTS Recorded Dec 3, 2024
From: MONROE CAPITAL MANAGEMENT ADVISORS, LLC, AS COLLATERAL AGENT
To: SOUNDHOUND, INC.
Reel/Frame 069480/0312 →
SECURITY INTEREST Recorded Aug 9, 2024
From: SOUNDHOUND, INC.
To: MONROE CAPITAL MANAGEMENT ADVISORS, LLC, AS COLLATERAL AGENT
Reel/Frame 068526/0413 →
RELEASE OF SECURITY INTEREST Recorded Jun 11, 2024
From: ACP POST OAK CREDIT II LLC, AS COLLATERAL AGENT
To: SOUNDHOUND, INC.; SOUNDHOUND AI IP, LLC
Reel/Frame 067698/0845 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jun 27, 2023
From: SOUNDHOUND AI IP HOLDING, LLC
To: SOUNDHOUND AI IP, LLC
Reel/Frame 064205/0676 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jun 23, 2023
From: SOUNDHOUND, INC.
To: SOUNDHOUND AI IP HOLDING, LLC
Reel/Frame 064083/0484 →
RELEASE OF SECURITY INTEREST Recorded Apr 21, 2023
From: FIRST-CITIZENS BANK & TRUST COMPANY, AS AGENT
To: SOUNDHOUND, INC.
Reel/Frame 063411/0396 →
RELEASE OF SECURITY INTEREST Recorded Apr 19, 2023
From: OCEAN II PLO LLC, AS ADMINISTRATIVE AGENT AND COLLATERAL AGENT
To: SOUNDHOUND, INC.
Reel/Frame 063380/0625 →
SECURITY INTEREST Recorded Apr 17, 2023
From: SOUNDHOUND, INC.; SOUNDHOUND AI IP, LLC
To: ACP POST OAK CREDIT II LLC
Reel/Frame 063349/0355 →
CORRECTIVE ASSIGNMENT TO CORRECT THE COVER SHEET PREVIOUSLY RECORDED AT REEL: 056627 FRAME: 0772. ASSIGNOR(S) HEREBY CONFIRMS THE SECURITY INTEREST. Recorded Apr 12, 2023
From: SOUNDHOUND, INC.
To: OCEAN II PLO LLC, AS ADMINISTRATIVE AGENT AND COLLATERAL AGENT
Reel/Frame 063336/0146 →
SECURITY INTEREST Recorded Jun 18, 2021
From: OCEAN II PLO LLC, AS ADMINISTRATIVE AGENT AND COLLATERAL AGENT
To: SOUNDHOUND, INC.
Reel/Frame 056627/0772 →
SECURITY INTEREST Recorded Apr 1, 2021
From: SOUNDHOUND, INC.
To: SILICON VALLEY BANK
Reel/Frame 055807/0539 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jun 6, 2016
From: STONEHOCKER, TIMOTHY; MOHAJER, KEYVAN; MONT-REYNAUD, BERNARD
To: SOUNDHOUND, INC.
Reel/Frame 038821/0977 →