IP Library Granted Patent US 12,223,963
Granted Patent B2
US 12,223,963 · App. 16/900,857 · Granted Feb 11, 2025

Performing speech recognition using a local language context including a set of words with descriptions in terms of components smaller than the words

Inventors: Keyvan Mohajer (Los Gatos, CA); Timothy Stonehocker (Sunnyvale, CA); Bernard Mont-Reynaud (Sunnyvale, CA)
Assignee: ScoutHound AI IP, LLC
G10L15/30G10L15/04G10L15/063G10L15/08G10L15/26G10L17/06G10L2015/0635G10L2015/081G10L15/34
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,223,963
App. No.
16/900,857
Granted
Feb 11, 2025
Kind
B2
Abstract

A method of a local recognition system controlling a host device to perform one or more operations is provided. The method includes receiving, by the local recognition system, a query, performing speech recognition on the received query by implementing, by the local recognition system, a local language context comprising a set of words comprising descriptions in terms of components smaller than the words, and performing speech recognition, using the local language context, to create a transcribed query. Further, the method includes controlling the host device in dependence upon the speech recognition performed on the transcribed query.

Claims (55)

1. A method of operating a local recognition system hosted by a host device, and controlling the host device to perform one or more operations, the method comprising:

receiving, by the local recognition system, a query;

routing the query using a control module within the local recognition system to a local speech recognition module;

performing speech recognition on the received query by:

accessing, by the local speech recognition module, a local language context comprising:

(i) a first module comprising a vocabulary that includes a set of words, and

(ii) a second module comprising a language model;

implementing, by the local recognition system, an update module external to the local language context, the update module performing adaptation of the first module of the local language context by adding words to, or removing words from, the vocabulary; and

performing speech recognition, by the local speech recognition module using the local language context, to produce a transcribed output that matches content of the received query; and

in dependence on the transcribed output, issuing, by the control module, a command to control the host device to perform the one or more operations.

2. The method of claim 1 , wherein the local language context includes phonetic strings.

3. The method of claim 2 , wherein the local language context includes one or more of the phonetic strings per pronunciation of a word.

4. The method of claim 1 , wherein the local language context includes phonetic lattices.

5. The method of claim 4 , wherein the local language context includes one phonetic lattice per word of the set of words.

6. The method of claim 1 , wherein the vocabulary comprises words and phrases, and the language model comprises a set of constraints on word sequences expressed as N-grams and grammars.

7. A non-transitory computer-readable recording medium having a program recorded thereon, the program, when executed by a processor of a local recognition system hosted by a host device, causes the processor to perform a method comprising:

receiving, by the local recognition system, a query;

routing the query using a control module within the local recognition system to a local speech recognition module;

performing speech recognition on the received query by:

accessing, by the local speech recognition module, a local language context comprising:

(i) a first module comprising a vocabulary including a set of words, and

(ii) a second module comprising a language model;

implementing, by the local recognition system, an update module external to the local language context, the update module performing adaptation of the first module by adding words to, or removing words from, the vocabulary; and

performing speech recognition, by the local speech recognition module using the local language context, to produce a transcribed output that matches content of the received query; and

in dependence on the transcribed output, issuing, by the control module, a command to control the host device to perform one or more operations.

8. The non-transitory computer-readable recording medium of claim 7 , wherein the local language context includes phonetic strings.

9. The non-transitory computer-readable recording medium of claim 8 , wherein the local language context includes one or more of the phonetic strings per pronunciation of a word.

10. The non-transitory computer-readable recording medium of claim 7 , wherein the local language context includes phonetic lattices.

11. The non-transitory computer-readable recording medium of claim 10 , wherein the local language context includes one phonetic lattice per word of the set of words.

12. The non-transitory computer-readable recording medium of claim 7 , wherein the method further comprises

updating speech recognition functionality of the host device over a wireless network.

13. The non-transitory computer-readable recording medium of claim 7 , wherein the method further comprises

performing garbage collection by the update module when available memory resources are about to run out.

14. A local recognition system hosted by a host device and including one or more processors coupled to memory, the memory being loaded with computer instructions to control the host device to perform one or more operations, the computer instructions, when executed on the one or more processors, causing the one or more processors to implement operations comprising:

receiving a query;

routing the query using a control module within the local recognition system to a local speech recognition module;

performing speech recognition on the received query by:

accessing, by the local speech recognition module from a local database, a local language context comprising:

(i) a first module comprising a vocabulary including a set of words, and

(ii) a second module comprising a language model;

implementing, by the local recognition system, an update module external to the local language context, the update module performing adaptation of the first module by adding words to, or removing words from, the vocabulary; and

performing speech recognition, by the local speech recognition module using the local language context, to produce a transcribed output that matches content of the received query; and

in dependence on the transcribed output, issuing, by the control module, a command to control the host device to perform one or more operations.

15. The local recognition system of claim 14 , wherein

the local language context includes phonetic strings.

16. The local recognition system of claim 15 , wherein

the local language context includes one or more of the phonetic strings per pronunciation of a word.

17. The local recognition system of claim 14 , wherein

the local language context includes phonetic lattices.

18. The local recognition system of claim 17 , wherein

the local language context includes one phonetic lattice per word of the set of words.

19. The method of claim 1 , further comprising

updating speech recognition functionality of the host device over a wireless network.

20. The method of claim 1 , further comprising

performing garbage collection by the update module when available memory resources are about to run out.

Assignments (12)
TERMINATION AND RELEASE OF SECURITY INTEREST IN PATENTS Recorded Dec 3, 2024
From: MONROE CAPITAL MANAGEMENT ADVISORS, LLC, AS COLLATERAL AGENT
To: SOUNDHOUND, INC.
Reel/Frame 069480/0312 →
SECURITY INTEREST Recorded Aug 9, 2024
From: SOUNDHOUND, INC.
To: MONROE CAPITAL MANAGEMENT ADVISORS, LLC, AS COLLATERAL AGENT
Reel/Frame 068526/0413 →
RELEASE OF SECURITY INTEREST Recorded Jun 11, 2024
From: ACP POST OAK CREDIT II LLC, AS COLLATERAL AGENT
To: SOUNDHOUND, INC.; SOUNDHOUND AI IP, LLC
Reel/Frame 067698/0845 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jun 27, 2023
From: SOUNDHOUND AI IP HOLDING, LLC
To: SOUNDHOUND AI IP, LLC
Reel/Frame 064205/0676 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jun 23, 2023
From: SOUNDHOUND, INC.
To: SOUNDHOUND AI IP HOLDING, LLC
Reel/Frame 064083/0484 →
RELEASE OF SECURITY INTEREST Recorded Apr 21, 2023
From: FIRST-CITIZENS BANK & TRUST COMPANY, AS AGENT
To: SOUNDHOUND, INC.
Reel/Frame 063411/0396 →
RELEASE OF SECURITY INTEREST Recorded Apr 19, 2023
From: OCEAN II PLO LLC, AS ADMINISTRATIVE AGENT AND COLLATERAL AGENT
To: SOUNDHOUND, INC.
Reel/Frame 063380/0625 →
SECURITY INTEREST Recorded Apr 17, 2023
From: SOUNDHOUND, INC.; SOUNDHOUND AI IP, LLC
To: ACP POST OAK CREDIT II LLC
Reel/Frame 063349/0355 →
CORRECTIVE ASSIGNMENT TO CORRECT THE COVER SHEET PREVIOUSLY RECORDED AT REEL: 056627 FRAME: 0772. ASSIGNOR(S) HEREBY CONFIRMS THE SECURITY INTEREST. Recorded Apr 12, 2023
From: SOUNDHOUND, INC.
To: OCEAN II PLO LLC, AS ADMINISTRATIVE AGENT AND COLLATERAL AGENT
Reel/Frame 063336/0146 →
SECURITY INTEREST Recorded Jun 18, 2021
From: OCEAN II PLO LLC, AS ADMINISTRATIVE AGENT AND COLLATERAL AGENT
To: SOUNDHOUND, INC.
Reel/Frame 056627/0772 →
SECURITY INTEREST Recorded Apr 1, 2021
From: SOUNDHOUND, INC.
To: SILICON VALLEY BANK
Reel/Frame 055807/0539 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Sep 11, 2020
From: STONEHOCKER, TIMOTHY; MOHAJER, KEYVAN; MONT-REYNAUD, BERNARD
To: SOUNDHOUND, INC.
Reel/Frame 053747/0146 →
Continuity (6)
Continuation 15603257 · May 23, 2017
Continuation 15085944 · Mar 30, 2016
Continuation 14621024 · Feb 12, 2015
Continuation 13530101 · Jun 21, 2012
Provisional Application 61561393 · Nov 18, 2011
Related Publication 20200312329A1 · Oct 1, 2020
References Cited (83)
US 5956683A · Jacobs et al. · 1999 [cited by applicant]
US 6092045A · Stubley · 2000 [cited by examiner]
US 6173266B1 · Marx · 2001 [cited by examiner]
US 6327568B1 · Joost · 2001 [cited by applicant]
US 6374222B1 · Kao · 2002 [cited by examiner]
US 6377913B1 · Coffman et al. · 2002 [cited by applicant]
US 6408272B1 · White et al. · 2002 [cited by applicant]
US 6456975B1 · Chang · 2002 [cited by applicant]
US 6487534B1 · Thelen et al. · 2002 [cited by applicant]
US 6697782B1 · Iso-Sipila et al. · 2004 [cited by applicant]
US 6701294B1 · Ball et al. · 2004 [cited by applicant]
US 6704708B1 · Pickering · 2004 [cited by applicant]
US 7027987B1 · Franz · 2006 [cited by examiner]
US 7058573B1 · Murveit et al. · 2006 [cited by applicant]
US 7277854B2 · Bennett et al. · 2007 [cited by applicant]
US 7286989B1 · Niedermair · 2007 [cited by examiner]
US 7472060B1 · Gorin et al. · 2008 [cited by applicant]
US 7899669B2 · Gadbois · 2011 [cited by applicant]
US 8521526B1 · Lloyd et al. · 2013 [cited by applicant]
US 8949130B2 · Phillips · 2015 [cited by applicant]
US 8972263B2 · Stonehocker et al. · 2015 [cited by applicant]
US 9330669B2 · Stonehocker et al. · 2016 [cited by applicant]
US 9678928B1 · Tung · 2017 [cited by applicant]
US 9691390B2 · Stonehocker et al. · 2017 [cited by applicant]
US 20020198706A1 · Kao et al. · 2002 [cited by applicant]
US 20030125869A1 · Adams, Jr. · 2003 [cited by examiner]
US 20040153306A1 · Tanner · 2004 [cited by examiner]
US 20040210437A1 · Baker · 2004 [cited by applicant]
US 20050010422A1 · Ikeda et al. · 2005 [cited by applicant]
US 20050149326A1 · Hogengout · 2005 [cited by examiner]
US 20060009980A1 · Burke et al. · 2006 [cited by applicant]
US 20060036438A1 · Chang · 2006 [cited by applicant]
US 20060190256A1 · Stephanick et al. · 2006 [cited by applicant]
US 20060190268A1 · Wang · 2006 [cited by applicant]
US 20070011010A1 · Dow et al. · 2007 [cited by applicant]
US 20070038450A1 · Josifovski · 2007 [cited by examiner]
US 20070083374A1 · Bates · 2007 [cited by examiner]
US 20070233487A1 · Cohen · 2007 [cited by examiner]
US 20070276651A1 · Bliss et al. · 2007 [cited by applicant]
US 20080095327A1 · Wlasiuk · 2008 [cited by examiner]
US 20080177534A1 · Wang et al. · 2008 [cited by applicant]
US 20100057451A1 · Carraux et al. · 2010 [cited by applicant]
US 20100106497A1 · Phillips · 2010 [cited by applicant]
US 20110015928A1 · Odell et al. · 2011 [cited by applicant]
US 20110223893A1 · Lau · 2011 [cited by examiner]
US 20120022874A1 · Lloyd et al. · 2012 [cited by applicant]
US 20120089394A1 · Teodosiu et al. · 2012 [cited by applicant]
US 20120150539A1 · Jeon et al. · 2012 [cited by applicant]
US 20120179457A1 · Newman et al. · 2012 [cited by applicant]
US 20130085753A1 · Bringert et al. · 2013 [cited by applicant]
US 20130132084A1 · Stonehocker et al. · 2013 [cited by applicant]
US 20140163977A1 · Hoffmeister et al. · 2014 [cited by applicant]
US 20140250378A1 · Stifelman et al. · 2014 [cited by applicant]
US 20140372122A1 · Harsham et al. · 2014 [cited by applicant]
US 20160217788A1 · Stonehocker et al. · 2016 [cited by applicant]
US 20170069308A1 · Aleksic et al. · 2017 [cited by applicant]
US 20170178623A1 · Shamir et al. · 2017 [cited by applicant]
EP 2930716A1 · 2015 [cited by applicant]
WO 2016209444A1 · 2016 [cited by applicant]
U.S. Appl. No. 14/621,024—Office Action dated Aug. 26, 2015, 10 pages. [cited by applicant]
U.S. Appl. No. 14/621,024—Response to Aug. 26 Office Action filed Nov. 13, 2016, 5 pages. [cited by applicant]
U.S. Appl. No. 14/621,024—Notice of Allowance dated Jan. 5, 2016, 9 pages. [cited by applicant]
U.S. Appl. No. 13/530,101—Office Action dated Mar. 26, 2014, 8 pages. [cited by applicant]
U.S. Appl. No. 13/530,101—Response to Mar. 26 Office Action filed Sep. 23, 2014, 11 pages. [cited by applicant]
U.S. Appl. No. 13/530,101—Notice of Allowance dated Oct. 24, 2014, 8 pages. [cited by applicant]
U.S. Appl. No. 15/085,944—Office Action dated Nov. 16, 2016, 11 pages. [cited by applicant]
U.S. Appl. No. 15/085,944—Notice of Allowance mailed Feb. 24, 2016, 10 pages. [cited by applicant]
Javier Gonzalez-Dominguez, et al., A Real-Time End-to-End Multilingual Speech Recognition Architecture, IEEE Journal of Selected Topics in Signal Processing, Jun. 2015, vol. 9, No. 4, IEEE. [cited by applicant]
Takuma Okamoto, et al., Reducing latency for language identification based on large-vocabulary continuous speech recognition, Acoust. Sci. & Tech., 2017, pp. 38-41, vol. 38, Issue 1, The Acoustical Society of Japan. [cited by applicant]
U.S. Appl. No. 15/619,304—Office Action mailed Apr. 16, 2018, 30 pages. [cited by applicant]
EP18177044.7—Extended Euorpean Search Report mailed Aug. 18, 2018, 7 pages. [cited by applicant]
U.S. Appl. No. 15/619,304—Notice of Allowance mailed Apr. 26, 2019, 18 pages. [cited by applicant]
U.S. Appl. No. 15/619,304—Response to Office Action mailed Apr. 16, 2018 filed Jul. 17, 2018, 221 pages. [cited by applicant]
U.S. Appl. No. 15/619,304—Final Office Action mailed Oct. 26, 2018, 22 pages. [cited by applicant]
U.S. Appl. No. 15/619,304—Response to Final Office Action mailed Oct. 26, 2018 filed Dec. 13, 2018, 17 pages. [cited by applicant]
U.S. Appl. No. 15/619,304—Office Action mailed Feb. 5, 2019, 18 pages. [cited by applicant]
U.S. Appl. No. 15/619,304—Response to Office Action mailed Feb. 5, 2019 filed Apr. 11, 2019, 13 pages. [cited by applicant]
EP18177044.7—Extended European Search Report mailed Aug. 18, 2018, 7 pages. [cited by applicant]
U.S. Appl. No. 13/530,101, filed Jun. 21, 2012, U.S. Pat. No. 8,972,263, Mar. 3, 2015, Issued. [cited by applicant]
U.S. Appl. No. 14/621,024, filed Feb. 12, 2015, U.S. Pat. No. 9,330,669, May 3, 2016, Issued. [cited by applicant]
U.S. Appl. No. 15/085,944, filed Mar. 30, 2016, U.S. Pat. No. 9,691,390, Jun. 27, 2017, Issued. [cited by applicant]
U.S. Appl. No. 15/603,257, filed May 23, 2017, Abandoned. [cited by applicant]
U.S. Appl. No. 15/619,304, filed Jun. 9, 2017, U.S. Pat. No. 10,410,635, Sep. 10, 2019, Issued. [cited by applicant]