IP Library › Granted Patent US 12,603,088
Granted Patent B2
US 12,603,088 · App. 18/379,618 · Granted Apr 14, 2026

Training a device specific acoustic model

Inventors: Keyvan Mohajer (Los Gatos, CA); Mehul Patel (Santa Clara, CA)
Assignee: SoundHound AI IP, LLC
G10L15/22G06F3/167G10L15/18
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,603,088
App. No.
18/379,618
Filed
Oct 12, 2023
Granted
Apr 14, 2026
Kind
B2
Art Unit
2657
USPC
704/257
Abstract

Custom acoustic models can be configured by developers by providing audio files with custom recordings. The custom acoustic model is trained by tuning a baseline model using the audio files. Audio files may contain custom noise to apply to clean speech for training. The custom acoustic model is provided as an alternative to a standard acoustic model. A speech recognition system can select an acoustic model for use upon receiving metadata about the device conditions or type. Speech recognition is performed on speech audio using one or more acoustic models. The result can be provided to developers through the user interface, and an error rate can be computed and also provided.

Claims (56)

1 . A method of performing speech recognition comprising:

storing a plurality of acoustic models associated with a device;

receiving metadata including information according to which one or more trained acoustic models, of the plurality of acoustic models that are already trained, is selected;

preselecting, based on the metadata, one or more trained acoustic models of the plurality of acoustic models;

receiving speech audio including natural language utterances;

selecting a trained acoustic model from the preselected one or more trained acoustic models, the selected trained acoustic model being trained with environmental features; and

employing the selected trained acoustic model to recognize speech from the natural language utterances included in the received speech audio.

2 . The method of claim 1 , wherein

the plurality of acoustic models are associated with different device conditions, and

the metadata is indicative of a device condition.

3 . The method of claim 2 , wherein

the device conditions include usage conditions of the device.

4 . The method of claim 3 , wherein

usage conditions of the device provide information regarding one or more hardware and software components of the device for receiving the speech audio or for providing audio feedback to a user of the device.

5 . The method of claim 1 , wherein

the received metadata is stored within the device.

6 . The method of claim 1 , wherein

the plurality of acoustic models are associated with different device types, and

the metadata is indicative of a device type.

7 . The method of claim 6 , wherein

the device type identifies at least one of a model number and serial number of the device.

8 . The method of claim 1 , wherein

employing the selected trained acoustic model to recognize speech comprises extraction of phonemes from the received speech audio.

9 . The method of claim 1 , wherein the device is an indoor appliance.

10 . The method of claim 1 , wherein the device is a mobile device.

11 . The method of claim 1 , wherein the plurality of acoustic models are stored, at least in part, in the cloud.

12 . A non-transitory computer readable medium storing code that, if executed by one or more computers, would cause the one or more computers to:

store a plurality of acoustic models associated with a device;

receive metadata including information according to which one or more trained acoustic models, of the plurality of acoustic models that are already trained, is selected;

preselect, based on the metadata, one or more trained acoustic models of the plurality of acoustic models;

receive speech audio including natural language utterances;

select a trained acoustic model from the preselected one or more trained acoustic models, the selected trained acoustic model being trained with environmental features; and

employ the selected trained acoustic model to recognize speech from the natural language utterances included in the received speech audio.

13 . A method of using a platform for configuring device-specific speech recognition, the method comprising:

receiving a selection of a set of at least two acoustic models appropriate for a specific type of a device, the selection of the set of at least two acoustic models being received by an interaction with a graphical user interface provided by a computer system; and

providing received speech audio and metadata to a speech recognition system associated with the platform.

14 . The method of claim 13 further comprising:

providing a custom acoustic model appropriate for the specific type of the device,

wherein the set of selected acoustic models includes the provided custom acoustic model.

15 . The method of claim 13 further comprising:

providing training data for training an acoustic model appropriate to the specific type of the device; and

selecting an acoustic model trained on the provided training data.

16 . The method of claim 13 , wherein

the metadata identifies an acoustic model of the set according to the specific type of the device.

17 . The method of claim 13 , wherein

the metadata identifies a specific device condition, and

a computer system selects an acoustic model from the set of acoustic models in dependence upon the specific device condition.

18 . The method of claim 13 , wherein

the at least two acoustic models recognize speech by extraction of phonemes from the received speech audio.

19 . The method of claim 13 , wherein the interaction is a physical interaction.

20 . A non-transitory computer readable medium storing code that, if executed by one or more computers, would cause the one or more computers to:

detect information useful for selecting an acoustic model from a plurality of acoustic models that are already trained and indicative of a device condition;

receive speech audio;

transmit the detected information and the received speech audio; and

receive information requested by speech in the speech audio,

wherein the detected information is capable of being employed to preselect and select a trained acoustic model from a plurality of acoustic models associated with different device conditions, the trained acoustic model being trained with environmental features, and the detected information being capable of being used to recognize speech from the received speech audio.

Assignments (6)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Mar 11, 2024
From: SOUNDHOUND, INC.
To: SOUNDHOUND AI IP HOLDING, LLC
Reel/Frame 066789/0796 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Mar 11, 2024
From: SOUNDHOUND AI IP HOLDING, LLC
To: SOUNDHOUND AI IP, LLC
Reel/Frame 066789/0852 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Mar 11, 2024
From: MOHAJER, KEYVAN; PATEL, MEHUL
To: SOUNDHOUND, INC.
Reel/Frame 066718/0458 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Mar 4, 2024
From: SOUNDHOUND, INC.
To: SOUNDHOUND AI IP HOLDING, LLC
Reel/Frame 066725/0221 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Mar 4, 2024
From: SOUNDHOUND AI IP HOLDING, LLC
To: SOUNDHOUND AI IP, LLC
Reel/Frame 066726/0826 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Mar 4, 2024
From: MOHAJER, KEYVAN; PATEL, MEHUL
To: SOUNDHOUND, INC.
Reel/Frame 066635/0602 →
Continuity (4)
Continuation 17573551 · Jan 11, 2022
Continuation 17237003 · Apr 21, 2021
Continuation 15996393 · Jun 1, 2018
Related Publication 20240038233A1 · Feb 1, 2024
References Cited (163)
US 5909666A · Gould · 1999 [cited by examiner]
US 6141641A · Hwang · 2000 [cited by examiner]
US 6263308B1 · Heckerman · 2001 [cited by examiner]
US 6442512B1 · Sengupta et al. · 2002 [cited by applicant]
US 6477493B1 · Brooks et al. · 2002 [cited by applicant]
US 6584439B1 · Geilhufe et al. · 2003 [cited by applicant]
US 6842734B2 · Yamada et al. · 2005 [cited by applicant]
US 7437294B1 · Thenthiruperai · 2008 [cited by applicant]
US 7720683B1 · Vermeulen et al. · 2010 [cited by applicant]
US 8489632B1 · Breckenridge · 2013 [cited by examiner]
US 8898063B1 · Sykes · 2014 [cited by examiner]
US 9159315B1 · Mengibar · 2015 [cited by examiner]
US 9208781B2 · Bell et al. · 2015 [cited by applicant]
US 9263040B2 · Tzirkel-Hancock et al. · 2016 [cited by applicant]
US 9443527B1 · Watanabe et al. · 2016 [cited by applicant]
US 9449283B1 · Purpura · 2016 [cited by examiner]
US 9460716B1 · Epstein et al. · 2016 [cited by applicant]
US 9691384B1 · Wang · 2017 [cited by examiner]
US 9786281B1 · Adams · 2017 [cited by examiner]
US 9881255B1 · Castellanos et al. · 2018 [cited by applicant]
US 10147442B1 · Panchapagesan · 2018 [cited by examiner]
US 10152968B1 · Agrusa · 2018 [cited by examiner]
US 10319250B2 · Lokeswarappa et al. · 2019 [cited by applicant]
US 10326657B1 · A et al. · 2019 [cited by applicant]
US 10410635B2 · Mont-Reynaud · 2019 [cited by applicant]
US 10424292B1 · Thimsen et al. · 2019 [cited by applicant]
US 10679621B1 · Sundaram · 2020 [cited by examiner]
US 20020049587A1 · Miyazawa · 2002 [cited by examiner]
US 20020055840A1 · Yamada · 2002 [cited by examiner]
US 20020055844A1 · L'Esperance · 2002 [cited by examiner]
US 20020169604A1 · Damiba et al. · 2002 [cited by applicant]
US 20030012347A1 · Steinbiss · 2003 [cited by examiner]
US 20030033144A1 · Silverman · 2003 [cited by examiner]
US 20030050783A1 · Yoshizawa · 2003 [cited by applicant]
US 20030120488A1 · Yoshizawa · 2003 [cited by examiner]
US 20030130840A1 · Forand · 2003 [cited by examiner]
US 20030191636A1 · Zhou · 2003 [cited by applicant]
US 20030236099A1 · Deisher · 2003 [cited by examiner]
US 20040138882A1 · Miyazawa · 2004 [cited by examiner]
US 20040158457A1 · Veprek · 2004 [cited by examiner]
US 20040230420A1 · Kadambe · 2004 [cited by examiner]
US 20050071159A1 · Boman · 2005 [cited by examiner]
US 20050075875A1 · Shozakai · 2005 [cited by examiner]
US 20050114128A1 · Hetherington et al. · 2005 [cited by applicant]
US 20050187763A1 · Arun · 2005 [cited by applicant]
US 20060053014A1 · Yoshizawa · 2006 [cited by examiner]
US 20060074651A1 · Arun · 2006 [cited by applicant]
US 20060178886A1 · Braho · 2006 [cited by examiner]
US 20060235687A1 · Carus et al. · 2006 [cited by applicant]
US 20080004875A1 · Chengalvarayan et al. · 2008 [cited by applicant]
US 20080167862A1 · Mohajer · 2008 [cited by examiner]
US 20080247577A1 · Dressler · 2008 [cited by examiner]
US 20090254753A1 · De Atley et al. · 2009 [cited by applicant]
US 20090313004A1 · Levi et al. · 2009 [cited by applicant]
US 20100049516A1 · Talwar · 2010 [cited by examiner]
US 20100145699A1 · Tian · 2010 [cited by examiner]
US 20100228548A1 · Liu et al. · 2010 [cited by applicant]
US 20100250240A1 · Shu · 2010 [cited by examiner]
US 20100268534A1 · Kishan Thambiratnam et al. · 2010 [cited by applicant]
US 20100312555A1 · Plumpe et al. · 2010 [cited by applicant]
US 20100312557A1 · Strom et al. · 2010 [cited by applicant]
US 20110066433A1 · Ljolje et al. · 2011 [cited by applicant]
US 20110295590A1 · Lloyd · 2011 [cited by examiner]
US 20110307253A1 · Lloyd · 2011 [cited by examiner]
US 20130013991A1 · Evans · 2013 [cited by applicant]
US 20130030802A1 · Jia · 2013 [cited by examiner]
US 20130185066A1 · Tzirkel-Hancock · 2013 [cited by examiner]
US 20140007222A1 · Qureshi et al. · 2014 [cited by applicant]
US 20140020061A1 · Popp et al. · 2014 [cited by applicant]
US 20140039888A1 · Taubman · 2014 [cited by examiner]
US 20140112556A1 · Kalinli-Akbacak · 2014 [cited by examiner]
US 20140142944A1 · Ziv et al. · 2014 [cited by applicant]
US 20140214414A1 · Poliak · 2014 [cited by applicant]
US 20140278415A1 · Ivanov · 2014 [cited by examiner]
US 20140365218A1 · Chang · 2014 [cited by examiner]
US 20140365221A1 · Ben-Ezra · 2014 [cited by applicant]
US 20140372118A1 · Yassa · 2014 [cited by examiner]
US 20150012268A1 · Nakadai et al. · 2015 [cited by applicant]
US 20150025890A1 · Jagatheesan · 2015 [cited by examiner]
US 20150058003A1 · Mohideen et al. · 2015 [cited by applicant]
US 20150081288A1 · Kim · 2015 [cited by examiner]
US 20150081300A1 · Kim · 2015 [cited by examiner]
US 20150100528A1 · Danson · 2015 [cited by examiner]
US 20150149167A1 · Beaufays · 2015 [cited by examiner]
US 20150149174A1 · Gollan et al. · 2015 [cited by applicant]
US 20150161999A1 · Kalluri et al. · 2015 [cited by applicant]
US 20150223001A1 · Choi · 2015 [cited by examiner]
US 20150286770A1 · Morishita · 2015 [cited by examiner]
US 20150301795A1 · Lebrun · 2015 [cited by applicant]
US 20150364139A1 · Dimitriadis et al. · 2015 [cited by applicant]
US 20150377667A1 · Ahmad · 2015 [cited by examiner]
US 20160019884A1 · Xiao · 2016 [cited by examiner]
US 20160234206A1 · Tunnell et al. · 2016 [cited by applicant]
US 20160253989A1 · Kuo · 2016 [cited by examiner]
US 20160358600A1 · Nallasamy · 2016 [cited by examiner]
US 20160372107A1 · Dow et al. · 2016 [cited by applicant]
US 20170053652A1 · Choi · 2017 [cited by examiner]
US 20170076725A1 · Kumar · 2017 [cited by examiner]
US 20170109368A1 · Mohajer · 2017 [cited by applicant]
US 20170116991A1 · Lin · 2017 [cited by examiner]
US 20170206903A1 · Kim · 2017 [cited by examiner]
US 20170213549A1 · Hassani · 2017 [cited by examiner]
US 20170213551A1 · Ji · 2017 [cited by examiner]
US 20170352346A1 · Paulik · 2017 [cited by examiner]
US 20180018959A1 · Des Jardins · 2018 [cited by examiner]
US 20180025721A1 · Li · 2018 [cited by examiner]
US 20180040318A1 · Gilbert · 2018 [cited by examiner]
US 20180061409A1 · Valentine · 2018 [cited by examiner]
US 20180213339A1 · Shah et al. · 2018 [cited by applicant]
US 20180286413A1 · Hassani et al. · 2018 [cited by applicant]
US 20180330737A1 · Paulik et al. · 2018 [cited by applicant]
US 20190051290A1 · Li et al. · 2019 [cited by applicant]
US 20190138940A1 · Feuz et al. · 2019 [cited by applicant]
US 20190185013A1 · Zhou et al. · 2019 [cited by applicant]
US 20190206389A1 · Kwon et al. · 2019 [cited by applicant]
US 20190287515A1 · Li et al. · 2019 [cited by applicant]
US 20190295539A1 · Mese et al. · 2019 [cited by applicant]
US 20210012769A1 · Vasconcelos et al. · 2021 [cited by applicant]
US 20210065712A1 · Holm · 2021 [cited by applicant]
US 20210118435A1 · Stahl · 2021 [cited by applicant]
US 20210256386A1 · Wieman et al. · 2021 [cited by applicant]
US 20210272552A1 · Lokeswarappa et al. · 2021 [cited by applicant]
US 20210312920A1 · Stahl · 2021 [cited by applicant]
US 20210335340A1 · Gowayyed et al. · 2021 [cited by applicant]
CN 101923854A · 2010 [cited by applicant]
CN 103038817A · 2013 [cited by applicant]
CN 103714812A · 2014 [cited by applicant]
CN 107958385A · 2018 [cited by applicant]
CN 113270091A · 2021 [cited by applicant]
EP 3783605A1 · 2021 [cited by applicant]
JP 2000353294A · 2000 [cited by applicant]
JP 2003177790A · 2003 [cited by applicant]
JP 2005181459A · 2005 [cited by applicant]
JP 2008158328A · 2008 [cited by applicant]
WO 2005010868A1 · 2005 [cited by applicant]
Microsoft.com, Microsoft Cognitive Services, Create custom acoustic models, 1 page (accessed Feb. 8, 2017, www.microsoft.com/cognitiveservices/en-us/customrecognitionintelligentservicecris). [cited by applicant]
Microsoft.com, Custom Speech Service, Creating a custom acoustic model, Jun. 27, 2017, 39 pages. [cited by applicant]
Microsoft.com, Design Windows 10 devices, May 2, 2107, 917 pages. [cited by applicant]
Amazon, SpeechRecognizer Interface, Profiles, 9 pages, (accessed Jan. 16, 2018, https://developer.amazon.com/docs/alexa-voice-service/speechrecognizer.html). [cited by applicant]
Amazon, Audio Hardware Configurations, Automatic Speech Recognition Profiles, 3 pages, (accessed Jan. 16, 2018, developer.amazon.com/docs/alexa-voice-service/audio-hardware-configurations.html#asr). [cited by applicant]
Freesound.org, About Freesound, 1 page, (accessed Feb. 8, 2017, freesound.org/help/about/). [cited by applicant]
Asma Rabaoui, Hidden Markov Model Environment Adaptation for Noisy Sounds in a Supervised Recognition System, 2nd International Symposium on Communication, Control and Signal Processing (ISCCSP). Mar. 13-15, 2006, 4 pag… [cited by applicant]
Goshu Nagino, Design of Ready-Made Acoustic Model Library by Two-Dimensional Visualization of Acoustic Space, Eighth International Conference on Spoken Language Processing. Oct. 4-8, 2004, 4 pages. [cited by applicant]
JP Application No. 2019-29710, Notice of Refusal dated Nov. 23, 2020, 16 pages (retrieved from Global Dossier). [cited by applicant]
JP Application No. 2019-29710, Response to Notice of Refusal dated Nov. 23, 2020, filed Jan. 3, 2021, 10 pages (retrieved from Global Dossier). [cited by applicant]
JP Application No. 2019-29710, Search Report dated Apr. 1, 2020, 46 pages (retrieved from Global Dossier). [cited by applicant]
JP Application No. 2019-29710, Notice of Grant dated Mar. 3, 2021, 5 pages (retrieved from Global Dossier). [cited by applicant]
Mirsamadi, S. et al., “On Multi-Domain Training and Adaptation of End-to-End RNN Acoustic Models for Distant Speech Recognition,” Interspeech 2017, Center for Robust Speech Systems, University of Texas, Aug. 20-24, 2017… [cited by applicant]
Mdhaffar et al., Retrieving Speaker Information from Personalized Acoustic Models for Speech Recognition, LIA, Avignon University, France, (arXiv preprint arXiv:2111.04194 ), dated Nov. 7, 2021, 5 pages. [cited by applicant]
McGraw, I. et al., “Personalized Speech Recognition on Mobile Devices,” 2016 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), IEEE Explore, Mar. 11, 2016, pp. 5955-5959. [cited by applicant]
Amazon, What devices does Amazon Transcribe work with, Amazon FAQs, 1 page. Retrieved on Jan. 6, 2022. Retrieved from the internet [URL: https://aws.amazon.com/transcribe/faqs/ ]. [cited by applicant]
Microsoft, What is Custom Speech?, dated Nov. 3, 2021, 2 pages. Retrieved on Nov. 4, 2021. Retrieved from the internet [URL: https://docs.microsoft.com/en-us/azure/cognitive-services/]. [cited by applicant]
Microsoft, Prepare data for Custom Speech, dated Nov. 3, 2021, 14 pages. Retrieved on Nov. 4, 2021. Retrieved from the internet [URL: https://docs.microsoft.com/en-US/azure/cognitive-services/]. [cited by applicant]
Microsoft, Inspect Custom Speech data, dated Nov. 3, 2021, 5 pages. Retrieved on Nov. 4, 2021. Retrieved from the internet [URL: https://docs.microsoft.com/en-us/azure/cognitive-services/]. [cited by applicant]
Microsoft, Evaluate and improve Custom Speech accuracy, dated Nov. 3, 2021, 9 pages. Retrieved on Nov. 4, 2021. Retrieved from the internet [URL: https://docs.microsoft.com/en-us/azure/cognitive-services/]. [cited by applicant]
Microsoft, Train and deploy a Custom Speech model, dated Nov. 3, 2021, 5 pages. Retrieved on Nov. 4, 2021. Retrieved from the internet [URL: https://docs.microsoft.com/en-us/azure/cognitive-services/]. [cited by applicant]
Siri Team, “Hey Siri: An On-device DNN-powered Voice Trigger for Apple's Personal Assistant,” Speech and Natural Language Processing, Apple, Oct. 2017, 14 pages, accessed Jan. 6, 2022, retrieved from the internet [URL: … [cited by applicant]
Jahgirdar et al., Build a custom speech-to-text model with speaker diarization capabilities, IBM, dated Jul. 20, 2020, 4 pages. Retrieved on Jan. 6, 2022. Retrieved from the internet [URL: https://developer.ibm.com/]. [cited by applicant]
Jahgirdar et al., Build custom Speech to Text model with speaker diarization capabilities, Github, 14 pages. Retrieved on Jan. 6, 2022. Retrieved from the internet [URL: https://github.com/IBM/build-custom-stt-model-wit… [cited by applicant]
Anchal Bhalla, Building Custom Speech Recognition Models Within Minutes, Medium.com, dated Aug. 26, 2019, 8 pages. Retrieved on Jan. 6, 2022. Retrieved from the internet [URL: https://medium.com/IBM-watson/building-cust… [cited by applicant]
CMUSphinx, Adapting the default acoustic model, Github, 4 pages. Retrieved on Jan. 6, 2022. Retrieved from the Internet [URL: https://cmusphinx.github.io ]. [cited by applicant]
CMUSphinx, Training an acoustic model for CMUSphinx, Github, 10 pages. Retrieved on Jan. 6, 2022. Retrieved from the internet [URL: https://cmusphinx.github.io ]. [cited by applicant]
CN Office Action from CN100738 dated Nov. 15, 2022, 11 pages. [cited by applicant]