IP Library › Granted Patent US 12,300,219
Granted Patent B2
US 12,300,219 · App. 18/474,853 · Granted May 13, 2025

Meaning inference from speech audio

Inventors: Sudharsan Krishnaswamy (San Jose, CA); Maisy Wieman (Boulder, CO); Jonah Probell (Alviso, CA)
Assignee: SoundHound AI IP, LLC
G10L15/063G10L13/02G10L15/16G10L15/1815G10L15/187G10L15/197G10L15/22
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,300,219
App. No.
18/474,853
Granted
May 13, 2025
Kind
B2
Abstract

A system and method invoke virtual assistant action, which may comprise an argument. From audio, a probability of an intent is inferred. A probability of a domain and a plurality of variable values may also be inferred. Invoking the action is in response to the intent probability exceeding a threshold. Invoking the action may also be in response to the domain probability exceeding a threshold, a variable value probability exceeding a threshold, detecting an end of utterance, and a specific amount of time having elapsed. The intent probability may increase when the audio includes speech of words with the same meaning in multiple natural languages. Invoking the action may also be conditional on the variable value exceeding its threshold within a certain period of time of the intent probability exceeding its threshold.

Claims (53)

1. A method comprising: obtaining audio;

inferring, from the audio, a plurality of intent probabilities;

inferring, from the audio, a plurality of entity value probabilities;

in response to an intent probability exceeding an intent threshold, invoking a virtual assistant action, wherein the virtual assistant action is conditional based on a variable value probability exceeding a variable value threshold;

wherein the virtual assistant action is further conditional based on t he variable value probability exceeding the variable value threshold within a specific time period of the intent probability exceeding the intent threshold; and

passing an entity value with the highest probability as an argument for the virtual assistant action.

2. The method of claim 1 wherein the inferring of intent probabilities uses a model trained to recognize speech in a plurality of human languages.

3. The method of claim 1 wherein invoking the virtual assistant action is conditional based on having not previously invoked an action within a specific amount of time.

4. The method of claim 1 wherein invoking the virtual assistant action is conditional based on end-of-utterance detection on the audio.

5. The method of claim 1 further comprising:

directly inferring, from the audio, a plurality of variable value probabilities,

wherein the virtual assistant action comprises an argument indicating which variable value probability is the highest.

6. The method of claim 1 further comprising:

directly inferring, from the audio, a domain probability,

wherein the virtual assistant action is conditional based on the domain probability exceeding a domain threshold.

7. The method of claim 1 further comprising:

directly inferring, from the audio, a domain probability,

wherein the virtual assistant action is conditional based on the domain probability exceeding a domain threshold;

directly inferring, from the audio, a plurality of variable value probabilities; and

the virtual assistant action comprising an argument indicating which variable value probability is the highest.

8. A method comprising:

obtaining audio;

inferring, from the audio:

(a) a domain probability;

(b) a plurality of intent probabilities; and

(c) a plurality of variable value probabilities; and

in response to:

(A) the domain probability exceeding a domain threshold;

(B) an intent probability exceeding an intent threshold when the audio includes speech of words in one of a plurality of recognized natural languages;

(C) a variable value probability exceeding a variable value threshold within a certain period of time of the intent probability exceeding the intent threshold;

(D) an end of an utterance detection signal; and

(E) a specific amount of time having elapsed,

invoking a virtual assistant action comprising an argument indicating which variable value probability is the highest.

9. A device comprising:

a microphone;

a speaker;

one or more computer processors; and

a memory storing code that causes the one or more computer processors to:

obtain audio from the microphone;

infer, from the audio, a plurality of intent probabilities;

infer, from the audio, a plurality of entity value probabilities;

in response to an intent probability exceeding an intent threshold, retrieve information, wherein retrieving the information is conditional based on which intent probability is the highest;

wherein the retrieving information is conditional based on having at least one entity value probability exceeding an entity value threshold;

wherein retrieving information is further conditional based on the entity value probability exceeding the entity value threshold within a specific time period of the intent probability exceeding the intent threshold;

pass an entity value with the highest probability as an argument for the information retrieval;

in response to retrieving the information, synthesize speed audio using text-to-speech; and

output the synthesized speech audio to the speaker.

10. The device of claim 9 wherein the inferring of intent probabilities uses a model trained to recognize speech in a plurality of human languages.

11. The device of claim 9 wherein retrieving information is conditional based on having not previously retrieved information within a specific amount of time.

12. The device of claim 9 wherein retrieving information is conditional based on end-of-utterance detection on the audio.

13. The device of claim 9 further comprising:

inferring, without transcription, from the audio, a domain probability,

wherein retrieving informamtion is conditional based on the domain probability exceeding a domain threshold.

Assignments (3)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Oct 31, 2023
From: KRISHNASWAMY, SUDHARSAN; WIEMAN, MAISY; PROBELL, JONAH
To: SOUNDHOUND, INC.
Reel/Frame 065409/0508 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Oct 31, 2023
From: SOUNDHOUND, INC.
To: SOUNDHOUND AI IP HOLDING, LLC
Reel/Frame 065414/0108 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Oct 31, 2023
From: SOUNDHOUND AI IP HOLDING, LLC
To: SOUNDHOUND AI IP, LLC
Reel/Frame 065414/0149 →
Continuity (3)
Continuation 17653365 · Mar 3, 2022
Continuation 16704216 · Dec 5, 2019
Related Publication 20240046918A1 · Feb 8, 2024
References Cited (64)
US 8527276B1 · Senior et al. · 2013 [cited by applicant]
US 10140977B1 · Raux et al. · 2018 [cited by applicant]
US 10867122B1 · Mertens et al. · 2020 [cited by applicant]
US 20080215322A1 · Fischer et al. · 2008 [cited by applicant]
US 20140278424A1 · Deng et al. · 2014 [cited by applicant]
US 20140365216A1 · Gruber et al. · 2014 [cited by applicant]
US 20150348549A1 · Giuli et al. · 2015 [cited by applicant]
US 20170083586A1 · Huang · 2017 [cited by examiner]
US 20170178626A1 · Gruber et al. · 2017 [cited by applicant]
US 20170287465A1 · Zhao et al. · 2017 [cited by applicant]
US 20180061394A1 · Kim et al. · 2018 [cited by applicant]
US 20180358005A1 · Tomar · 2018 [cited by examiner]
US 20190066668A1 · Lin et al. · 2019 [cited by applicant]
US 20190132451A1 · Kannan · 2019 [cited by examiner]
US 20190138269A1 · Dolph et al. · 2019 [cited by applicant]
US 20190179890A1 · Evermann · 2019 [cited by examiner]
US 20190251952A1 · Arik et al. · 2019 [cited by applicant]
US 20190281159A1 · Segalis et al. · 2019 [cited by applicant]
US 20200035228A1 · Seo et al. · 2020 [cited by applicant]
US 20200103963A1 · Kelly · 2020 [cited by examiner]
US 20200410976A1 · Zhou et al. · 2020 [cited by applicant]
US 20210074295A1 · Moreno · 2021 [cited by examiner]
US 20210082406A1 · Kim · 2021 [cited by applicant]
US 20210082412A1 · Kennewick · 2021 [cited by applicant]
US 20210225357A1 · Zhao et al. · 2021 [cited by applicant]
WO 2017091883 · 2017 [cited by applicant]
Y.-P. Chen, R. Price and S. Bangalore, “Spoken Language Understanding without Speech Recognition,” 2018 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), 2018, pp. 6189-6193, doi: 10.110… [cited by examiner]
Alex Graves, Neural Turing machines. arXiv preprint arXiv:1410.5401. Oct. 20, 2014. [cited by applicant]
Anonymous, Artificial Design: Modeling Artificial Super Intelligence with Extended General Relativity and Universal Darwinism via Geometrization for Universal Design Automation. [cited by applicant]
Ashish Vaswani, Attention is all you need. InAdvances in neural information processing systems 2017 {pp. 0998-6008). [cited by applicant]
Bing Liu et al: Attention-Based Recurrent Neural Network Models for Joint Intent Detection and Slot Filling, Interspeech 2016, Sep. 8, 2016 {Sep. 8, 2016), pp. 1-5, XP080724506, DOI: 10.21437/Interspeech.2016-1352. [cited by applicant]
Bing Liu, Joint online spoken language understanding and language modeling with recurrent neural networks. arXiv preprint arXiv:1609.01462. Sep. 6, 2016. [cited by applicant]
Chen, Qian, Zhu Zhuo, and Wen Wang. “Bert for joint intent classification and slot filling_” arXiv preprint arXiv:1902.10909 (2019). [cited by applicant]
Chenwei Zhang, Joint slot filling and intent detection via capsule neural networks_ arXiv preprint arXiv:1812.09471. Dec. 22, 2018. [cited by applicant]
Coucke, Alice, Alaa Saade, Adrien Ball, Theodore Bluche, Alexandre Caulier, David Leroy, Clement Doumouro et al. Snips voice platform: an embedded spoken language understanding system for private-by-design voice interfa… [cited by applicant]
Douglas Coimbra De Andrade, A neural attention model for speech command recognition. arXiv preprint arXiv:1808.08929_ Aug. 27, 2018. [cited by applicant]
Dzmitry Bahdanau, Neural machine translation by jointly learning to align and translate_ arXiv preprint arXiv:1409.0473. [cited by applicant]
Homa B Hashemi, Query intent detection using convolutional neural networks. In International Conference on Web Search and Data Mining, Workshop on Query Understanding 2016. [cited by applicant]
Hong-Kwang J Kuo et al: End-to-End Spoken Language Understanding Without Full Transcripts, Arxiv.org, Sep. 30, 2020 {Sep. 30, 2020). [cited by applicant]
James Kirkpatrick, Overcoming catastrophic forgetting in neural networks. Proceedings of the national academy of sciences. Mar. 28, 2017; 114(13):3521-6. [cited by applicant]
Kiaodong Liu. Multi-task deep neural networks for natural language understanding. arXiv preprint arXiv: 1901.11504 Jan. 31, 2019. [cited by applicant]
Neil Zeghidour, End-to-end speech recognition from the raw waveform. arXiv preprint arXiv:18O6.O7O98. Jun. 19, 2018. [cited by applicant]
Oriol Vinyals, Grammar as a foreign language_ InAdvances in neural information processing systems 2015 (pp. 2773-2781). [cited by applicant]
Shah, Darsh J., Raghav Gupta, Amir A_ Fayazi, and Dilek Hakkani-Tur. “Robust zero-shot cross-domain slot filling with example values.” arXiv preprint arXiv:1906.06870 (2019). [cited by applicant]
Shiyu Chang, Dilated recurrent neural networks. InAdvances in Neural Information Processing Systems 2017 (pp. 77-87). [cited by applicant]
Tara N Sainath, Learning the speech front-end with raw waveform CLDNNs. In Sixteenth Annual Conference of the International Speech Communication Association 2015. [cited by applicant]
Mctor Zhong, Seq2sql: Generating structured queries from natural language using reinforcement learning. arXiv preprinl arXiv:1709.00103. Aug. 31, 2017. [cited by applicant]
Yanzhang He, Streaming End-to-end Speech Recognition For Mobile Devices. In ICASSP 2019-2019 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) May 12, 2019 (pp. 6381-6385). IEEE. [cited by applicant]
Zoltan Tuske, Acoustic modeling with deep neural networks using raw lime signal for LVCSR. In Fifteenth annual Conference of the international speech communication association 2014. [cited by applicant]
Office Action dated Dec. 7, 2022 in U.S. Appl. No. 17/653,365. [cited by applicant]
Y.-P. Chen, R. Price and S. Bangalore, “Spoken Language Understanding without Speech Recognition,” 2018 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), 2018, pp. 6189-6193, doi: 10.110… [cited by applicant]
Response to Office Action dated Dec. 19, 2022 in U.S. Appl. No. 17/653,365. [cited by applicant]
Notice of Allowance and Fee(s) Due dated Jun. 6, 2023 in U.S. Appl. No. 17/653,365. [cited by applicant]
Ahmed Ali et al., “Automatic Dialect Detection in Arabic Broadcast Speech”, arXiv preprint arXiv: 1509.06928, Aug. 2016. [cited by applicant]
Oliver Bender et al., “Maximum entropy models for named entity recognition,” In Proceedings of the Seventh Conference on Natural Language Learning at HLT-NAACL 2003—vol. 4 (CON LL '03) Association for Computational Ling… [cited by applicant]
D. Serdyuk et al., “Towards End-to-End Spoken Language Understanding,” 2018 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), pp. 5754-5758, Apr. 20, 2018. [cited by applicant]
Office Action dated Jul. 9, 2021 in U.S. Appl. No. 16/704,216. [cited by applicant]
Jisung Wang, Speech Augmentation Using Wavenet in Speech Recognition. InICASSP 2019-2019 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) May 12, 2019 (pp. 6770-6774). IEEE. [cited by applicant]
Rohit Prabhavalkar, A Comparison of Sequence-lo-Sequence Models for Speech Recognition. InInterspeech Aug. 2017 (pp. 939-943). [cited by applicant]
Response to Office Action dated Sep. 13, 2021 in U.S. Appl. No. 16/704,216. [cited by applicant]
Notice of Allowance and Fee(s) Due dated Dec. 16, 2021 in U.S. Appl. No. 16/704,216. [cited by applicant]
Extended European Search Report dated Mar. 27, 2024, in European Patent Application No. 24156087.9. [cited by applicant]
Yinghui Huang et al., “Leveraging Unpaired Text Data For Training End-to-End Speech-to-Intent Systems”, ICASSP 2020—2020 IEEE International Conference On Acoustics, Speech and Signal Processing (ICASSP), IEEE, May 4, 20… [cited by applicant]
Office Action dated Nov. 5, 2024, in U.S. Appl. No. 18/461,212. [cited by applicant]