IP Library Granted Patent US 11,749,281
Granted Patent B2
US 11,749,281 · App. 16/703,783 · Granted Sep 5, 2023

Neural speech-to-meaning

Inventors: Sudharsan Krishnaswamy (San Jose, CA); Maisy Wieman (Boulder, CO); Jonah Probell (Alviso, CA)
Assignee: SoundHound AI IP, LLC
G10L15/26G06F3/167G10L15/183G10L15/1815G10L15/22G10L15/30G10L2015/223
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,749,281
App. No.
16/703,783
Granted
Sep 5, 2023
Kind
B2
Abstract

A neural speech-to-meaning system is trained on speech audio expressing specific intents. The system receives speech audio and produces indications of when the speech in the audio matches the intent. Intents may include variables that can have a large range of values, such as the names of places. The neural speech-to-meaning system simultaneously recognizes enumerated values of variables and general intents. Recognized variable values can serve as arguments to API requests made in response to recognized intents. Accordingly, neural speech-to-meaning supports voice virtual assistants that serve users based on API hits.

Claims (44)

1. A machine for recognizing an intent in speech audio, the machine comprising:

a variable recognizer that processes speech audio features, computes a probability of the speech audio having any of a plurality of enumerated variable values, and outputs the value of the plurality of enumerated variable values with the highest probability; and

an intent recognizer that processes speech audio features, computes a probability of the speech audio having the intent, and in response to the probability being above an intent threshold, produces a request for a virtual assistant action,

wherein the machine has no lexical representation of the speech audio;

the variable recognizer indicates the probability of the speech audio having a value from the plurality of enumerated variable values;

the intent recognizer conditions its output of a request for an action on the probability of the speech audio having a value from the plurality of enumerated variable values; and

the conditioning is based on a delayed indication of the probability of the speech audio having a value from the plurality of enumerated variable values.

2. The machine of claim 1 wherein the intent recognizer conditions its output of a request for an action on which value of the plurality of enumerated variable values has the highest probability.

3. The machine of claim 2 wherein the conditioning is based on a delayed indication of which value of the plurality of enumerated variable values has the highest probability.

4. The machine of claim 1 wherein one of the recognizers produces a score and the other recognizer is called in response to the score being above a score threshold.

5. The machine of claim 1 further comprising a domain recognizer that processes speech audio features and computes a probability of the speech audio referring to a specific domain, wherein the intent recognizer is associated with the domain and called in response to the probability of the speech audio referring to a specific domain being above a domain threshold.

6. The machine of claim 1 wherein no human-readable speech transcription is computed.

7. The machine of claim 1 further comprising a network client with access to a web application programming interface (API), wherein, in response to the intent recognizer producing a request for a virtual assistant action, the network client performs a request to the web API, the request having, as an argument, the value output by the variable recognizer.

8. The machine of claim 7 further comprising a speech synthesis engine, wherein, in response to receiving a response from the web API, the speech synthesis engine synthesizes speech audio containing information from the web API response and outputs the synthesized speech audio for a user of the virtual assistant.

9. A method of recognizing an intent from speech audio by a computer system, the method comprising:

obtaining speech audio;

processing features of the speech audio to compute a probability of the speech audio having any of a plurality of enumerated variable values that are numbers;

outputting the value of the plurality of enumerated variable values with the highest probability;

processing the features of the speech audio to compute a probability of the speech audio having the intent; and

in response to the probability being above an intent threshold, outputting a request for a virtual assistant action,

wherein outputting a request is conditioned on the probability of the value of the plurality of enumerated variable values with the highest probability; and

the conditioning is based on a delayed indication of the probability of the speech audio having a value from the plurality of enumerated variable values.

10. The method of claim 9 wherein outputting a request is conditioned on which value of the plurality of enumerated variable values has the highest probability.

11. The method of claim 10 wherein an indication of the value of the plurality of enumerated variable values is delayed.

12. The method of claim 9 wherein one of the probability computations is performed in response to the other probability computation having a result that is above a threshold.

13. The method of claim 9 further comprising processing features of the speech audio to compute a probability of the speech audio referring to a specific domain, wherein the intent is associated with the domain and computing the probability of the speech audio having the intent is performed in response to the probability of the speech audio referring to the domain being above a threshold.

14. The method of claim 9 wherein no human-readable speech transcription is computed.

15. The method of claim 9 wherein outputting a request for a virtual assistant action is performed by making a request to a web API, the request having as an argument, the value with the highest probability.

16. The method of claim 15 further comprising receiving a response from the web API; and synthesizing speech audio containing information from the web API response.

17. A non-transitory computer readable medium storing code capable of causing one or more computer processors to recognize an intent from speech audio by:

obtaining speech audio;

processing features of the speech audio to compute a probability of the speech audio having any of a plurality of enumerated variable values that are numbers;

outputting the value of the plurality of enumerated variable values with the highest probability;

processing the features of the speech audio to compute a probability of the speech audio having the intent; and

in response to the probability being above an intent threshold, outputting a request for a virtual assistant action,

wherein outputting a request is conditioned on the probability of the value of the plurality of enumerated variable values with the highest probability; and

the conditioning is based on a delayed indication of the probability of the speech audio having a value from the plurality of enumerated variable values.

18. The medium of claim 17 wherein outputting a request is conditioned on which value of the plurality of enumerated variable values has the highest probability.

19. The medium of claim 18 wherein an indication of the value of the plurality of enumerated variable values is delayed.

20. The medium of claim 17 wherein one of the probability computations is performed in response to the other probability computation having a result that is above a threshold.

21. The medium of claim 17 further comprising processing features of the speech audio to compute a probability of the speech audio referring to a specific domain, wherein the intent is associated with the domain and computing the probability of the speech audio having the intent is performed in response to the probability of the speech audio referring to the domain being above a threshold.

22. The medium of claim 17 wherein no human-readable speech transcription is computed.

23. The medium of claim 17 wherein outputting a request for a virtual assistant action is performed by making a request to a web API, the request having as an argument, the value with the highest probability.

24. The medium of claim 23 further comprising receiving a response from the web API; and synthesizing speech audio containing information from the web API response.

Assignments (7)
TERMINATION AND RELEASE OF SECURITY INTEREST IN PATENTS Recorded Dec 3, 2024
From: MONROE CAPITAL MANAGEMENT ADVISORS, LLC, AS COLLATERAL AGENT
To: SOUNDHOUND, INC.
Reel/Frame 069480/0312 →
SECURITY INTEREST Recorded Aug 9, 2024
From: SOUNDHOUND, INC.
To: MONROE CAPITAL MANAGEMENT ADVISORS, LLC, AS COLLATERAL AGENT
Reel/Frame 068526/0413 →
RELEASE OF SECURITY INTEREST Recorded Jun 11, 2024
From: ACP POST OAK CREDIT II LLC, AS COLLATERAL AGENT
To: SOUNDHOUND, INC.; SOUNDHOUND AI IP, LLC
Reel/Frame 067698/0845 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jun 27, 2023
From: SOUNDHOUND AI IP HOLDING, LLC
To: SOUNDHOUND AI IP, LLC
Reel/Frame 064205/0676 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jun 23, 2023
From: SOUNDHOUND, INC.
To: SOUNDHOUND AI IP HOLDING, LLC
Reel/Frame 064083/0484 →
SECURITY INTEREST Recorded Apr 17, 2023
From: SOUNDHOUND, INC.; SOUNDHOUND AI IP, LLC
To: ACP POST OAK CREDIT II LLC
Reel/Frame 063349/0355 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Dec 5, 2019
From: KRISHNASWAMY, SUDHARSAN; WIEMAN, MAISY; PROBELL, JONAH
To: SOUNDHOUND, INC.
Reel/Frame 051195/0637 →
Continuity (1)
Related Publication 20210174806A1 · Jun 10, 2021
Cited By (1)
US 12,561,536