IP Library Granted Patent US 11,769,488
Granted Patent B2
US 11,769,488 · App. 17/653,365 · Granted Sep 26, 2023

Meaning inference from speech audio

Inventors: Sudharsan Krishnaswamy (San Jose, CA); Maisy Wieman (Boulder, CO); Jonah Probell (Menlo Park, CA)
Assignee: SoundHound AI IP, LLC
G10L15/063G10L13/02G10L15/16G10L15/187G10L15/1815G10L15/197G10L15/22
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,769,488
App. No.
17/653,365
Granted
Sep 26, 2023
Kind
B2
Abstract

A system and method invoke virtual assistant action, which may comprise an argument. From audio, a probability of an intent is inferred. A probability of a domain and a plurality of variable values may also be inferred. Invoking the action is in response to the intent probability exceeding a threshold. Invoking the action may also be in response to the domain probability exceeding a threshold, a variable value probability exceeding a threshold, detecting an end of utterance, and a specific amount of time having elapsed. The intent probability may increase when the audio includes speech of words with the same meaning in multiple natural languages. Invoking the action may also be conditional on the variable value exceeding its threshold within a certain period of time of the intent probability exceeding its threshold.

Claims (47)

1. A method comprising:

obtaining audio;

inferring, without transcription, from the audio, a plurality of intent probabilities;

inferring, without transcription, from the audio, a plurality of entity value probabilities;

in response to an intent probability exceeding an intent threshold, invoking a virtual assistant action, wherein the virtual assistant action is conditional based on which intent probability is the highest;

wherein the virtual assistant action is conditional based on having at least one entity value probability exceeding an entity value threshold; and

wherein the virtual assistant action is further conditional based on the entity value probability exceeding the entity value threshold within a specific time period of the intent probability exceeding the intent threshold; and

passing the entity value with the highest probability as an argument for the virtual assistant action.

2. The method of claim 1 wherein the inferring of intent probabilities uses a model trained to recognize speech in a plurality of human languages.

3. The method of claim 1 wherein invoking the virtual assistant action is conditional based on having not previously invoked an action within a specific amount of time.

4. The method of claim 1 wherein invoking the virtual assistant action is conditional based on end-of-utterance detection on the audio.

5. The method of claim 1 further comprising:

inferring, without transcription, from the audio, a domain probability,

wherein the virtual assistant action is conditional based on the domain probability exceeding a domain threshold.

6. A method comprising:

obtaining audio;

inferring, without transcription, from the audio:

(a) a domain probability;

(b) a plurality of intent probabilities; and

(c) a plurality of variable value probabilities; and

in response to:

(A) the domain probability exceeding a domain threshold;

(B) an intent probability exceeding an intent threshold when the audio includes speech of words in one of a plurality of recognized natural languages;

(C) a variable value probability exceeding a variable value threshold within a certain period of time of the intent probability exceeding the intent threshold;

(D) an end of an utterance detection signal; and

(E) a specific amount of time having elapsed,

invoking a virtual assistant action comprising an argument indicating which variable value probability is the highest.

7. A device comprising:

a microphone;

a speaker;

a computer processor; and

a memory storing code that causes the computer processor to:

obtain audio from the microphone;

infer, without transcription, from the audio, a plurality of intent probabilities;

infer, without transcription, from the audio, a plurality of entity value probabilities;

in response to an intent probability exceeding an intent threshold, retrieve information, wherein retrieving the information is conditional based on which intent probability is the highest;

wherein the retrieving information is conditional based on having at least one entity value probability exceeding an entity value threshold;

wherein retrieving information is further conditional based on the entity value probability exceeding the entity value threshold within a specific time period of the intent probability exceeding the intent threshold;

pass the entity value with the highest probability as an argument for the information retrieval;

in response to retrieving the information, synthesize speech audio using text-to-speech; and

output the synthesized speech audio to the speaker.

8. The device of claim 7 wherein the inferring of intent probabilities uses a model trained to recognize speech in a plurality of human languages.

9. The device of claim 7 wherein retrieving information is conditional based on having not previously retrieved information within a specific amount of time.

10. The device of claim 7 wherein retrieving information is conditional based on end-of-utterance detection on the audio.

11. The device of claim 7 further comprising:

inferring, without transcription, from the audio, a domain probability,

wherein retrieving information is conditional based on the domain probability exceeding a domain threshold.

Assignments (7)
TERMINATION AND RELEASE OF SECURITY INTEREST IN PATENTS Recorded Dec 3, 2024
From: MONROE CAPITAL MANAGEMENT ADVISORS, LLC, AS COLLATERAL AGENT
To: SOUNDHOUND, INC.
Reel/Frame 069480/0312 →
SECURITY INTEREST Recorded Aug 9, 2024
From: SOUNDHOUND, INC.
To: MONROE CAPITAL MANAGEMENT ADVISORS, LLC, AS COLLATERAL AGENT
Reel/Frame 068526/0413 →
RELEASE OF SECURITY INTEREST Recorded Jun 11, 2024
From: ACP POST OAK CREDIT II LLC, AS COLLATERAL AGENT
To: SOUNDHOUND, INC.; SOUNDHOUND AI IP, LLC
Reel/Frame 067698/0845 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jun 27, 2023
From: SOUNDHOUND AI IP HOLDING, LLC
To: SOUNDHOUND AI IP, LLC
Reel/Frame 064205/0676 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jun 23, 2023
From: SOUNDHOUND, INC.
To: SOUNDHOUND AI IP HOLDING, LLC
Reel/Frame 064083/0484 →
SECURITY INTEREST Recorded Apr 17, 2023
From: SOUNDHOUND, INC.; SOUNDHOUND AI IP, LLC
To: ACP POST OAK CREDIT II LLC
Reel/Frame 063349/0355 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Mar 3, 2022
From: KRISHNASWAMY, SUDHARSAN; WIEMAN, MAISY; PROBELL, JONAH
To: SOUNDHOUND, INC.
Reel/Frame 059162/0321 →