IP Library Granted Patent US 10,229,683
Granted Patent B2
US 10,229,683 · App. 15/456,354 · Granted Mar 12, 2019

Speech-enabled system with domain disambiguation

Inventor: Rainer Leeb (San Jose, CA)
Assignee: SoundHound, Inc.
G10L15/22G10L15/02G10L15/18G10L2015/221
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 10,229,683
App. No.
15/456,354
Granted
Mar 12, 2019
Kind
B2
Abstract

Systems perform methods of interpreting spoken utterances from a user and responding to the utterances by providing requested information or performing a requested action. The utterances are interpreted in the context of multiple domains. Each interpretation is assigned a relevancy score based on how well the interpretation represents what the speaker intended. Interpretations having a relevancy score below a threshold for its associated domain are discarded. A remaining interpretation is chosen based on choosing the most relevant domain for the utterance. The user may be prompted to provide disambiguation information that can be used to choose the best domain. Storing past associations of utterance representation and domain choice allows for measuring the strength of correlation between uttered words and phrases with relevant domains. This correlation strength information may allow the system to automatically disambiguate alternate interpretations without requiring user input.

Claims (64)

1. A method of providing a result to a user of a speech-enabled system, the method comprising:

using at least one computer to:

interpret a natural language utterance according to a plurality of domains to automatically create both a particular interpretation and a corresponding relevancy score for each domain, where a domain represents a subject area and comprises a set of grammar rules, such that (i) a first particular interpretation of the natural language utterance has a particular meaning that is unique to a first domain of the domains and (ii) a second particular interpretation of the natural language utterance has a particular meaning that is both unique to a second domain of the domains and different from the particular meaning that is unique to the first domain;

compare each relevancy score against a threshold for its domain to determine a list of candidate domains having relevancy scores exceeding their thresholds;

present the list of candidate domains to the user, wherein the presented list only identifies domains having relevancy scores exceeding their thresholds;

ask the user to choose a domain from the presented list;

receive a choice of domain from the user;

store the choice in a database;

train a neural network implementing domain scoring, the neural network being trained by processing the database; and

implement the trained neural network to automatically identify a most likely domain for a natural language utterance.

2. The method of claim 1 further comprising using the at least one computer to increment a value of a counter representing the chosen domain.

3. The method of claim 2 wherein the candidate domains are presented to the user in an order based on the value of the counter.

4. The method of claim 1 further comprising using the at least one computer to store an indication of the most recently chosen domain.

5. The method of claim 4 wherein the candidate domains are presented to the user in an order based on the indication of the most recently chosen domain.

6. The method of claim 1 further comprising:

using the at least one computer to:

store a record in a database, the record comprising:

a representation of the natural language utterance; and

the choice of domain for the utterance.

7. The method of claim 1 further comprising:

using the at least one computer to:

store a record in a database, the record comprising:

the interpretation of the utterance according to the chosen domain; and

the choice of domain.

8. The method of claim 1 , wherein a particular threshold is assigned to each domain and at least two particular thresholds are different for at least two domains.

9. The method of claim 1 , wherein audio speech is used to (i) present the list of candidate domains to the user and (ii) ask the user to choose the domain from the presented list.

10. A method of providing a result to a user of a speech-enabled system, the method comprising:

using at least one computer to:

interpret a natural language utterance according to a plurality of domains to automatically create both a particular interpretation and a corresponding relevancy score for each domain, where a domain represents a subject area and comprises a set of grammar rules, such that (i) a first particular interpretation of the natural language utterance has a particular meaning that is unique to a first domain of the domains and (ii) a second particular interpretation of the natural language utterance has a particular meaning that is both unique to a second domain of the domains and different from the particular meaning that is unique to the first domain;

compare each relevancy score against a threshold for its domain to determine a number of candidate domains having relevancy scores exceeding their thresholds; and

responsive to the number of candidate domains being greater than a maximum number of domains that is reasonable to present to the user for disambiguation, ask the user to provide a general clarification;

receive a response utterance with clarification information to determine a domain;

store the determined domain in a database;

train a neural network implementing domain scoring, the neural network being trained by processing the database; and

implement the trained neural network to automatically identify a most likely domain for a natural language utterance.

11. The method of claim 10 wherein the maximum number of domains that is reasonable to present to the user is based on environment information.

12. A non-transitory computer readable medium storing code that, when executed by at least one computer, causes the one or more computer to:

interpret a natural language utterance according to a plurality of domains to automatically create both a particular interpretation and a corresponding relevancy score for each domain, where a domain represents a subject area and comprises a set of grammar rules, such that (i) a first particular interpretation of the natural language utterance has a particular meaning that is unique to a first domain of the domains and (ii) a second particular interpretation of the natural language utterance has a particular meaning that is both unique to a second domain of the domains and different from the particular meaning that is unique to the first domain;

compare each relevancy score against a threshold for its domain to determine a list of candidate domains having relevancy scores exceeding their thresholds;

present the list of candidate domains to a user, wherein the presented list only identifies domains having relevancy scores exceeding their thresholds;

ask the user to choose a domain from the presented list;

receive a choice of domain from the user;

store the choice in a database;

train a neural network implementing domain scoring, the neural network being trained by processing the database; and

implement the trained neural network to automatically identify a most likely domain for a natural language utterance.

13. A speech-enabled system comprising:

means for performing disambiguation by:

interpreting a natural language utterance according to a plurality of domains to automatically produce a particular interpretation for each domain, where a domain represents a subject area and comprises a set of grammar rules, such that (i) a first particular interpretation of the natural language utterance has a particular meaning that is unique to a first domain of the domains and (ii) a second particular interpretation of the natural language utterance has a particular meaning that is both unique to a second domain of the domains and different from the particular meaning that is unique to the first domain;

determining that the natural language utterance is sensible in a plurality of domains;

asking a user for disambiguation;

receiving a response utterance with clarification information;

storing the clarification in a database;

training a neural network implementing domain scoring, the neural network being trained by processing the database; and

implementing the trained neural network to automatically identify a most likely domain for a natural language utterance.

14. An automotive platform comprising:

a speech capture module enabled to capture a spoken utterance from a user;

a speech recognition module that interprets the utterance according to a plurality of domains to automatically produce both a particular interpretation and a corresponding relevancy score for each domain, where a domain represents a subject area and comprises a set of grammar rules, such that (i) a first particular interpretation of the utterance has a particular meaning that is unique to a first domain of the domains and (ii) a second particular interpretation of the utterance has a particular meaning that is both unique to a second domain of the domains and different from the particular meaning that is unique to the first domain; and

a speech generation module enabled to produce speech,

wherein, responsive to only relevancy scores exceeding an associated threshold, the speech generation module:

produces speech comprising a list of domains;

asks the user to choose a domain from the list

stores the choice in a database;

train a neural network implementing domain scoring, the neural network being trained by processing the database; and

implement the trained neural network to automatically identify a most likely domain for a natural language utterance.

Assignments (11)
TERMINATION AND RELEASE OF SECURITY INTEREST IN PATENTS Recorded Dec 3, 2024
From: MONROE CAPITAL MANAGEMENT ADVISORS, LLC, AS COLLATERAL AGENT
To: SOUNDHOUND, INC.
Reel/Frame 069480/0312 →
SECURITY INTEREST Recorded Aug 9, 2024
From: SOUNDHOUND, INC.
To: MONROE CAPITAL MANAGEMENT ADVISORS, LLC, AS COLLATERAL AGENT
Reel/Frame 068526/0413 →
RELEASE OF SECURITY INTEREST Recorded Jun 11, 2024
From: ACP POST OAK CREDIT II LLC, AS COLLATERAL AGENT
To: SOUNDHOUND, INC.; SOUNDHOUND AI IP, LLC
Reel/Frame 067698/0845 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jun 27, 2023
From: SOUNDHOUND AI IP HOLDING, LLC
To: SOUNDHOUND AI IP, LLC
Reel/Frame 064205/0676 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jun 23, 2023
From: SOUNDHOUND, INC.
To: SOUNDHOUND AI IP HOLDING, LLC
Reel/Frame 064083/0484 →
RELEASE OF SECURITY INTEREST Recorded Apr 21, 2023
From: FIRST-CITIZENS BANK & TRUST COMPANY, AS AGENT
To: SOUNDHOUND, INC.
Reel/Frame 063411/0396 →
RELEASE OF SECURITY INTEREST Recorded Apr 19, 2023
From: OCEAN II PLO LLC, AS ADMINISTRATIVE AGENT AND COLLATERAL AGENT
To: SOUNDHOUND, INC.
Reel/Frame 063380/0625 →
SECURITY INTEREST Recorded Apr 17, 2023
From: SOUNDHOUND, INC.; SOUNDHOUND AI IP, LLC
To: ACP POST OAK CREDIT II LLC
Reel/Frame 063349/0355 →
CORRECTIVE ASSIGNMENT TO CORRECT THE COVER SHEET PREVIOUSLY RECORDED AT REEL: 056627 FRAME: 0772. ASSIGNOR(S) HEREBY CONFIRMS THE SECURITY INTEREST. Recorded Apr 12, 2023
From: SOUNDHOUND, INC.
To: OCEAN II PLO LLC, AS ADMINISTRATIVE AGENT AND COLLATERAL AGENT
Reel/Frame 063336/0146 →
SECURITY INTEREST Recorded Jun 18, 2021
From: OCEAN II PLO LLC, AS ADMINISTRATIVE AGENT AND COLLATERAL AGENT
To: SOUNDHOUND, INC.
Reel/Frame 056627/0772 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Apr 11, 2017
From: LEEB, RAINER
To: SOUNDHOUND, INC.
Reel/Frame 041969/0632 →
Continuity (1)
Related Publication 20180261216A1 · Sep 13, 2018