IP Library Granted Patent US 10,096,317
Granted Patent B2
US 10,096,317 · App. 15/131,833 · Granted Oct 9, 2018

Hierarchical speech recognition decoder

Inventors: Ethan Selfridge (Jamaica Plain, MA); Michael Johnston (New York, NY)
Assignee: INTERACTIONS LLC
G10L15/197G10L15/02G10L15/063G10L2015/0631
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 10,096,317
App. No.
15/131,833
Granted
Oct 9, 2018
Kind
B2
Abstract

A speech interpretation module interprets the audio of user utterances as sequences of words. To do so, the speech interpretation module parameterizes a literal corpus of expressions by identifying portions of the expressions that correspond to known concepts, and generates a parameterized statistical model from the resulting parameterized corpus. When speech is received the speech interpretation module uses a hierarchical speech recognition decoder that uses both the parameterized statistical model and language sub-models that specify how to recognize a sequence of words. The separation of the language sub-models from the statistical model beneficially reduces the size of the literal corpus needed for training, reduces the size of the resulting model, provides more fine-grained interpretation of concepts, and improves computational efficiency by allowing run-time incorporation of the language sub-models.

Claims (32)

1. A computer-implemented method of a voice server for producing user-specific interpretations of user utterances, the method comprising:

accessing, by the voice server, a literal speech recognition corpus comprising a plurality of expressions;

accessing, by the voice server, a concept tagging module for identifying instances of a plurality of concepts within an expression;

generating, by the voice server, a parameterized speech recognition corpus by applying the concept tagging module to the expressions of the literal speech recognition model corpus in order to identify, within the expressions, portions of the expressions that are instances of the concepts;

generating, by the voice server, a parameterized statistical model based on the parameterized speech recognition model corpus, the parameterized statistical model indicating a plurality of probability scores for a corresponding plurality of n-grams, some of the plurality of n-grams including a placeholder indicating one of the plurality of concepts;

accessing, by the voice server, a plurality of language sub-models corresponding to the plurality of concepts, at least some of the plurality of language sub-models being customized for a user;

receiving, by the voice server over a computer network, an utterance of the user, the utterance having been accepted from the user at a client device as spoken input; and

generating, by the voice server, a user-specific interpretation of the utterance using both the parameterized statistical model and ones of the plurality of language sub-models corresponding to instances of the concepts in the utterance, the interpretation comprising a sequence of recognized words.

2. A computer-implemented method, comprising:

accessing a literal speech recognition corpus comprising a plurality of expressions, an expression comprising a sequence of word tokens;

accessing a concept tagging module for identifying instances of a plurality of concepts within an expression and for replacing expressions with placeholders that indicate classes associated with concepts;

generating a parameterized speech recognition corpus by using the concept tagging module to identify, within the expressions of the literal speech recognition corpus, portions of the expressions that are instances of the concepts and to replace the identified portions of the expressions with placeholders; and

generating a parameterized statistical model based on the parameterized speech recognition corpus

receiving, over a computer network, an utterance of a user, the utterance having been accepted from the user at a client device as spoken input; and

generating a text interpretation of the utterance using the parameterized statistical model together with a language sub-model corresponding to one of the plurality of concepts.

3. The computer-implemented method of claim 2 , wherein the language sub-model is customized for the user, and wherein generating the interpretation of the utterance comprises generating a user-specific interpretation of the utterance using the language sub-model customized for the user.

4. The computer-implemented method of claim 2 , wherein the interpretation comprises a plurality of phrases and a corresponding plurality of probability scores for the phrases.

5. The computer-implemented method of claim 2 , wherein the interpretation comprises a lattice in which nodes of the lattice are literal word tokens and edges between the nodes have weights indicating probabilities that the corresponding literal word tokens occur in sequence.

6. The computer-implemented method of claim 2 , wherein the parameterized statistical model indicates a plurality of probability scores for a corresponding plurality of n-grams, some of the plurality of n-grams including a placeholder indicating one of the plurality of concepts.

7. The computer-implemented method of claim 2 , further comprising training the concept tagging module to identify the instances of the plurality of concepts by analyzing expressions of the literal speech recognition corpus that are labeled with concepts that the expressions represent.

8. A non-transitory computer-readable storage medium storing instructions executable by a computer processor, the instructions comprising:

instructions for accessing a literal speech recognition corpus comprising a plurality of expressions, an expression comprising a sequence of word tokens;

instructions for accessing a concept tagging module for identifying instances of a plurality of concepts within an expression and for replacing expressions with placeholders that indicate classes associated with concepts;

instructions for generating a parameterized speech recognition corpus by using the concept tagging module to identify, within the expressions of the literal speech recognition corpus, portions of the expressions that are instances of the concepts and to replace the identified portions of the expressions with placeholders; and

instructions for generating a parameterized statistical model based on the parameterized speech recognition corpus

instructions for receiving, over a computer network, an utterance of a user, the utterance having been accepted from the user at a client device as spoken input; and

instructions for generating a text interpretation of the utterance using the parameterized statistical model together with a language sub-model corresponding to one of the plurality of concepts.

9. The non-transitory computer-readable storage medium of claim 8 , wherein the language sub-model is customized for the user, and wherein generating the interpretation of the utterance comprises generating a user-specific interpretation of the utterance using the language sub-model customized for the user.

10. The non-transitory computer-readable storage medium of claim 8 , wherein the interpretation comprises a plurality of phrases and a corresponding plurality of probability scores for the phrases.

11. The non-transitory computer-readable storage medium of claim 8 , wherein the interpretation comprises a lattice in which nodes of the lattice are literal word tokens and edges between the nodes have weights indicating probabilities that the corresponding literal word tokens occur in sequence.

12. The non-transitory computer-readable storage medium of claim 8 , wherein the parameterized statistical model indicates a plurality of probability scores for a corresponding plurality of n-grams, some of the plurality of n-grams including a placeholder indicating one of the plurality of concepts.

13. The non-transitory computer-readable storage medium of claim 8 , the instructions further comprising instructions for training the concept tagging module to identify the instances of the plurality of concepts by analyzing expressions of the literal speech recognition corpus that are labeled with concepts that the expressions represent.

Assignments (8)
RELEASE OF SECURITY INTEREST Recorded Sep 4, 2025
From: RUNWAY GROWTH FINANCE CORP., AS AGENT
To: INTERACTIONS CORPORATION; INTERACTIONS LLC
Reel/Frame 072802/0931 →
CORRECTIVE ASSIGNMENT TO CORRECT THE THE APPLICATION NUMBER PREVIOUSLY RECORDED AT REEL: 060445 FRAME: 0733. ASSIGNOR(S) HEREBY CONFIRMS THE ASSIGNMENT. Recorded Feb 1, 2023
From: INTERACTIONS LLC; INTERACTIONS CORPORATION
To: RUNWAY GROWTH FINANCE CORP.
Reel/Frame 062919/0063 →
RELEASE OF SECURITY INTEREST IN INTELLECTUAL PROPERTY RECORDED AT REEL/FRAME: 049388/0152 Recorded Jul 1, 2022
From: SILICON VALLEY BANK
To: INTERACTIONS LLC
Reel/Frame 060558/0719 →
RELEASE OF SECURITY INTEREST IN INTELLECTUAL PROPERTY RECORDED AT REEL/FRAME: 049388/0082 Recorded Jun 30, 2022
From: SILICON VALLEY BANK
To: INTERACTIONS LLC
Reel/Frame 060558/0474 →
INTELLECTUAL PROPERTY SECURITY AGREEMENT Recorded Jun 27, 2022
From: INTERACTIONS LLC; INTERACTIONS CORPORATION
To: RUNWAY GROWTH FINANCE CORP.
Reel/Frame 060445/0733 →
INTELLECTUAL PROPERTY SECURITY AGREEMENT Recorded Jun 5, 2019
From: INTERACTIONS LLC
To: SILICON VALLEY BANK
Reel/Frame 049388/0082 →
FIRST AMENDMENT TO AMENDED AND RESTATED INTELLECTUAL PROPERTY SECURITY AGREEMENT Recorded Jun 5, 2019
From: INTERACTIONS LLC
To: SILICON VALLEY BANK
Reel/Frame 049388/0152 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jul 15, 2016
From: SELFRIDGE, ETHAN; JOHNSTON, MICHAEL J.
To: INTERACTIONS LLC
Reel/Frame 039170/0792 →
Continuity (1)
Related Publication 20170301346A1 · Oct 19, 2017
Cited By (1)
US 12,573,401