IP Library Granted Patent US 11,295,732
Granted Patent B2
US 11,295,732 · App. 16/529,730 · Granted Apr 5, 2022

Dynamic interpolation for hybrid language models

Inventors: Steffen Holm (San Francisco, CA); Terry Kong (Sunnyvale, CA); Kiran Garaga Lokeswarappa (Mountain View, CA)
Assignee: SoundHound, Inc.
G10L15/197G10L15/02G10L15/16G10L15/1815G10L15/22G10L2015/025
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,295,732
App. No.
16/529,730
Granted
Apr 5, 2022
Kind
B2
Abstract

In order to improve the accuracy of ASR, an utterance is transcribed using a plurality of language models, such as for example, an N-gram language model and a neural language model. The language models are trained separately. They each output a probability score or other figure of merit for a partial transcription hypothesis. Model scores are interpolated to determine a hybrid score. While recognizing an utterance, interpolation weights are chosen or updated dynamically, in the specific context of processing. The weights are based on dynamic variables associated with the utterance, the partial transcription hypothesis, or other aspects of context.

Claims (31)

1. A method of speech transcription, the method comprising:

receiving a speech utterance;

inferring a sequence of phonemes from the speech utterance according to an acoustic model;

computing a plurality of transcription hypotheses from the sequence of phonemes;

computing, according to a first model, a first model score for the transcription;

computing, according to a second model, a second model score for the transcription;

computing a hybrid score by interpolation between the first model score and second model score using interpolation weights, where a first interpolation weight is applied to the first model score and a second interpolation weight is applied to the second model score and at least one of the weights is dynamic and conditioned on the content of the transcription based on word syntax structure and semantic information;

choosing one of the transcriptions based on the hybrid score; and

outputting the chosen hypothesis as the transcription of the speech.

2. The method of claim 1 , wherein the conditioning is based on word presence.

3. The method of claim 1 , wherein the conditioning is based on semantic information.

4. The method of claim 1 , wherein the first model is an n-gram model and the second model is a neural network.

5. The method of claim 1 , wherein the interpolation weights are generated using rule-based logic.

6. The method of claim 1 , wherein the interpolation weights are generated using a neural network.

7. The method of claim 1 , wherein the first model score and second model score are generated using a function that compresses the values of the first model score and the second model score.

8. The method of claim 1 , wherein the interpolation generates the hybrid score using a weighted sum function.

9. A non-transitory computer readable storage medium having embodied thereon a program, the program being executable by a processor to perform a method for speech transcription, the method comprising:

receiving a speech utterance;

computing a plurality of transcription hypotheses from a sequence of phonemes inferred from an acoustic model;

computing, according to a first model, a first model score for the transcription;

computing, according to a second model, a second model score for the transcription;

computing a hybrid score by interpolation between the first model score and second model score using interpolation weights, where a first interpolation weight is applied to the first model score and a second interpolation weight is applied to the second model score and at least one of the weights is dynamic and conditioned on the content of the transcription based on word syntax structure and semantic information;

choosing one of the transcriptions based on the hybrid score; and

outputting the chosen hypothesis as the transcription of the speech.

10. The non-transitory computer readable storage medium of claim 9 , wherein the conditioning is based on word presence.

11. The non-transitory computer readable storage medium of claim 9 , wherein the conditioning is based on semantic information.

12. The non-transitory computer readable storage medium of claim 9 , wherein the first model is an n-gram model and the second model is a neural network.

13. The non-transitory computer readable storage medium of claim 9 , wherein the interpolation weights are generated using rule-based logic.

14. The non-transitory computer readable storage medium of claim 9 , wherein the interpolation weights are generated using a neural network.

15. The non-transitory computer readable storage medium of claim 9 , wherein the first model score and second model score are generated using a function that compresses the values of the first model score and the second model score.

16. The non-transitory computer readable storage medium of claim 9 , wherein the interpolation generates the hybrid score using a weighted sum function.

Assignments (12)
TERMINATION AND RELEASE OF SECURITY INTEREST IN PATENTS Recorded Dec 3, 2024
From: MONROE CAPITAL MANAGEMENT ADVISORS, LLC, AS COLLATERAL AGENT
To: SOUNDHOUND, INC.
Reel/Frame 069480/0312 →
SECURITY INTEREST Recorded Aug 9, 2024
From: SOUNDHOUND, INC.
To: MONROE CAPITAL MANAGEMENT ADVISORS, LLC, AS COLLATERAL AGENT
Reel/Frame 068526/0413 →
RELEASE OF SECURITY INTEREST Recorded Jun 11, 2024
From: ACP POST OAK CREDIT II LLC, AS COLLATERAL AGENT
To: SOUNDHOUND, INC.; SOUNDHOUND AI IP, LLC
Reel/Frame 067698/0845 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jun 27, 2023
From: SOUNDHOUND AI IP HOLDING, LLC
To: SOUNDHOUND AI IP, LLC
Reel/Frame 064205/0676 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jun 23, 2023
From: SOUNDHOUND, INC.
To: SOUNDHOUND AI IP HOLDING, LLC
Reel/Frame 064083/0484 →
RELEASE OF SECURITY INTEREST Recorded Apr 21, 2023
From: FIRST-CITIZENS BANK & TRUST COMPANY, AS AGENT
To: SOUNDHOUND, INC.
Reel/Frame 063411/0396 →
RELEASE OF SECURITY INTEREST Recorded Apr 19, 2023
From: OCEAN II PLO LLC, AS ADMINISTRATIVE AGENT AND COLLATERAL AGENT
To: SOUNDHOUND, INC.
Reel/Frame 063380/0625 →
SECURITY INTEREST Recorded Apr 17, 2023
From: SOUNDHOUND, INC.; SOUNDHOUND AI IP, LLC
To: ACP POST OAK CREDIT II LLC
Reel/Frame 063349/0355 →
CORRECTIVE ASSIGNMENT TO CORRECT THE COVER SHEET PREVIOUSLY RECORDED AT REEL: 056627 FRAME: 0772. ASSIGNOR(S) HEREBY CONFIRMS THE SECURITY INTEREST. Recorded Apr 12, 2023
From: SOUNDHOUND, INC.
To: OCEAN II PLO LLC, AS ADMINISTRATIVE AGENT AND COLLATERAL AGENT
Reel/Frame 063336/0146 →
SECURITY INTEREST Recorded Jun 18, 2021
From: OCEAN II PLO LLC, AS ADMINISTRATIVE AGENT AND COLLATERAL AGENT
To: SOUNDHOUND, INC.
Reel/Frame 056627/0772 →
SECURITY INTEREST Recorded Apr 1, 2021
From: SOUNDHOUND, INC.
To: SILICON VALLEY BANK
Reel/Frame 055807/0539 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Oct 6, 2019
From: STEFFEN, HOLM; KONG, TERRY; LOKESWARAPPA, KIRAN GARAGA
To: SOUNDHOUND, INC.
Reel/Frame 050633/0806 →
Continuity (1)
Related Publication 20210035569A1 · Feb 4, 2021