IP Library › Granted Patent US 12,512,096
Granted Patent B2
US 12,512,096 · App. 17/340,378 · Granted Dec 30, 2025

Speech recognition using dialog history

Inventors: Behnam Hedayatnia (San Francisco, CA); Anirudh Raju (San Jose, CA); Ankur Gandhe (Bothell, WA); Chandra Prakash Khatri (San Jose, CA); Ariya Rastrow (Seattle, WA); Anushree Venkatesh (San Mateo, CA); Arindam Mandal (Redwood City, CA); Raefer Christopher Gabriel (San Jose, CA); Ahmad Shikib Mehri (Sunnyvale, CA)
Assignee: Amazon Technologies, Inc.
G10L15/19G10L15/22G10L15/30G10L19/00
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,512,096
App. No.
17/340,378
Filed
Jun 7, 2021
Granted
Dec 30, 2025
Kind
B2
Art Unit
2656
USPC
704/235
Abstract

Described herein is a system for rescoring automatic speech recognition hypotheses for conversational devices that have multi-turn dialogs with a user. The system leverages dialog context by incorporating data related to past user utterances and data related to the system generated response corresponding to the past user utterance. Incorporation of this data improves recognition of a particular user utterance within the dialog.

Claims (62)

1 . A computer-implemented method, comprising:

receiving input audio data corresponding to a first utterance;

performing speech recognition using the input audio data to determine first data indicating at a least a first speech recognition hypothesis representing the first utterance;

encoding the first data to generate first encoded data;

receiving second encoded data corresponding to a past utterance;

processing the first encoded data and the second encoded data using a machine learning model to determine model output data indicating a second speech recognition hypothesis as representing the first utterance, the second speech recognition hypothesis being different from the first speech recognition hypothesis;

processing the second speech recognition hypothesis to generate second data indicating an action to be performed in response to the first utterance; and

generating, using the second data, output data responsive to the first utterance.

2 . The computer-implemented method of claim 1 , further comprising:

receiving third data corresponding to the past utterance; and

encoding the third data to generate the second encoded data.

3 . The computer-implemented method of claim 2 , further comprising:

determining fourth data representing content of a user input corresponding to the past utterance; and

including the fourth data in the third data.

4 . The computer-implemented method of claim 2 , further comprising:

determining fourth data representing content of a system response to the past utterance; and

including the fourth data in the third data.

5 . The computer-implemented method of claim 1 , further comprising:

determining third encoded data representing parts-of-speech of at least one of the first utterance or the past utterance,

wherein the machine learning model further processes the third encoded data to determine the model output data.

6 . The computer-implemented method of claim 1 , further comprising:

determining third encoded data representing a device corresponding to the first utterance,

wherein the machine learning model further processes the third encoded data to determine the model output data.

7 . The computer-implemented method of claim 1 , further comprising:

determining weight data based at least in part on the second encoded data,

wherein the machine learning model uses the weight data to determine the model output data.

8 . The computer-implemented method of claim 1 , further comprising:

determining third encoded data representing a topic of at least one of the first utterance or the past utterance,

wherein the machine learning model further processes the third encoded data to determine the model output data.

9 . The computer-implemented method of claim 1 , wherein the first data represents a plurality of speech recognition hypotheses.

10 . A system comprising:

at least one processor; and

at least one memory comprising instructions that, when executed by the at least one processor, cause the system to:

receive input audio data corresponding to a first utterance;

perform speech recognition using the input audio data to determine first data indicating at a least a first speech recognition hypothesis representing the first utterance;

encode the first data to generate first encoded data;

receive second encoded data corresponding to a past utterance;

process the first encoded data and the second encoded data using a machine learning model to determine model output data indicating a second speech recognition hypothesis representing the first utterance, the second speech recognition hypothesis being different from the first speech recognition hypothesis;

process the second speech recognition hypothesis to generate second data indicating an action to be performed in response to the first utterance; and

generate, using the second data, output data responsive to the first utterance.

11 . The system of claim 10 , wherein the at least one memory further comprises instructions that, when executed by the at least one processor, further cause the system to:

receive third data corresponding to the past utterance; and

encode the third data to generate the second encoded data.

12 . The system of claim 11 , wherein the at least one memory further comprises instructions that, when executed by the at least one processor, further cause the system to:

determine fourth data representing content of a user input corresponding to the past utterance; and

include the fourth data in the third data.

13 . The system of claim 11 , wherein the at least one memory further comprises instructions that, when executed by the at least one processor, further cause the system to:

determine fourth data representing content of a system response to the past utterance; and

include the fourth data in the third data.

14 . The system of claim 10 , wherein the at least one memory further comprises instructions that, when executed by the at least one processor, further cause the system to:

determine third encoded data representing parts-of-speech of at least one of the first utterance or the past utterance,

wherein the machine learning model further processes the third encoded data to determine the model output data.

15 . The system of claim 10 , wherein the at least one memory further comprises instructions that, when executed by the at least one processor, further cause the system to:

determine third encoded data representing a device corresponding to the first utterance,

wherein the machine learning model further processes the third encoded data to determine the model output data.

16 . The system of claim 10 , wherein the at least one memory further comprises instructions that, when executed by the at least one processor, further cause the system to:

determine weight data based at least in part on the second encoded data,

wherein the machine learning model uses the weight data to determine the model output data.

17 . The system of claim 10 , wherein the at least one memory further comprises instructions that, when executed by the at least one processor, further cause the system to:

determine third encoded data representing a topic of at least one of the first utterance or the past utterance,

wherein the machine learning model further processes the third encoded data to determine the model output data.

18 . The system of claim 10 , wherein the first data represents a plurality of speech recognition hypotheses.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jun 7, 2021
From: HEDAYATNIA, BEHNAM; RAJU, ANIRUDH; GANDHE, ANKUR; KHATRI, CHANDRA PRAKASH; RASTROW, ARIYA; VENKATESH, ANUSHREE; MANDAL, ARINDAM; GABRIEL, RAEFER CHRISTOPHER; MEHRI, AHMAD SHIKIB
To: AMAZON TECHNOLOGIES, INC.
Reel/Frame 056454/0293 →
Continuity (2)
Continuation 16204670 · Nov 29, 2018
Related Publication 20210312914A1 · Oct 7, 2021
References Cited (20)
US 9691384B1 · Wang · 2017 [cited by examiner]
US 10789943B1 · Lapshina · 2020 [cited by examiner]
US 10860629B1 · Gangadharaiah · 2020 [cited by examiner]
US 11043230B1 · Riding · 2021 [cited by examiner]
US 11056107B2 · Nahamoo · 2021 [cited by examiner]
US 20120084086A1 · Gilbert · 2012 [cited by examiner]
US 20140358537A1 · Gilbert · 2014 [cited by examiner]
US 20160055240A1 · Tur · 2016 [cited by examiner]
US 20160104478A1 · Seo · 2016 [cited by examiner]
US 20160300573A1 · Carbune · 2016 [cited by examiner]
US 20170177716A1 · Perez · 2017 [cited by examiner]
US 20170213551A1 · Ji · 2017 [cited by examiner]
US 20180254035A1 · Kulkarni · 2018 [cited by examiner]
US 20180329957A1 · Frazzingaro · 2018 [cited by examiner]
US 20190035385A1 · Lawson · 2019 [cited by examiner]
US 20190043527A1 · Georges · 2019 [cited by examiner]
US 20190115027A1 · Shah · 2019 [cited by examiner]
US 20190318724A1 · Chao · 2019 [cited by examiner]
US 20200152184A1 · Steedman Henderson · 2020 [cited by examiner]
US 20210350209A1 · Wang · 2021 [cited by examiner]