IP Library › Granted Patent US 12,400,663
Granted Patent B2
US 12,400,663 · App. 18/425,465 · Granted Aug 26, 2025

Speech interface device with caching component

Inventor: Stanislaw Ignacy Pasko (Gdansk, PL)
Assignee: Amazon Technologies, Inc.
G10L15/30G10L15/18H04L67/5683
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,400,663
App. No.
18/425,465
Granted
Aug 26, 2025
Kind
B2
Abstract

A speech interface device is configured to receive response data from a remote speech processing system for responding to user speech. This response data may be enhanced with information such as a remote ASR result(s) and a remote NLU result(s). The response data from the remote speech processing system may include one or more cacheable status indicators associated with the NLU result(s) and/or remote directive data, which indicate whether the remote NLU result(s) and/or the remote directive data are individually cacheable. A caching component of the speech interface device allows for caching at least some of this cacheable remote speech processing information, and using the cached information locally on the speech interface device when responding to user speech in the future. This allows for responding to user speech, even when the speech interface device is unable to communicate with a remote speech processing system over a wide area network.

Claims (42)

1. A method comprising:

generating, by a device, audio data that represents user speech;

performing, by the device, speech processing using the audio data to generate a first result;

determining, by the device that a remote system is unavailable for processing the first result;

determining, by the device and while the remote system is unavailable, that the first result matches with a second result previously stored in a memory of the first device;

determining, by the device, that command information associated with the second result is stored in the memory, the command information being previously received from the remote system and being based on previous interactions between the device and the remote system; and

performing, by the device, an action specified in the command information, the action being associated with the second result.

2. The method of claim 1 , further comprising:

generating, by the device, second audio data that represents second user speech;

performing the speech processing using the second audio data to determine the action;

retrieving the second result from the memory of the device, and

performing, by the device and using the second result, the action.

3. The method of claim 1 , wherein performing the speech processing comprises performing automatic speech recognition (ASR) processing on the audio data to generate text data.

4. The method of claim 1 , wherein the second result includes a model trained by one or more remote components, and wherein performing the speech processing comprises performing the speech processing on the audio data using the model.

5. The method of claim 1 , wherein the second result is used to identify a second device collocated in an environment with the device, and further comprising:

sending a command to the second device, the command instructing an action to be performed at the second device.

6. The method of claim 1 , wherein each of the previous interactions comprises an interaction where a previous user speech is spoken in an environment of the device and a microphone of the device generates respective audio data representing the previous user speech.

7. The method of claim 1 , wherein the second result is associated with a previous user speech that was detected by a second device.

8. The method of claim 7 , wherein the second device is collocated in an environment with the device.

9. A device comprising:

one or more processors; and

memory storing computer-executable instructions that, when executed by the one or more processors, cause the device to:

generate audio data that represents user speech;

perform speech processing using the audio data to generate a first result;

determine that a remote system is unavailable for processing the first result;

determine that the first result matches a second result stored in a memory of the device;

determine that command information associated with the second result is stored in the memory, the command information being previously received from the remote system and being based on previous interactions between the device and the remote system; and

perform an action specified in the command information, the action being associated with the second result.

10. The device of claim 9 , wherein the instructions, when executed by the one or more processors, further cause the device to:

generate second audio data that represents second user speech;

perform the speech processing using the second audio data to determine the action;

retrieve the second result from the memory of the device, and

perform, using the second result, the action.

11. The device of claim 9 , wherein performing the speech processing comprises performing automatic speech recognition (ASR) processing on the audio data to generate text data.

12. The device of claim 9 , wherein the second result includes a model trained by one or more remote components, and wherein performing the speech processing comprises performing the speech processing on the audio data using the model.

13. The device of claim 9 , wherein the second result is used to identify a second device collocated in an environment with the device, and wherein the instructions, when executed by the one or more processors, further cause the device to:

send a command to the second device, the command instructing an action to be performed at the second device.

14. The device of claim 9 , wherein each of the previous interactions comprises an interaction where a previous user speech is spoken in an environment of the device and a microphone of the device generates respective audio data representing the previous user speech.

15. The device of claim 9 , further comprising a microphone, and wherein each of the previous interactions comprises an interaction where a user spoke in an environment of the device and the microphone generated respective audio data representing the user speech.

16. The device of claim 9 , wherein the second result is associated with a previous user speech that was detected by a second device.

17. The device of claim 16 , wherein the second device is collocated in an environment with the device.

18. The device of claim 9 , further comprising a microphone, and wherein the microphone generated the audio data that represents the user speech.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jan 29, 2024
From: PASKO, STANISLAW IGNACY
To: AMAZON TECHNOLOGIES, INC.
Reel/Frame 066281/0140 →
Continuity (4)
Continuation 17902519 · Sep 2, 2022
Continuation 17018279 · Sep 11, 2020
Continuation 15934761 · Mar 23, 2018
Related Publication 20240249725A1 · Jul 25, 2024
References Cited (97)
US 6204763B1 · Sone · 2001 [cited by applicant]
US 6226749B1 · Carloganu · 2001 [cited by examiner]
US 6408272B1 · White et al. · 2002 [cited by applicant]
US 6836760B1 · Bellegarda · 2004 [cited by examiner]
US 8364694B2 · Volkert · 2013 [cited by applicant]
US 8515736B1 · Duta · 2013 [cited by examiner]
US 8761373B1 · Raghavan · 2014 [cited by examiner]
US 8838434B1 · Liu · 2014 [cited by examiner]
US 8983840B2 · Deshmukh · 2015 [cited by examiner]
US 9131369B2 · Ganong, III · 2015 [cited by examiner]
US 9355110B1 · Chi · 2016 [cited by applicant]
US 9405832B2 · Edwards · 2016 [cited by examiner]
US 9484021B1 · Mairesse · 2016 [cited by examiner]
US 9514747B1 · Bisani et al. · 2016 [cited by applicant]
US 9558735B2 · Roberts et al. · 2017 [cited by applicant]
US 9607617B2 · Hebert · 2017 [cited by examiner]
US 9619459B2 · Hebert · 2017 [cited by examiner]
US 9966065B2 · Gruber et al. · 2018 [cited by applicant]
US 10018977B2 · Cipollo et al. · 2018 [cited by applicant]
US 10388277B1 · Ghosh · 2019 [cited by examiner]
US 10515637B1 · Devries et al. · 2019 [cited by applicant]
US 10521189B1 · Ryabov et al. · 2019 [cited by applicant]
US 10629186B1 · Slifka · 2020 [cited by applicant]
US 20040049389A1 · Marko et al. · 2004 [cited by applicant]
US 20060025995A1 · Erhart · 2006 [cited by examiner]
US 20060149544A1 · Hakkani-Tur · 2006 [cited by examiner]
US 20080059188A1 · Konopka et al. · 2008 [cited by applicant]
US 20080189390A1 · Heller et al. · 2008 [cited by applicant]
US 20090030697A1 · Cerra et al. · 2009 [cited by applicant]
US 20090030698A1 · Cerra et al. · 2009 [cited by applicant]
US 20100088100A1 · Lindahl · 2010 [cited by applicant]
US 20100169075A1 · Raffa et al. · 2010 [cited by applicant]
US 20100268536A1 · Suendermann · 2010 [cited by examiner]
US 20100332234A1 · Agapi · 2010 [cited by examiner]
US 20110038613A1 · Buchheit · 2011 [cited by applicant]
US 20110066634A1 · Phillips et al. · 2011 [cited by applicant]
US 20130007208A1 · Tsui et al. · 2013 [cited by applicant]
US 20130085586A1 · Parekh · 2013 [cited by applicant]
US 20130151250A1 · VanBlon · 2013 [cited by examiner]
US 20130159000A1 · Ju · 2013 [cited by examiner]
US 20130173765A1 · Korbecki · 2013 [cited by applicant]
US 20130326353A1 · Singhal · 2013 [cited by examiner]
US 20140019573A1 · Swift · 2014 [cited by applicant]
US 20140039899A1 · Cross, Jr. · 2014 [cited by examiner]
US 20140046876A1 · Zhang et al. · 2014 [cited by applicant]
US 20140058732A1 · Labsky · 2014 [cited by examiner]
US 20140207442A1 · Ganong, III · 2014 [cited by examiner]
US 20140274203A1 · Ganong, III et al. · 2014 [cited by applicant]
US 20150012271A1 · Peng · 2015 [cited by examiner]
US 20150120288A1 · Thomson · 2015 [cited by examiner]
US 20150120296A1 · Stern et al. · 2015 [cited by applicant]
US 20150134334A1 · Sachidanandam et al. · 2015 [cited by applicant]
US 20150199967A1 · Reddy et al. · 2015 [cited by applicant]
US 20150279352A1 · Willett · 2015 [cited by examiner]
US 20150348548A1 · Piernot et al. · 2015 [cited by applicant]
US 20160012819A1 · Willett · 2016 [cited by examiner]
US 20160098998A1 · Wang et al. · 2016 [cited by applicant]
US 20160103652A1 · Kuniansky · 2016 [cited by applicant]
US 20160342383A1 · Parekh · 2016 [cited by applicant]
US 20160378747A1 · Orr et al. · 2016 [cited by applicant]
US 20160379626A1 · Deisher · 2016 [cited by examiner]
US 20170097618A1 · Cipollo et al. · 2017 [cited by applicant]
US 20170177716A1 · Perez · 2017 [cited by examiner]
US 20170213546A1 · Gilbert · 2017 [cited by examiner]
US 20170236512A1 · Williams · 2017 [cited by examiner]
US 20170263253A1 · Thomson et al. · 2017 [cited by applicant]
US 20170278511A1 · Willett · 2017 [cited by examiner]
US 20170278514A1 · Mathias · 2017 [cited by examiner]
US 20170294184A1 · Bradley · 2017 [cited by examiner]
US 20180018959A1 · Des Jardins et al. · 2018 [cited by applicant]
US 20180060326A1 · Kuo · 2018 [cited by examiner]
US 20180061403A1 · Devaraj · 2018 [cited by examiner]
US 20180061404A1 · Devaraj · 2018 [cited by examiner]
US 20180197545A1 · Willett · 2018 [cited by examiner]
US 20180211663A1 · Shin · 2018 [cited by examiner]
US 20180211668A1 · Willett · 2018 [cited by examiner]
US 20180233141A1 · Solomon et al. · 2018 [cited by applicant]
US 20180247065A1 · Rhee · 2018 [cited by examiner]
US 20180268818A1 · Schoenmackers et al. · 2018 [cited by applicant]
US 20180294001A1 · Kayama · 2018 [cited by applicant]
US 20180314689A1 · Wang · 2018 [cited by examiner]
US 20180330728A1 · Gruenstein · 2018 [cited by examiner]
US 20190027147A1 · Diamant et al. · 2019 [cited by applicant]
US 20190042539A1 · Kalsi et al. · 2019 [cited by applicant]
US 20190043509A1 · Suppappola · 2019 [cited by examiner]
US 20190043529A1 · Muchlinski · 2019 [cited by examiner]
US 20190057693A1 · Fry · 2019 [cited by examiner]
US 20190103101A1 · Danila · 2019 [cited by examiner]
US 20190295552A1 · Pasko · 2019 [cited by examiner]
US 20190371307A1 · Zhao et al. · 2019 [cited by applicant]
US 20190392836A1 · Kang et al. · 2019 [cited by applicant]
US 20200051547A1 · Shanmugam et al. · 2020 [cited by applicant]
US 20200410996A1 · Brandel et al. · 2020 [cited by applicant]
Office Action for U.S. Appl. No. 15/934,761, mailed on Feb. 6, 2020, Pasko, “Speech Interface Device With Caching Component”, 11 Pages. [cited by applicant]
Office Action for U.S. Appl. No. 17/018,279, mailed 12/20/202, Pasko, “Speech Interface Device With Caching Component ”, 24 pages. [cited by applicant]
Office Action for U.S. Appl. No. 17/018,279, mailed Aug. 18, 2021, Pasko, “Speech Interface Device With Caching Component ”, 20 pages. [cited by applicant]
Office Action for U.S. Appl. No. 17/902,519, mailed on May 22, 2023, Inventor #1Stanislaw Ignacy Pasko, “Speech Interface Device With Caching Component ,” 8 pages. [cited by applicant]