IP Library › Granted Patent US 12,475,880
Granted Patent B2
US 12,475,880 · App. 18/663,831 · Granted Nov 18, 2025

Non-speech input to speech processing system

Inventor: Travis Grizzel (Snoqualmie, WA)
Assignee: Amazon Technologies, Inc.
G10L15/01G06F3/017G10L13/00G10L15/18G10L15/187G10L15/24G10L2015/088
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,475,880
App. No.
18/663,831
Granted
Nov 18, 2025
Kind
B2
Abstract

A system and method for associating motion data with utterance audio data for use with a speech processing system. A device, such as a wearable device, may be capable of capturing utterance audio data and sending it to a remote server for speech processing, for example for execution of a command represented in the utterance. The device may also capture motion data using motion sensors of the device. The motion data may correspond to gestures, such as head gestures, that may be interpreted by the speech processing system to determine and execute commands. The device may associate the motion data with the audio data so the remote server knows what motion data corresponds to what portion of audio data for purposes of interpreting and executing commands. Metadata sent with the audio data and/or motion data may include association data such as timestamps, session identifiers, message identifiers, etc.

Claims (66)

1 . A wearable glasses device configured to be worn on a head of a user, the wearable glasses device comprising:

at least one sensor;

at least one microphone;

at least one processor; and

at least one memory comprising instructions that, when executed by the at least one processor, cause the wearable glasses device to:

capture, by the at least one microphone, first input audio of an utterance;

determine input audio data representing the first input audio;

receive, from the at least one sensor, sensor data corresponding to a physical attribute of the head of the user;

determine, based at least in part on an operational status of the wearable glasses device and the sensor data, gesture data;

determine time data that indicates receipt of the sensor data with respect to receipt of the first input audio;

determine, based at least in part on the time data, first data indicating the gesture data corresponds to a portion of the input audio data;

cause, based at least in part on the input audio data, the gesture data and the first data, speech processing to be performed using the input audio data; and

present an output representing a result of the speech processing.

2 . The wearable glasses device of claim 1 , wherein:

they at least one sensor comprises a motion sensor; and

the sensor data represents movement of the head.

3 . The wearable glasses device of claim 1 , wherein:

the at least one sensor comprises a camera; and

the sensor data represents a direction of the head.

4 . The wearable glasses device of claim 1 , wherein:

the at least one sensor comprises a camera; and

the sensor data represents movement of the head.

5 . The wearable glasses device of claim 1 , wherein the at least one memory further comprises instructions that, when executed by the at least one processor, cause the wearable glasses device to:

process the sensor data to determine a physical gesture was performed using the head of the user; and

determine, based at least in part on the physical gesture, the gesture data,

wherein the instructions that cause the wearable glasses device to cause speech processing to be performed comprise instructions that, when executed by the at least one processor, cause the wearable glasses device to cause the speech processing to be performed based at least in part on the gesture data.

6 . The wearable glasses device of claim 1 , wherein the at least one memory further comprises instructions that, when executed by the at least one processor, cause the wearable glasses device to:

send the sensor data and the input audio data to at least one second device,

wherein the instructions that cause the wearable glasses device to cause speech processing to be performed comprise instructions that, when executed by the at least one processor, cause the wearable glasses device to cause the at least one second device to perform speech processing using the input audio data and the sensor data.

7 . The wearable glasses device of claim 1 , further comprising an audio output component wherein the instructions that cause the wearable glasses device to present the output comprise instructions that, when executed by the at least one processor, cause the wearable glasses device to present audio using the audio output component, the audio representing the result of the speech processing.

8 . The wearable glasses device of claim 1 , wherein the instructions that cause the wearable glasses device to cause speech processing to be performed comprise instructions that, when executed by the at least one processor, cause the wearable glasses device to perform the speech processing using the input audio data.

9 . The wearable glasses device of claim 1 , wherein the instructions that cause the wearable glasses device to cause speech processing to be performed comprise instructions that, when executed by the at least one processor, cause the wearable glasses device to cause speech processing to include natural language understanding (NLU) processing, wherein the NLU processing is based at least in part on the sensor data.

10 . The wearable glasses device of claim 1 , wherein the at least one memory further comprises instructions that, when executed by the at least one processor, cause the wearable glasses device to:

determine second time data corresponding to second sensor data,

wherein the instructions that cause the wearable glasses device to cause speech processing to be performed comprise instructions that, when executed by the at least one processor, cause the wearable glasses device to cause speech processing to be performed based at least in part on the second time data.

11 . A computer-implemented method for execution by a wearable glasses device, the method comprising:

capturing, by at least one microphone of a wearable glasses device, first input audio of an utterance;

determining input audio data representing the first input audio;

receiving, from at least one sensor of the wearable glasses device, sensor data corresponding to a physical attribute of a head of a user;

determining, based at least in part on an operational status of the wearable glasses device and the sensor data, gesture data;

determining time data that indicates receipt of the sensor data with respect to receipt of the first input audio;

determining, based at least in part on the time data, first data indicating the gesture data corresponds to a portion of the input audio data;

causing, based at least in part on the input audio data, the gesture data and the first data, speech processing to be performed using the input audio data; and

presenting an output representing a result of the speech processing.

12 . The computer-implemented method of claim 11 , wherein:

the at least one sensor comprises a motion sensor; and

the sensor data represents movement of the head.

13 . The computer-implemented method of claim 11 , wherein:

the at least one sensor comprises a camera; and

the sensor data represents a direction of the head.

14 . The computer-implemented method of claim 11 , wherein:

the at least one sensor comprises a camera; and

the sensor data represents movement of the head.

15 . The computer-implemented method of claim 11 , further comprising:

processing the sensor data to determine a physical gesture was performed using the head of the user; and

determining, based at least in part on the physical gesture, the gesture data,

wherein causing the speech processing to be performed comprises causing the speech processing to be performed based at least in part on the gesture data.

16 . The computer-implemented method of claim 11 , further comprising:

sending the sensor data and the input audio data to at least one second device,

wherein causing the speech processing to be performed comprises causing the at least one second device to perform the speech processing.

17 . The computer-implemented method of claim 11 , wherein presenting the output comprises causing an audio output component of the wearable glasses device to present audio representing the result of the speech processing.

18 . The computer-implemented method of claim 11 , wherein causing the speech processing to be performed comprises performing the speech processing by the wearable glasses device.

19 . The computer-implemented method of claim 11 , wherein causing the speech processing to be performed comprises causing natural language understanding (NLU) processing to be performed, wherein the NLU processing is based at least in part on the sensor data.

20 . The computer-implemented method of claim 11 , further comprising:

determining second time data corresponding to second sensor data,

wherein causing the speech processing to be performed comprises causing the speech processing to be performed based at least in part on the second time data.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded May 14, 2024
From: GRIZZEL, TRAVIS
To: AMAZON TECHNOLOGIES, INC.
Reel/Frame 067409/0418 →
Continuity (3)
Continuation 16902992 · Jun 16, 2020
Continuation 15389742 · Dec 23, 2016
Related Publication 20240296829A1 · Sep 5, 2024
References Cited (9)
US 10282057B1 · Binder · 2019 [cited by examiner]
US 20080086754A1 · Chen · 2008 [cited by examiner]
US 20120236025A1 · Jacobsen · 2012 [cited by examiner]
US 20130288753A1 · Jacobsen · 2013 [cited by examiner]
US 20140173440A1 · Dal Mutto · 2014 [cited by examiner]
US 20140371955A1 · Vaghefinazari · 2014 [cited by examiner]
US 20150012426A1 · Purves · 2015 [cited by examiner]
US 20150177841A1 · VanBlon · 2015 [cited by examiner]
US 20150336588A1 · Ebner · 2015 [cited by examiner]