IP Library Granted Patent US 12,531,064
Granted Patent B1
US 12,531,064 · App. 18/620,703 · Granted Jan 20, 2026

Audio-based user engagement detection

Inventor: Wai Chung Chu (San Jose, CA)
Assignee: Amazon Technologies, Inc.
G10L15/22G10L25/21H04R3/005
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,531,064
App. No.
18/620,703
Granted
Jan 20, 2026
Kind
B1
Abstract

A system can operate a speech-controlled device to perform user engagement detection (UED) processing to detect when speech represented in audio data is directed to the device. For example, the device may extract audio features from the audio data and process these audio features using a classifier to estimate an orientation of the user's head, which may be used as a proxy for user engagement. Thus, if the head orientation is within an engagement zone (which varies based on distance to the user), the device may determine that the user is engaged with the device and perform language processing on input speech. In contrast, if the head orientation is outside of the engagement zone, the device may determine that the user is not engaged and ignore the input speech. To enable additional functionality, the classifier may optionally output a coarse estimate of the head orientation along with the UED determination.

Claims (73)

1 . An electronic device comprising:

a plurality of microphones;

a loudspeaker;

one or more processors; and

one or more non-transitory computer readable media storing processor executable instructions which, when executed using the one or more processors, cause the electronic device to perform operations comprising:

determining, based on first audio data corresponding to sound captured by at least two microphones of the plurality of microphones, first sound source localization data comprising:

first cell data indicating a first three-dimensional vector and a first power value associated with the first three-dimensional vector, and

second cell data indicating a second three-dimensional vector and a second power value associated with the second three-dimensional vector;

determining, based on the first sound source localization data:

a first metric representing direct-to-reverberant ratio information, and

a second metric representing direction variance information;

based on the first metric and the second metric, using a first machine learning model to determine model output estimating user engagement; and

executing, based on the model output, a first operation.

2 . The electronic device of claim 1 , wherein:

the first sound source localization data comprises data indicating, for each respective cell of a plurality of cells, a respective three-dimensional vector and a respective power value associated with the respective three-dimensional vector, the plurality of cells including a first cell corresponding to the first cell data and a second cell corresponding to the second cell data; and

the one or more computer readable media store further processor executable instructions which, when executed using the one or more processors, cause the electronic device to perform operations comprising:

determining, based on the first sound source localization data, that the first power value is a peak power value among the plurality of cells,

determining, based on the first sound source localization data, a mean power value for the plurality of cells, and

determining a peak-to-mean ratio value based on the first power value and the mean power value,

wherein the peak-to-mean ratio value is the first metric.

3 . The electronic device of claim 1 , wherein:

the first sound source localization data comprises data indicating, for each respective cell of a plurality of cells, a respective three-dimensional vector and a respective power value associated with the respective three-dimensional vector, the plurality of cells including a first cell corresponding to the first cell data and a second cell corresponding to the second cell data; and

the one or more computer readable media store further processor executable instructions which, when executed using the one or more processors, cause the electronic device to perform operations comprising:

determining, based on the first sound source localization data, first data indicating a power-weighted mean direction vector,

wherein the second metric is determined based on the first data.

4 . The electronic device of claim 1 , wherein the first sound source localization data represents sound source localization data for a first frame of audio data, and wherein the one or more computer readable media store further processor executable instructions which, when executed using the one or more processors, cause the electronic device to perform operations comprising:

determining, based on first audio data corresponding to sound captured by a second set of microphones of the plurality of microphones, second sound source localization data comprising data indicating, for each respective cell of a plurality of cells, a respective three-dimensional vector and a respective power value associated with the respective three-dimensional vector;

wherein the first metric is determined based on the first sound source localization data and the second sound source localization data; and

wherein the second metric is determined based on the first sound source localization data and the second sound source localization data.

5 . The electronic device of claim 1 , wherein the first sound source localization data represents sound source localization data for a first frame of audio data, and wherein the one or more computer readable media store further processor executable instructions which, when executed using the one or more processors, cause the electronic device to perform operations comprising:

determining, based on first audio data corresponding to sound captured by a second set of microphones of the plurality of microphones, second sound source localization data comprising data indicating, for each respective cell of a plurality of cells, a respective three-dimensional vector and a respective power value associated with the respective three-dimensional vector; and

determining, based on the second sound source localization data:

a third metric representing direct-to-reverberant ratio information, and

a fourth metric representing direction variance information,

wherein the determining of the model output using the first machine learning model is based on the third metric and the fourth metric.

6 . The electronic device of claim 1 , wherein the model output indicates that a user is engaging with the electronic device.

7 . The electronic device of claim 1 , wherein the model output indicates that a user is not engaging with the electronic device.

8 . The electronic device of claim 1 , wherein the first operation comprises turning off a light of the electronic device.

9 . The electronic device of claim 1 , wherein the model output comprises an engagement value indicating an engagement level and a confidence value indicating a confidence level in the engagement value.

10 . The electronic device of claim 1 , wherein the first operation comprises sending at least a subset of the first audio data to a remote system, and wherein the one or more computer readable media store further processor executable instructions which, when executed using the one or more processors, cause the electronic device to perform operations comprising:

outputting, using a speaker of the electronic device and based on first response data received from the remote system, audio representing speech responding to user speech.

11 . The electronic device of claim 1 , wherein the model output indicates that a user is not engaging with the electronic device, and wherein the first operation comprises powering down a component of the electronic device.

12 . The electronic device of claim 1 , wherein the first operation comprises transitioning to a different state.

13 . The electronic device of claim 1 , wherein the model output indicates that a user is engaging with the electronic device, and wherein the first operation comprises transitioning to an awake or active state.

14 . The electronic device of claim 1 , wherein the model output indicates that a user is not engaging with the electronic device, and wherein the first operation comprises transitioning to an inactive or sleep state.

15 . The electronic device of claim 1 , wherein the model output is a numerical value indicating an engagement level on a scale of engagement.

16 . An electronic device comprising:

a plurality of microphones;

a loudspeaker;

one or more processors; and

one or more non-transitory computer readable media storing processor executable instructions which, when executed using the one or more processors, cause the electronic device to perform operations comprising:

determining, based on first audio data corresponding to sound captured by at least two microphones of the plurality of microphones, first sound source localization data indicating, for each respective cell of a plurality of cells:

a respective three-dimensional vector, and

a respective power value associated with the respective three-dimensional vector;

determining, based on the first sound source localization data, a first power value that is a peak power value among the plurality of cells;

determining, based on the first sound source localization data, a mean power value for the plurality of cells;

determining a peak-to-mean ratio value based on the first power value and the mean power value;

determining, using a first machine learning model and based on the peak-to-mean ratio value, model output estimating user engagement; and

executing, based on the model output, a first operation.

17 . The electronic device of claim 16 , wherein the first operation comprises performing speech transcription on at least a portion of the first audio data.

18 . The electronic device of claim 16 , wherein the first operation comprises generating, based on at least a portion of the first audio data and using a second machine learning model, transcription data.

19 . An electronic device comprising:

a plurality of microphones;

a loudspeaker;

one or more processors; and

one or more non-transitory computer readable media storing processor executable instructions which, when executed using the one or more processors, cause the electronic device to perform operations comprising:

determining, based on first audio data corresponding to sound captured by at least two microphones of the plurality of microphones, first sound source localization data indicating, for each respective cell of a plurality of cells:

a respective three-dimensional vector, and

a respective power value associated with the respective three-dimensional vector;

determining, based on the first sound source localization data, a power-weighted mean direction vector representing direction variance information;

determining, using a first machine learning model and based on the power-weighted mean direction vector, model output estimating user engagement; and

executing, based on the model output, a first operation.

20 . The electronic device of claim 19 , wherein the first operation comprises performing speech transcription on at least a portion of the first audio data.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Oct 21, 2025
From: CHU, WAI CHUNG
To: AMAZON TECHNOLOGIES, INC.
Reel/Frame 072631/0009 →
References Cited (35)
US 5485524A · Kuusama · 1996 [cited by examiner]
US 7197456B2 · Haverinen · 2007 [cited by examiner]
US 8213598B2 · Bendersky · 2012 [cited by examiner]
US 8271279B2 · Hetherington · 2012 [cited by examiner]
US 8750491B2 · Prakash · 2014 [cited by examiner]
US 8798992B2 · Gay · 2014 [cited by examiner]
US 8914282B2 · Konchitsky · 2014 [cited by examiner]
US 9043203B2 · Rettelbach · 2015 [cited by examiner]
US 9088328B2 · Gunzelmann · 2015 [cited by examiner]
US 9202463B2 · Newman · 2015 [cited by examiner]
US 9613612B2 · Perkmann · 2017 [cited by examiner]
US 9640179B1 · Hart · 2017 [cited by examiner]
US 9653060B1 · Hilmes · 2017 [cited by examiner]
US 9659555B1 · Hilmes · 2017 [cited by examiner]
US 9792897B1 · Kaskari · 2017 [cited by examiner]
US 9818425B1 · Ayrapetian · 2017 [cited by examiner]
US 9959886B2 · Anhari · 2018 [cited by examiner]
US 10134425B1 · Johnson, Jr. · 2018 [cited by examiner]
US 10192567B1 · Kamdar · 2019 [cited by examiner]
US 11950062B1 · Chu et al. · 2024 [cited by applicant]
US 20090265169A1 · Dyba · 2009 [cited by examiner]
US 20130332175A1 · Setiawan · 2013 [cited by examiner]
US 20140214676A1 · Bukai · 2014 [cited by examiner]
US 20150302845A1 · Nakano · 2015 [cited by examiner]
US 20180357995A1 · Lee · 2018 [cited by examiner]
US 20190108837A1 · Christoph · 2019 [cited by examiner]
US 20190222943A1 · Andersen · 2019 [cited by examiner]
US 20190259381A1 · Ebenezer · 2019 [cited by examiner]
US 20200058320A1 · Liu · 2020 [cited by examiner]
US 20200243061A1 · Sun · 2020 [cited by examiner]
US 20220246161A1 · Verbeke · 2022 [cited by examiner]
Amit Chhetri, et al. “Multichannel Audio Front-end for Far-field Automatice Speech Recognition,” 26th European Signal Processing Conference (EUSIPCO), 2018, 5 pages. https://www.eurasip.org/Proceedings/Eusipco/Eusipco20… [cited by applicant]
Karan Ahuja, et al. “Direction-of-Voice (DoV) Estimation for Intuitive Speech Interaction with Smart Device Ecosystems,” UIST Oct. 2020, Virtual Event, pp. 1121-1131. Retrieved from https://dl.acm.org/doi/10.1145/337933… [cited by applicant]
Mark Barnard and Wenwu Wang. “Audio head pose estimation using the direct to reverberant speech ratio,” Speech Communication, vol. 85, 2016, pp. 98-108, https://dx.doi.org/10.1016/j.specom.2016.09.005. [cited by applicant]
Alberto Abad, et al. “Audio-based approaches to head orientation estimation in a smart-room,” Conference paper at Interspeech 2007, Aug. 2007, pp. 590-593. Retrieved at https://www.researchgate.net/publication/221480124. [cited by applicant]