IP Library Granted Patent US 12,334,067
Granted Patent B2
US 12,334,067 · App. 17/883,224 · Granted Jun 17, 2025

Voice trigger based on acoustic space

Inventors: Prateek Murgai (Palo Alto, CA); Ashrith Deshpande (San Jose, CA)
Assignee: Apple Inc.
G10L15/22G10L15/16G10L15/20G10L25/84H04R1/406H04R3/005G10L2015/223
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,334,067
App. No.
17/883,224
Granted
Jun 17, 2025
Kind
B2
Abstract

A plurality of microphone signals can be obtained. In the plurality of microphone signals, speech of a user can be detected. A gaze of a user can be determined based on the plurality of microphone signals. A voice activated response of the computing device can be performed in response to the gaze of the user being directed at the computing device. Other aspects are described and claimed.

Claims (37)

1. A method, performed by a computing device, comprising:

obtaining a plurality of microphone signals generated from a plurality of microphones;

detecting, in the plurality of microphone signals, speech of a user;

determining whether the speech is originating in a shared acoustic space with the computing device based on the plurality of microphone signals;

determining an acoustic gaze of the user based on the plurality of microphone signals, wherein the acoustic gaze comprises a direction in which a head and mouth of the user is pointing with respect to the computing device; and

triggering a voice activated response of the computing device in response to a determination that the speech originates in the shared acoustic space with the computing device and the acoustic gaze of the user being directed at the computing device.

2. The method of claim 1 , wherein determining the acoustic gaze of the user includes estimating a direct to reverberant ratio (DRR) using the plurality of microphone signals.

3. The method of claim 2 , wherein the acoustic gaze of the user is determined to be directed at the computing device when the DRR satisfies a threshold or when the DRR is higher than a second DRR that is determined from microphone signals of a second device.

4. The method of claim 1 , wherein determining the acoustic gaze of the user includes generating a plurality of acoustic pickup beams from the plurality of microphone signals and measuring direct and reverberant acoustic sound in the plurality of acoustic pickup beams.

5. The method of claim 1 , wherein determining, whether the speech is originating in the shared acoustic space with the computing device is performed based on a trained neural network.

6. The method of claim 5 , wherein the trained neural network is trained to output a confidence score that indicates whether the speech is originating in the shared acoustic space with the computing device.

7. The method of claim 5 , wherein the computing device is one of a plurality of computing devices that sense the speech, and a selected one of the plurality of computing devices is triggered in response to:

determining that the speech originates in the shared acoustic space with the selected one of the plurality of computing devices, and

determining that the acoustic gaze of the user is directed at the selected one of the plurality of computing devices.

8. The method of claim 1 , further comprising performing blind room estimation using at least one of the microphone signals to determine reverberation time of an acoustic space of the computing device, wherein the reverberation time is used to track the acoustic space of the computing device.

9. The method of claim 1 , wherein the voice activated response includes at least one of: a wake-up of the computing device, processing the speech to detect a voice command, responding to a voice command in the speech, or determining an identity of the user based on the speech.

10. A processor of a computing device, the processor configured to:

obtain a plurality of microphone signals generated from a plurality of microphones;

detect, in the plurality of microphone signals, speech of a user;

determine, with a trained neural network, whether the speech originates in a shared acoustic space with the computing device based on the plurality of microphone signals;

determine an acoustic gaze of the user based on the plurality of microphone signals, wherein the acoustic gaze comprises a direction in which a head and mouth of the user is pointing with respect to the computing device; and

triggering a voice activated response of the computing device in response to a determination that the speech originates in the shared acoustic space with the computing device and the acoustic gaze of the user being directed at the computing device.

11. The processor of claim 10 determines whether the speech originates in the shared acoustic space as output of a trained neural network responsive to input based on the plurality of microphone signals, wherein the trained neural network is trained to output a confidence score indicating whether the speech is originating in the shared acoustic space with the computing device.

12. The processor of claim 10 , wherein the voice activated response of the computing device is not triggered when the speech does not originate in the shared acoustic space with the computing device.

13. The processor of claim 10 , wherein the acoustic gaze of the user is determined by estimating a direct to reverberant ratio (DRR) using the plurality of microphone signals.

14. The processor of claim 13 , wherein the acoustic gaze of the user is determined to be directed at the computing device when the DRR satisfies a threshold or when the DRR is higher than a second DRR that is determined from microphone signals of a second device.

15. The processor of claim 13 , wherein estimating the DRR includes generating a plurality of acoustic pickup beams from the plurality of microphone signals and measuring direct and reverberant acoustic sound for each of the plurality of acoustic pickup beams.

16. The processor of claim 10 , wherein the computing device is one of a plurality of computing devices, and a selected one of the plurality of computing devices is triggered based on:

a determination that the speech originates in the shared acoustic space with the selected one of the plurality of computing devices, and

a determination that the acoustic gaze of the user is directed at the selected one of the plurality of computing devices.

17. The processor of claim 10 , wherein the voice activated response includes at least one of: a wake-up of the computing device, detecting a voice command in the speech, responding to a voice command in the speech, and determining an identity of the user based on the speech.

18. A method, performed by a computing device, comprising:

obtaining, from each of one or more computing devices, an indication of an acoustic gaze of a user relative to a respective one of the one or more computing devices and an indication of whether speech of the user originates in a shared acoustic space with the respective one of the one or more computing devices, wherein the acoustic gaze comprises a direction in which a head and mouth of the user is pointing with respect to the respective one of the one or more computing devices; and

selecting one of the one or more computing devices to trigger a voice activated response based on:

an indication from the selected one of the one or more computing devices that the acoustic gaze of the user is directed at the selected one of the one or more computing devices, and

an indication from the selected one of the one or more computing devices that the speech of the user originates in the shared acoustic space with the selected one of the one or more computing devices.

19. The method of claim 18 , further comprising obtaining, a reverberation time of a respective acoustic space from each of the one or more computing devices to track the respective acoustic space of each of the one or more computing devices.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Aug 8, 2022
From: MURGAI, PRATEEK; DESHPANDE, ASHRITH
To: APPLE INC.
Reel/Frame 060746/0772 →
Continuity (2)
Provisional Application 63239567 · Sep 1, 2021
Related Publication 20230062634A1 · Mar 2, 2023
References Cited (26)
US 7191090B1 · Cunningham · 2007 [cited by applicant]
US 8885442B2 · Agevik et al. · 2014 [cited by applicant]
US 9449613B2 · Peters et al. · 2016 [cited by applicant]
US 9734845B1 · Liu et al. · 2017 [cited by applicant]
US 10395667B2 · Ebenezer · 2019 [cited by applicant]
US 10410651B2 · Lou · 2019 [cited by examiner]
US 11132991B2 · Park et al. · 2021 [cited by applicant]
US 11289086B2 · Burton · 2022 [cited by examiner]
US 11482217B2 · Golikov · 2022 [cited by examiner]
US 11508378B2 · Kim et al. · 2022 [cited by applicant]
US 11694685B2 · Carbune · 2023 [cited by examiner]
US 20120263020A1 · Taylor et al. · 2012 [cited by applicant]
US 20150333819A1 · Candelore · 2015 [cited by applicant]
US 20160234595A1 · Goran et al. · 2016 [cited by applicant]
US 20180294000A1 · Steele · 2018 [cited by examiner]
US 20180336905A1 · Kim et al. · 2018 [cited by applicant]
US 20200312315A1 · Li et al. · 2020 [cited by applicant]
US 20210281965A1 · Malik et al. · 2021 [cited by applicant]
KR 20190096861A · 2019 [cited by applicant]
KR 20200009035A · 2020 [cited by applicant]
KR 20200052804A · 2020 [cited by applicant]
WO 2017138934 · 2017 [cited by applicant]
WO 2018112643 · 2018 [cited by applicant]
Examination Report under Section 18(3) for UK Application No. GB2211193.4 mailed Jan. 31, 2023, 2 pages. [cited by applicant]
Search Report under Section 17 for UK Application No. GB2211193.4 mailed Jan. 31, 2023, 1 page. [cited by applicant]
Bryan, Nicholas J. , Impulse Response Data Augmentation and Deep Neural Networks for Blind Room Acoustic Parameter Estimation, Acoustics, Speech and Signal Processing (ICASSP), ICASSP 2020—2020 IEEE International Confer… [cited by applicant]