IP Library › Granted Patent US 12,217,736
Granted Patent B2
US 12,217,736 · App. 18/367,859 · Granted Feb 4, 2025

Simultaneous acoustic event detection across multiple assistant devices

Inventors: Matthew Sharifi (Kilchberg, CH); Victor Carbune (Zurich, CH)
Assignee: GOOGLE LLC
G10L15/01G01S3/8006G10L15/08G10L15/32H04R29/006G10L2015/088
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,217,736
App. No.
18/367,859
Granted
Feb 4, 2025
Kind
B2
Abstract

Implementations can detect respective audio data that captures an acoustic event at multiple assistant devices in an ecosystem that includes a plurality of assistant devices, process the respective audio data locally at each of the multiple assistant devices to generate respective measures that are associated with the acoustic event using respective event detection models, process the respective measures to determine whether the detected acoustic event is an actual acoustic event, and cause an action associated with the actional acoustic event to be performed in response to determining that the detected acoustic event is the actual acoustic event. In some implementations, the multiple assistant devices that detected the respective audio data are anticipated to detect the respective audio data that captures the actual acoustic event based on a plurality of historical acoustic events being detected at each of the multiple assistant devices.

Claims (52)

1. A method implemented by one or more processors, the method comprising:

detecting, via one or more microphones of an assistant device located in an ecosystem that includes a plurality of assistant devices, audio data that captures an acoustic event, wherein the acoustic event comprises a hotword detection event;

processing, using an event detection model that is stored locally at the assistant device, the audio data that captures the acoustic event to generate a measure associated with the acoustic event, wherein the event detection model that is stored locally at the assistant device comprises a hotword detection model that is trained to detect whether a particular word or phrase is captured in the audio data;

in response to detecting the audio data via the one or more microphones of the assistant device:

anticipating detection of additional audio data via one or more additional microphones of an additional assistant device based on a plurality of historical acoustic events being detected at both the assistant device and the additional assistant device, the additional assistant device being in addition to the assistant device, and the additional assistant device being co-located in the ecosystem with the assistant device;

detecting, via the one or more additional microphones of the additional assistant device located in the ecosystem, the additional audio data that also captures the acoustic event;

processing, using an additional event detection model that is stored locally at the additional assistant device, the additional audio data that captures the acoustic event to generate an additional measure associated with the acoustic event, wherein the additional event detection model that is stored locally at the additional assistant device comprises an additional hotword detection model that is trained to detect whether the particular word or phrase is captured in the additional audio data;

determining, based on the measure satisfying a threshold indicating that the particular word or phrase is captured in the audio data and based on the additional measure satisfying the threshold indicating that the particular word or phrase is captured in the additional audio data, that the acoustic event detected by at least both the assistant device and the additional assistant device corresponds to an occurrence of an actual acoustic event; and

in response to determining that the acoustic event corresponds to an occurrence of the actual acoustic event, causing one or more components of an automated assistant to be activated at one or more of: the assistant device, the additional assistant device, or a further additional assistant device.

2. The method of claim 1 , wherein the measure associated with the acoustic event comprises a corresponding confidence level corresponding to whether the audio data captures the particular word or phrase, and wherein the additional measure associated with the acoustic event comprises a corresponding additional confidence level corresponding to whether the additional audio data captures the particular word or phrase.

3. The method of claim 2 , wherein determining that the acoustic event corresponds to the occurrence of the actual acoustic event comprises determining the particular word or phrase is captured in both the audio data and the additional audio data based on the corresponding confidence level and the corresponding additional confidence level satisfying the threshold.

4. The method of claim 1 , wherein the hotword detection model that is stored locally at the assistant device is a distinct hotword model that is distinct from the additional hotword detection model that is stored locally at the additional assistant device.

5. The method of claim 1 , wherein the assistant device, the additional assistant device, and the further additional client device are co-located in an ecosystem of devices.

6. The method of claim 1 , wherein the audio data temporally corresponds to the additional audio data.

7. The method of claim 6 , wherein determining that the acoustic event detected by at least both the assistant device and the additional assistant device corresponds to an occurrence of an actual acoustic event based on the measure satisfying the threshold indicating that the particular word or phrase is captured in the audio data and based on the additional measure satisfying the threshold indicating that the particular word or phrase is captured in the audio data is in response to determining that the audio data temporally corresponds to the additional audio data.

8. The method of claim 6 , wherein determining that the acoustic event detected by both the assistant device and the additional assistant device corresponds to the occurrence of the actual acoustic event is in response to determining that a timestamp associated with the audio data temporally corresponds to an additional timestamp associated with the additional audio data.

9. The method of claim 1 , further comprising:

transmitting, by the assistant device, and to a remote system, the audio data; and

transmitting, by the additional assistant device, and to the remote system, the additional audio data,

wherein determining whether the acoustic event detected by both the assistant device and the additional assistant device corresponds to the occurrence of the actual acoustic event is by the remote system.

10. The method of claim 1 , wherein the one or more components of the automated assistant comprise one or more of:

a speech-to-text component; or

a natural language understanding component.

11. A system comprising:

one or more processors; and

memory storing instructions that, when executed, the one or more processors are operable to:

detect, via one or more microphones of an assistant device located in an ecosystem that includes a plurality of assistant devices, audio data that captures an acoustic event, wherein the acoustic event comprises a hotword detection event;

process, using an event detection model that is stored locally at the assistant device, the audio data that captures the acoustic event to generate a measure associated with the acoustic event, wherein the event detection model that is stored locally at the assistant device comprises a hotword detection model that is trained to detect whether a particular word or phrase is captured in the audio data;

in response to detecting the audio data via the one or more microphones of the assistant device:

anticipate detection of additional audio data via one or more additional microphones of an additional assistant device based on a plurality of historical acoustic events being detected at both the assistant device and the additional assistant device, the additional assistant device being in addition to the assistant device, and the additional assistant device being co-located in the ecosystem with the assistant device;

detect, via the one or more additional microphones of the additional assistant device located in the ecosystem, the additional audio data that also captures the acoustic event;

process, using an additional event detection model that is stored locally at the additional assistant device, the additional audio data that captures the acoustic event to generate an additional measure associated with the acoustic event, wherein the additional event detection model that is stored locally at the additional assistant device comprises an additional hotword detection model that is trained to detect whether the particular word or phrase is captured in the additional audio data;

determine, based on the measure satisfying a threshold indicating that the particular word or phrase is captured in the audio data and based on the additional measure satisfying the threshold indicating that the particular word or phrase is captured in the additional audio data, that the acoustic event detected by at least both the assistant device and the additional assistant device corresponds to an occurrence of an actual acoustic event; and

in response to determining that the acoustic event corresponds to an occurrence of the actual acoustic event, cause one or more components of an automated assistant to be activated at one or more of: the assistant device, the additional assistant device, or a further additional assistant device.

12. A non-transitory computer-readable storage medium storing instructions that, when executed, cause one or more processors to perform operations, the operations comprising:

detecting, via one or more microphones of an assistant device located in an ecosystem that includes a plurality of assistant devices, audio data that captures an acoustic event, wherein the acoustic event comprises a hotword detection event;

processing, using an event detection model that is stored locally at the assistant device, the audio data that captures the acoustic event to generate a measure associated with the acoustic event, wherein the event detection model that is stored locally at the assistant device comprises a hotword detection model that is trained to detect whether a particular word or phrase is captured in the audio data;

in response to detecting the audio data via the one or more microphones of the assistant device:

anticipating detection of additional audio data via one or more additional microphones of an additional assistant device based on a plurality of historical acoustic events being detected at both the assistant device and the additional assistant device, the additional assistant device being in addition to the assistant device, and the additional assistant device being co-located in the ecosystem with the assistant device;

detecting, via the one or more additional microphones of the additional assistant device located in the ecosystem, the additional audio data that also captures the acoustic event, the additional assistant device being in addition to the assistant device, and the additional assistant device being co-located in the ecosystem with the assistant device;

processing, using an additional event detection model that is stored locally at the additional assistant device, the additional audio data that captures the acoustic event to generate an additional measure associated with the acoustic event, wherein the additional event detection model that is stored locally at the additional assistant device comprises an additional hotword detection model that is trained to detect whether the particular word or phrase is captured in the additional audio data;

determining, based on the measure satisfying a threshold indicating that the particular word or phrase is captured in the audio data and based on the additional measure satisfying the threshold indicating that the particular word or phrase is captured in the additional audio data, that the acoustic event detected by at least both the assistant device and the additional assistant device corresponds to an occurrence of an actual acoustic event; and

in response to determining that the acoustic event corresponds to an occurrence of the actual acoustic event, causing one or more components of an automated assistant to be activated at one or more of: the assistant device, the additional assistant device, or a further additional assistant device.

13. The system of claim 11 , wherein the measure associated with the acoustic event comprises a corresponding confidence level corresponding to whether the audio data captures the particular word or phrase, and wherein the additional measure associated with the acoustic event comprises a corresponding additional confidence level corresponding to whether the additional audio data captures the particular word or phrase, and wherein determining that the acoustic event corresponds to the occurrence of the actual acoustic event comprises determining the particular word or phrase is captured in both the audio data and the additional audio data based on the corresponding confidence level and the corresponding additional confidence level satisfying the threshold.

14. The system of claim 11 , wherein the hotword detection model that is stored locally at the assistant device is a distinct hotword model that is distinct from the additional hotword detection model that is stored locally at the additional assistant device.

15. The system of claim 11 , wherein the assistant device, the additional assistant device, and the further additional client device are co-located in an ecosystem of devices.

16. The system of claim 11 , wherein the audio data temporally corresponds to the additional audio data, and wherein determining that the acoustic event detected by at least both the assistant device and the additional assistant device corresponds to an occurrence of an actual acoustic event based on the measure satisfying the threshold indicating that the particular word or phrase is captured in the audio data and based on the additional measure satisfying the threshold indicating that the particular word or phrase is captured in the audio data is in response to determining that the audio data temporally corresponds to the additional audio data.

17. The system of claim 11 , wherein the audio data temporally corresponds to the additional audio data, and wherein determining that the acoustic event detected by both the assistant device and the additional assistant device corresponds to the occurrence of the actual acoustic event is in response to determining that a timestamp associated with the audio data temporally corresponds to an additional timestamp associated with the additional audio data.

18. The system of claim 11 , wherein the one or more processors are further operable to:

transmit, by the assistant device, and to a remote system, the audio data; and

transmit, by the additional assistant device, and to the remote system, the additional audio data,

wherein determining that the acoustic event detected by both the assistant device and the additional assistant device corresponds to the occurrence of the actual acoustic event is by the remote system.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Sep 15, 2023
From: SHARIFI, MATTHEW; CARBUNE, VICTOR
To: GOOGLE LLC
Reel/Frame 064915/0954 →
Continuity (2)
Continuation 17085926 · Oct 30, 2020
Related Publication 20230419951A1 · Dec 28, 2023
References Cited (37)
US 9293136B2 · Aleksic · 2016 [cited by examiner]
US 9728188B1 · Rosen · 2017 [cited by examiner]
US 10769203B1 · Sonasath et al. · 2020 [cited by applicant]
US 11195522B1 · Makashir · 2021 [cited by examiner]
US 20160055850A1 · Nakadaï · 2016 [cited by applicant]
US 20160078286A1 · Tani · 2016 [cited by applicant]
US 20170004132A1 · Sharifi · 2017 [cited by examiner]
US 20170025124A1 · Mixter · 2017 [cited by examiner]
US 20170083285A1 · Meyers et al. · 2017 [cited by applicant]
US 20170186433A1 · Alvarez Guevara · 2017 [cited by applicant]
US 20180114538A1 · Kakadiaris · 2018 [cited by applicant]
US 20180158453A1 · Campbell · 2018 [cited by applicant]
US 20180330589A1 · Horling · 2018 [cited by applicant]
US 20180330735A1 · Strope · 2018 [cited by applicant]
US 20190103094A1 · Neil · 2019 [cited by applicant]
US 20190115026A1 · Sharifi · 2019 [cited by examiner]
US 20190147873A1 · Voigt · 2019 [cited by applicant]
US 20190287528A1 · Hughes · 2019 [cited by applicant]
US 20190385594A1 · Park · 2019 [cited by applicant]
US 20200105256A1 · Fainberg et al. · 2020 [cited by applicant]
US 20200160867A1 · Roeck · 2020 [cited by applicant]
US 20220139371A1 · Sharifi et al. · 2022 [cited by applicant]
CN 108351872 · 2018 [cited by applicant]
CN 110364151 · 2019 [cited by applicant]
CN 110534102 · 2019 [cited by applicant]
CN 110853620 · 2020 [cited by applicant]
EP 3407348 · 2018 [cited by applicant]
EP 3975171 · 2022 [cited by applicant]
JP 2017072857 · 2017 [cited by applicant]
WO 2014174738 · 2014 [cited by applicant]
European Patent Office; Communication pursuant to Article 94(3) issued in Application No. 20829156.7; 7 pages; dated Jun. 1, 2023. [cited by applicant]
European Patent Office; International Search Report and Written Opinion of App. No. PCT/US2020/064988; 14 pages; dated Jul. 6, 2021. [cited by applicant]
Intellectual Property India; First Examination Report issued in Application No. 202227062653; 8 pages; dated Sep. 19, 2023. [cited by applicant]
Japanese Patent Office; Notice of Reasons of Rejection issued in Application No. 2022-569600; 10 pages; dated Jan. 15, 2024. [cited by applicant]
China National Intellectual Property Administration; Notification of First Office Action issued in Application No. 202080100908.3; 42 pages; dated May 24, 2024. [cited by applicant]
Intellectual Property India; Hearing Notice issued in Application No. 202227062653; 4 pages; dated Aug. 13, 2024. [cited by applicant]
China National Intellectual Property Administration; Notice of Grant issued in Application No. 202080100908.3; 6 pages; dated Sep. 30, 2024. [cited by applicant]
Cited By (1)
US 12,579,980