IP Library › Granted Patent US 12,322,391
Granted Patent B2
US 12,322,391 · App. 18/770,316 · Granted Jun 3, 2025

Multiple concurrent voice assistants

Inventors: Jonathan Hayden Gomes (San Diego, CA); Shashank Goel (Milpitas, CA); Oscar Armando Azucena (Mountain View, CA); Patrick Berny (Los Altos, CA); Keun-Young Park (Santa Clara, CA); Matthew William Crowley (Los Altos, CA)
Assignee: GOOGLE LLC
G10L15/22G06F3/14G06F3/165G10L15/16G10L21/0208H04M1/271H04R3/00G10L2015/088G10L15/30G10L2021/02082
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,322,391
App. No.
18/770,316
Granted
Jun 3, 2025
Kind
B2
Abstract

Techniques are described herein for concurrent voice assistants. A method includes: providing first and second automated assistants with access to one or more microphones; receiving, from the first automated assistant, an indication that the first automated assistant has initiated a first session, and in response: continuing providing, to the first automated assistant, access to the one or more microphones; discontinuing providing, to the second automated assistant, access to the one or more microphones; and preventing the second automated assistant from accessing one or more portions of an output audio data stream; receiving, from the first automated assistant, an indication that the first session has ended, and in response: continuing providing, to the first automated assistant, access to the one or more microphones; resuming providing, to the second automated assistant, access to the one or more microphones; and resuming providing, to the second automated assistant, the output audio data stream.

Claims (80)

1. A computer program product comprising one or more non-transitory computer-readable storage media having program instructions collectively stored on the one or more computer-readable storage media, the program instructions executable to:

concurrently provide a first automated assistant and a second automated assistant with access to one or more microphones;

receive, from the first automated assistant, an indication that the first automated assistant has initiated a first session;

in response to receiving, from the first automated assistant, the indication that the first automated assistant has initiated the first session:

continue providing, to the first automated assistant, access to the one or more microphones;

discontinue providing, to the second automated assistant, access to the one or more microphones; and

prevent the second automated assistant from accessing one or more portions of an output audio data stream provided for rendering via one or more speakers, the one or more portions including output audio data of the first automated assistant;

receive, from the first automated assistant, an indication that the first session has ended; and

in response to receiving, from the first automated assistant, the indication that the first session has ended:

continue providing, to the first automated assistant, access to the one or more microphones;

resume providing, to the second automated assistant, access to the one or more microphones; and

resume providing, to the second automated assistant, the output audio data stream, wherein the second automated assistant uses the output audio data stream in noise cancellation.

2. The computer program product according to claim 1 , wherein in preventing the second automated assistant from accessing the one or more portions of the output audio data stream the program instructions are executable to:

prevent the second automated assistant from accessing the output audio data stream while the first automated assistant is providing audible output during the first session.

3. The computer program product according to claim 1 , wherein the program instructions are further executable to:

determine that a phone call has been initiated;

in response to determining that the phone call has been initiated, discontinue providing, to the first automated assistant and to the second automated assistant, access to the one or more microphones;

determine that the phone call has ended; and

in response to determining that the phone call has ended, resume providing, to the first automated assistant and to the second automated assistant, access to the one or more microphones.

4. The computer program product according to claim 1 , wherein the program instructions are further executable to:

subsequent to receiving, from the first automated assistant, the indication that the first automated assistant has initiated the first session, and prior to receiving, from the first automated assistant, the indication that the first session has ended:

identify a notification, from the second automated assistant, to be output;

determine a priority of the notification; and

in response to the priority of the notification satisfying a threshold, prior to receiving, from the first automated assistant, the indication that the first session has ended, output the notification.

5. The computer program product according to claim 4 , wherein the program instructions are further executable to:

buffer, in a buffer, the one or more portions of the output audio data stream that include the output audio data of the first automated assistant;

pause presentation of the one or more portions of the output audio data stream that include the output audio data of the first automated assistant, while outputting the notification from the second automated assistant; and

after outputting the notification from the second automated assistant, resume, from the buffer, presentation of the one or more portions of the output audio data stream that include the output audio data of the first automated assistant.

6. The computer program product according to claim 1 , wherein the program instructions are further executable to:

subsequent to receiving, from the first automated assistant, the indication that the first automated assistant has initiated the first session, and prior to receiving, from the first automated assistant, the indication that the first session has ended:

identify a notification, from the second automated assistant, to be output;

determining a priority of the notification; and

in response to the priority of the notification not satisfying a threshold, suppress the notification until receiving, from the first automated assistant, the indication that the first session has ended.

7. The computer program product according to claim 6 , wherein the program instructions are further executable to:

buffer, in a buffer, the notification;

in response to receiving, from the first automated assistant, the indication that the first session has ended, output, from the buffer, the notification.

8. The computer program product according to claim 1 , wherein the program instructions are further executable to:

receive, via a graphical user interface, user interface input that is a request to initiate, on the second automated assistant, a second session;

in response to receiving the user interface input:

cause the second automated assistant to initiate the second session;

continue providing, to the second automated assistant, access to the one or more microphones;

discontinue providing, to the first automated assistant, access to the one or more microphones; and

prevent the first automated assistant from accessing output audio data of the second automated assistant.

9. The computer program product according to claim 1 , wherein the program instructions are further executable to, subsequent to receiving, from the first automated assistant, the indication that the first automated assistant has initiated the first session, and prior to receiving, from the first automated assistant, the indication that the first session has ended:

display, on a graphical user interface, a visual indication that access to the one or more microphones is being provided to the first automated assistant.

10. The computer program product according to claim 1 , wherein the program instructions are further executable to, subsequent to receiving, from the first automated assistant, the indication that the first automated assistant has initiated the first session:

provide an audible indication that access to the one or more microphones is being provided to the first automated assistant.

11. The computer program product according to claim 1 , wherein the program instructions are further executable to, in response to receiving, from the first automated assistant, the indication that the first automated assistant has initiated the first session:

reassign a physical button, that, upon activation, is configured to initiate a session of the second automated assistant, to instead initiate a session of the first automated assistant.

12. A method implemented by one or more processors, the method comprising:

receiving, via one or more microphones, first audio data that captures a first spoken utterance of a user;

concurrently providing the first audio data to a first hotword detector of a first automated assistant and to a second hotword detector of a second automated assistant;

receiving a first confidence score based on a probability, determined by the first hotword detector of the first automated assistant, of a first hotword being present in the first audio data, and a second confidence score based on a probability, determined by the second hotword detector of the second automated assistant, of a second hotword being present in the first audio data;

based on the first confidence score and the second confidence score:

providing, to the first automated assistant, second audio data, received via the one or more microphones, that captures a second spoken utterance of the user that follows the first spoken utterance of the user;

providing an indication that audio from the one or more microphones is being provided to the first automated assistant; and

preventing the second automated assistant from obtaining the second audio data.

13. The method according to claim 12 , wherein:

the first hotword detector processes the first audio data using one or more machine learning models of the first hotword detector to generate a first predicted output that indicates the probability of the first hotword being present in the first audio data; and

the second hotword detector processes the first audio data using one or more machine learning models of the second hotword detector to generate a second predicted output that indicates the probability of the second hotword being present in the first audio data.

14. The method according to claim 13 , wherein:

the first confidence score is higher than the second confidence score; and

the first confidence score satisfies a threshold.

15. The method according to claim 14 , wherein the indication that audio from the one or more microphones is being provided to the first automated assistant is a visual indication that is displayed on a graphical user interface, or an audible indication.

16. A system comprising:

a processor, a computer-readable memory, one or more non-transitory computer-readable storage media, and program instructions collectively stored on the one or more computer-readable storage media, the program instructions executable to:

receive, via one or more microphones, first audio data that captures a first spoken utterance of a user;

concurrently provide the first audio data to a first hotword detector of a first automated assistant and to a second hotword detector of a second automated assistant;

receive a first confidence score based on a probability, determined by the first hotword detector of the first automated assistant, of a first hotword being present in the first audio data, and a second confidence score based on a probability, determined by the second hotword detector of the second automated assistant, of a second hotword being present in the first audio data;

based on the first confidence score and the second confidence score:

provide, to the first automated assistant, second audio data, received via the one or more microphones, that captures a second spoken utterance of the user that follows the first spoken utterance of the user;

provide an indication that audio from the one or more microphones is being provided to the first automated assistant; and

prevent the second automated assistant from obtaining the second audio data.

17. The system according to claim 16 , wherein:

the first hotword detector processes the first audio data using one or more machine learning models of the first hotword detector to generate a first predicted output that indicates the probability of the first hotword being present in the first audio data; and

the second hotword detector processes the first audio data using one or more machine learning models of the second hotword detector to generate a second predicted output that indicates the probability of the second hotword being present in the first audio data.

18. The system according to claim 16 , wherein:

the first confidence score is higher than the second confidence score; and

the first confidence score satisfies a threshold.

19. The system according to claim 16 , wherein the indication that audio from the one or more microphones is being provided to the first automated assistant is a visual indication that is displayed on a graphical user interface, or an audible indication.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Aug 26, 2024
From: GOMES, JONATHAN HAYDEN; GOEL, SHASHANK; AZUCENA, OSCAR ARMANDO; BERNY, PATRICK; PARK, KEUN-YOUNG; CROWLEY, MATTHEW WILLIAM
To: GOOGLE LLC
Reel/Frame 068401/0495 →
Continuity (2)
Continuation 17722213 · Apr 15, 2022
Related Publication 20240363115A1 · Oct 31, 2024
References Cited (24)
US 10887734B2 · Dar · 2021 [cited by examiner]
US 20140135952A1 · Maehara · 2014 [cited by examiner]
US 20170105095A1 · Um · 2017 [cited by examiner]
US 20170329572A1 · Shah · 2017 [cited by examiner]
US 20180033431A1 · Newendorp · 2018 [cited by examiner]
US 20180151180A1 · Yehuday · 2018 [cited by examiner]
US 20190066672A1 · Wood · 2019 [cited by examiner]
US 20190206411A1 · Li · 2019 [cited by examiner]
US 20190215184A1 · Emigh · 2019 [cited by examiner]
US 20190251960A1 · Maker · 2019 [cited by examiner]
US 20200028734A1 · Emigh · 2020 [cited by examiner]
US 20200184964A1 · Myers et al. · 2020 [cited by applicant]
US 20200202853A1 · Ramic et al. · 2020 [cited by applicant]
US 20210289607A1 · Wilberding · 2021 [cited by applicant]
US 20220115009A1 · Sharifi · 2022 [cited by examiner]
US 20220172727A1 · Sharifi · 2022 [cited by examiner]
US 20230335127A1 · Gomes et al. · 2023 [cited by applicant]
US 20230352010A1 · Sharifi · 2023 [cited by examiner]
CN 113593550 · 2021 [cited by applicant]
WO 2020025186 · 2020 [cited by applicant]
WO 2022067345 · 2022 [cited by applicant]
European Patent Office; International Search Report and Written Opinion issued in Application No. PCT/US2022/052425; 22 pages; dated Jun. 20, 2023. [cited by applicant]
European Patent Office; Invitation to Pay Additional Fees issued in Application No. PCT/US2022/052425; 15 pages; dated Apr. 28, 2023. [cited by applicant]
“Future of In-Vehicle AI Powered Voice-Controlled Personal Assistant” FutureBridge. Retrieved from https://www.futurebridge.com/industry/perspectives-mobility/future-of-in-vehicle-ai-powered-voice-controlled-personal-as… [cited by applicant]