Notification management for multi-assistant artificial intelligence
A speech-processing system may provide access to multiple virtual assistants via a user device. A user may invoke a particular virtual assistant by speaking its wakeword. The virtual assistants may send notifications to the user via the user device, which may present an indication of the available notifications. The user may request delivery of the notifications using, for example, a voice user interface. When delivering notifications from multiple virtual assistants, the system may use distinct voice and/or visual characteristics to indicate a source of the notifications. For example, after delivering a first notification corresponding to a first virtual assistant (e.g., the one invoked by the user at the beginning of the interaction) the first virtual assistant may offer to deliver a second notification from a second virtual assistant. The second virtual assistant may deliver the second notification using voice/visual characteristics different from those used for the first notification.
1 . A computer-implemented method comprising:
receiving first data representing a first notification to be delivered to a first user, the first notification corresponding to a first virtual assistant;
receiving second data representing a second notification to be delivered to the first user, the second notification corresponding to a second virtual assistant different from the first virtual assistant;
causing a first device to present a first indication that the first notification is available and a second indication that the second notification is available;
receiving, from the first device, first audio data representing a first spoken request to receive notifications associated with the first virtual assistant;
determining that the first audio data includes a representation of a first wakeword corresponding to the first virtual assistant;
selecting first voice characteristics corresponding to the first virtual assistant;
receiving first synthesized speech corresponding to the first voice characteristics, the first synthesized speech representing content of the first notification;
in response to receiving the first audio data, causing, using the first data, the first device to output the first synthesized speech;
causing the first device to output second synthesized speech having the first voice characteristics and indicating availability of the second notification;
receiving, from the first device, second audio data representing a second spoken request to receive notifications associated with the second virtual assistant; and
in response to receiving the second audio data, causing, using the second data, the first device to output third synthesized speech representing the second notification, the third synthesized speech having second voice characteristics, different from the first voice characteristics, corresponding to the second virtual assistant.
2 . The computer-implemented method of claim 1 , further comprising:
in response to receiving the first audio data, sending, to a speech generation component, the first data and a first token corresponding to the first virtual assistant;
generating, by the speech generation component using the first data and the first token, the first synthesized speech;
in response to receiving the second audio data, sending, to the speech generation component, the second data and a second token corresponding to the first virtual assistant; and
generating, by the speech generation component using the second data and the second token, the second synthesized speech.
3 . The computer-implemented method of claim 1 , further comprising:
determining, by a notification component, that the second notification corresponds to the second virtual assistant different from the first virtual assistant; and
based at least in part on the second notification corresponding to the second virtual assistant, selecting the second voice characteristics.
4 . The computer-implemented method of claim 1 , further comprising:
causing the first device to present a third indication that a third notification corresponding to a third virtual assistant is available;
receiving, from the first device, third audio data representing a third spoken request to receive notifications associated with the third virtual assistant;
detecting, in the third audio data, a first representation of the first wakeword;
in response to detecting the first representation, causing the first device to output fourth synthesized speech having the first voice characteristics and indicating that the first virtual assistant cannot satisfy the third spoken request;
receiving, from the first device, fourth audio data representing a second wakeword corresponding to the third virtual assistant and a fourth spoken request to receive notifications associated with the third virtual assistant; and
in response to receiving the fourth audio data, causing the first device to output fifth synthesized speech having third voice characteristics corresponding to the third virtual assistant and representing the third notification.
5 . A computer-implemented method comprising:
receiving first data representing a first notification to be delivered to a user, the first notification corresponding to a first virtual assistant;
receiving second data representing a second notification to be delivered to the user, the second notification corresponding to a second virtual assistant different from the first virtual assistant;
causing a first device associated with the user to present a first indication that at least one notification is available;
receiving first input data representing a first request to receive notifications associated with the first virtual assistant;
determining that the first input data includes a representation of a first wakeword corresponding to the first virtual assistant;
selecting first voice characteristics corresponding to the first virtual assistant;
receiving first synthesized speech corresponding to the first voice characteristics, the first synthesized speech representing content of the first notification;
in response to receiving the first input data, causing, using the first data, the first device to present a first output representing the first notification, the first output including the first synthesized speech and having first characteristics corresponding to the first virtual assistant;
causing the first device to present a second output indicating availability of the second notification, the second output having the first characteristics;
receiving second input data representing a second request to receive notifications associated with the second virtual assistant; and
in response to receiving the second input data, causing, using the second data, the first device to present a third output representing the second notification, the third output having second characteristics, different from the first characteristics, corresponding to the second virtual assistant.
6 . The computer-implemented method of claim 5 , further comprising:
in response to determining that the first input data includes the representation of the first wakeword corresponding to the first virtual assistant, selecting the first voice characteristics.
7 . The computer-implemented method of claim 5 , further comprising:
in response to receiving the first input data, sending the first data to a speech generation component; and
generating, by the speech generation component using the first data, the first synthesized speech.
8 . The computer-implemented method of claim 5 , further comprising:
receiving third input data representing the first wakeword and a third request to receive notifications associated with a third virtual assistant;
causing the first device to present a fourth output having the first characteristics and indicating that a potential error involving the first virtual assistant and the third request;
receiving fourth input data representing a second wakeword corresponding to the third virtual assistant and a fourth request to receive notifications associated with the third virtual assistant; and
in response to receiving the fourth input data, causing the first device to present a fifth output having third characteristics corresponding to the third virtual assistant and representing a third notification corresponding to the third virtual assistant.
9 . The computer-implemented method of claim 5 , further comprising:
determining that the second notification corresponds to the second virtual assistant different from the first virtual assistant; and
based at least in part on the second notification corresponding to the second virtual assistant, selecting the second characteristics.
10 . The computer-implemented method of claim 9 , further comprising:
generating, using the first data and the first voice characteristics, the first synthesized speech, wherein the first output additionally includes first visual content and the first synthesized speech; and
generating, using the second data and second voice characteristics corresponding to the second virtual assistant, second synthesized speech representing a second summary of the second notification, wherein the third output includes second visual content and the second synthesized speech.
11 . The computer-implemented method of claim 5 , further comprising:
in response to receiving the first input data, causing the first device to present a first visual indication corresponding to the first virtual assistant during presentation of the first output;
in response to receiving the second input data, generating second synthesized speech representing the second notification, wherein the third output includes the second synthesized speech; and
causing the first device to present a second visual indication, different from the first visual indication, corresponding to the second virtual assistant during presentation of the third output.
12 . The computer-implemented method of claim 5 , further comprising:
in response to receiving the second data, determining that the user corresponds to the first device and a second device;
determining that a second wakeword associated with the second virtual assistant is enabled for the first device but not the second device; and
in response to determining that the second wakeword is enabled for the first device, causing the first device to present a second indication that the second notification is available.
13 . A system, comprising:
at least one processor; and
at least one memory comprising instructions that, when executed by the at least one processor, cause the system to:
receive first data representing a first notification to be delivered to a user, the first notification corresponding to a first virtual assistant;
receive second data representing a second notification to be delivered to the user, the second notification corresponding to a second virtual assistant different from the first virtual assistant;
cause a first device associated with the user to present a first indication that at least one notification is available;
receive first input data representing a first request to receive notifications associated with the first virtual assistant;
determine that the first input data includes a representation of a first wakeword corresponding to the first virtual assistant;
select first voice characteristics corresponding to the first virtual assistant;
receive first synthesized speech corresponding to the first voice characteristics, the first synthesized speech representing content of the first notification;
in response to receiving the first input data, cause, using the first data, the first device to present a first output representing the first notification, the first output including the first synthesized speech and having first characteristics corresponding to the first virtual assistant;
cause the first device to present a second output indicating availability of the second notification, the second output having the first characteristics;
receive second input data representing a second request to receive notifications associated with the second virtual assistant; and
in response to receiving the second input data, cause, using the second data, the first device to present a third output representing the second notification, the third output having second characteristics, different from the first characteristics, corresponding to the second virtual assistant.
14 . The system of claim 13 , wherein the at least one memory further comprises instructions that, when executed by the at least one processor, further cause the system to:
in response to determining that the first input data includes the representation of the first wakeword corresponding to the first virtual assistant, select the first voice characteristics.
15 . The system of claim 13 , wherein the at least one memory further comprises instructions that, when executed by the at least one processor, further cause the system to:
in response to receiving the first input data, send the first data to a speech generation component; and
generate, by the speech generation component using the first data, the first synthesized speech.
16 . The system of claim 13 , wherein the at least one memory further comprises instructions that, when executed by the at least one processor, further cause the system to:
receive third input data representing the first wakeword and a third request to receive notifications associated with a third virtual assistant;
cause the first device to present a fourth output having the first characteristics and indicating that a potential error involving the first virtual assistant and the third request;
receive fourth input data representing a second wakeword corresponding to the third virtual assistant and a fourth request to receive notifications associated with the third virtual assistant; and
in response to receiving the fourth input data, cause the first device to present a fifth output having third characteristics corresponding to the third virtual assistant and representing a third notification corresponding to the third virtual assistant.
17 . The system of claim 13 , wherein the at least one memory further comprises instructions that, when executed by the at least one processor, further cause the system to:
determine that the second notification corresponds to the second virtual assistant different from the first virtual assistant; and
based at least in part on the second notification corresponding to the second virtual assistant, select the second characteristics.
18 . The system of claim 17 , wherein the at least one memory further comprises instructions that, when executed by the at least one processor, further cause the system to:
generate, using the first data and the first voice characteristics, the first synthesized speech, wherein the first output additionally includes first visual content and the first synthesized speech; and
generate, using the second data and second voice characteristics corresponding to the second virtual assistant, second synthesized speech representing a second summary of the second notification, wherein the third output includes second visual content and the second synthesized speech.
19 . The system of claim 13 , wherein the at least one memory further comprises instructions that, when executed by the at least one processor, further cause the system to:
in response to receiving the first input data, cause the first device to present a first visual indication corresponding to the first virtual assistant during presentation of the first output;
in response to receiving the second input data, generate second synthesized speech representing the second notification, wherein the third output includes the second synthesized speech; and
cause the first device to present a second visual indication, different from the first visual indication, corresponding to the second virtual assistant during presentation of the third output.
20 . The system of claim 13 , wherein the at least one memory further comprises instructions that, when executed by the at least one processor, further cause the system to:
in response to receiving the second data, determine that the user corresponds to the first device and a second device;
determine that a second wakeword associated with the second virtual assistant is enabled for the first device but not the second device; and
in response to determining that the second wakeword is enabled for the first device, cause the first device to present a second indication that the second notification is available.