Optimizing display engagement in action automation
Embodiments of the present invention provide systems, methods, and computer storage media directed to optimizing engagement with a display during digital assistant-performed operations in response to a received command. The digital assistant generates an overlay having user interface elements that present information determined to be relevant to a user based on the received command and contextual data. The overlay is presented over the underlying operations performed on corresponding applications to mask the visible steps of the operations being performed. In this way, the digital assistant optimizes display resources that are typically rendered useless during the processing of digital assistant-performed operations.
1 . A system supported by a user device, the system comprising:
at least one processor;
computer storage media storing computer-usable instructions that, when used by the at least one processor, cause the at least one processor to:
receive a command by a digital assistant of a mobile device;
select an action dataset associated with the received command;
determine that a first automated action of multiple automated operations defined by the action dataset and associated with an application is to be executed in response to the received command;
generate an overlay interface that includes a first user interface element configured to present content relevant to one or more parameters of the received command and masks a visual output generated by the application;
cause performance of the multiple automated operations defined by the action dataset; and
during performance of the multiple automated operations, cause the generated overlay interface to be presented via the application of the user device.
2 . The system of claim 1 , wherein the generated overlay interface is displayed for at least a latency period that corresponds to a duration of the performance of the multiple automated operations.
3 . The system of claim 1 , wherein the content relevant to one or more parameters of the received command is determined to be semantically or contextually relevant to the first automated action.
4 . The system of claim 1 , wherein the performance of the multiple automated operations defined by the action dataset is caused by initiating the action dataset.
5 . The system of claim 1 , wherein the at least one processor removes the generated overlay interface, after presenting the generated overlay interface, to cause the application to display resulting visual output data based on a completed performance of the multiple automated operations.
6 . The system of claim 1 , wherein the multiple automated operations include an emulated touch event, an invocation of a deep link, an automated inclusion of at least one parameter, or an automated entry of input data.
7 . The system of claim 1 , wherein the command is received based on generated speech-to-text data.
8 . The system of claim 1 , wherein the overlay interface is based on the action dataset.
9 . A method performed by a digital assistant of a user device, the method comprising:
receiving a command from a user of the user device;
selecting an action dataset associated with the received command and that defines multiple automated operations to be performed by an application of the user device;
determining that a first automated action associated with the application is to be executed in response to the received command;
generating an overlay interface that includes a first user interface element configured to present content and to mask a visual output generated by the application;
causing performance of the multiple automated operations defined by the action dataset; and
during performance of the multiple automated operations, causing the generated overlay interface to be presented via the application of the user device.
10 . The method of claim 9 , wherein the content is semantically or contextually relevant to the first automated action.
11 . The method of claim 9 , wherein the generated overlay interface is displayed for at least a latency period that corresponds to a duration of the performance of the multiple automated operations.
12 . The method of claim 9 , wherein the performance of the multiple automated operations defined by the action dataset is caused by initiating the action dataset.
13 . The method of claim 9 , wherein the at least one processor removes the generated overlay interface, after presenting the generated overlay interface, to cause the application to display resulting visual output data based on a completed performance of the multiple automated operations.
14 . The method of claim 9 , wherein the multiple automated operations include an emulated touch event, an invocation of a deep link, an automated inclusion of at least one parameter, or an automated entry of input data.
15 . The method of claim 9 , wherein the command is received based on generated speech-to-text data.
16 . The method of claim 9 , wherein the overlay interface is based on the action dataset.
17 . The method of claim 9 , wherein the first user interface element has a size that only masks the visual output generated by the first automated action.
18 . The method of claim 9 , wherein the first user interface element has a size that masks an entire display of the user device.
19 . A non-transitory computer-readable medium, whose contents, when executed by a digital assistant of a user device, cause the digital assistant to perform a method, the method comprising:
receiving a command from a user of the user device;
selecting an action dataset associated with the received command that defines multiple automated operations to be performed by an application of the user device;
generating a masking overlay interface that includes a user interface element configured to present content relevant to one or more parameters of the received command and masks a visual output generated by the application;
causing performance of the multiple automated operations defined by the action dataset; and
during performance of the multiple automated operations, causing the generated masking overlay interface to be presented via the application of the user device.
20 . The non-transitory computer-readable medium of claim 19 , wherein content of the masking overlay interface is semantically or contextually relevant to a user of the user device.