Proactive assistance via a cascade of LLMS
A method for providing proactive assistance includes obtaining, by a digital assistant, a contextual event associated with a user of a user device. The method includes determining, using a local large language model (LLM) executing on the user device, a remote LLM prompt confidence. The method includes determining that the remote LLM prompt confidence satisfies a threshold. Based on determining that the remote LLM prompt confidence satisfies the threshold, the method includes generating a remote LLM prompt for a remote LLM executing remote from the user device. The method includes transmitting, to the remote LLM, the remote LLM prompt. The method includes receiving, at the digital assistant, from the remote LLM, response content providing the proactive assistance associated with the contextual event. The method includes providing, for output from the user device, presentation content based on the response content received from the remote LLM.
1 . A computer-implemented method executed by data processing hardware that causes the data processing hardware to perform operations comprising:
obtaining, by a digital assistant, a contextual event associated with a user of a user device;
generating, using a local large language model (LLM) executing on the user device, initial response content associated with the contextual event, the initial response content providing an offer for the digital assistant to interact with a remote LLM to perform an action on the user's behalf based on the contextual event;
providing, for output from the user device, initial presentation content based on the initial response content, the initial presentation content prompting the user to consent to the offer for the digital assistant to interact with the remote LLM to perform the action on the user's behalf;
receiving an initial presentation content interaction indicating user interaction with the initial presentation content:
determining, using the local LLM, a remote LLM prompt confidence based on the received initial presentation content, the remote LLM prompt confidence indicating a likelihood of prompting a remote LLM for proactive assistance associated with the contextual event;
determining that the remote LLM prompt confidence satisfies a threshold;
based on determining that the remote LLM prompt confidence satisfies the threshold, generating a remote LLM prompt for the remote LLM executing remote from the user device;
transmitting, to the remote LLM, the remote LLM prompt;
receiving, at the digital assistant, from the remote LLM, response content providing the proactive assistance associated with the contextual event; and
providing, for output from the user device, presentation content based on the response content received from the remote LLM.
2 . The method of claim 1 , wherein the remote LLM prompt confidence comprises a probability generated by the local LLM.
3 . The method of claim 1 , wherein
the remote LLM prompt confidence comprises the initial presentation content interaction.
4 . The method of claim 1 , wherein the initial presentation content interaction comprises user consent for transmitting the remote LLM prompt to the remote LLM.
5 . The method of claim 1 , wherein the remote LLM prompt is based on output from the local LLM.
6 . The method of claim 5 , wherein the output comprises a summary of the contextual event.
7 . The method of claim 5 , wherein:
the contextual event comprises personal identification information associated with the user; and
generating the remote LLM prompt comprises redacting, using the output from the local LLM, the personal identification information.
8 . The method of claim 1 , wherein the contextual event comprises at least one of:
sensor data captured by a sensor of the user device; or
application-specific data generated by another application executing on the user device.
9 . The method of claim 1 , wherein:
the contextual event comprises non-textual data; and
the operations further comprise transforming the non-textual data into textual data.
10 . The method of claim 1 , wherein:
the operations further comprise batching a plurality of contextual events together; and
determining the remote LLM prompt confidence is further based on the batched plurality of contextual events.
11 . A system comprising:
data processing hardware; and
memory hardware in communication with the data processing hardware, the memory hardware storing instructions that when executed on the data processing hardware cause the data processing hardware to perform operations comprising:
obtaining, by a digital assistant, a contextual event associated with a user of a user device;
generating, using a local large language model (LLM) executing on the user device, initial response content associated with the contextual event, the initial response content providing an offer for the digital assistant to interact with a remote LLM to perform an action on the user's behalf based on the contextual event;
providing, for output from the user device, initial presentation content based on the initial response content, the initial presentation content prompting the user to consent to the offer for the digital assistant to interact with the remote LLM to perform the action on the user's behalf;
receiving an initial presentation content interaction indicating user interaction with the initial presentation content;
determining, using the local LLM, a remote LLM prompt confidence based on the received initial presentation content, the remote LLM prompt confidence indicating a likelihood of prompting a remote LLM for proactive assistance associated with the contextual event;
determining that the remote LLM prompt confidence satisfies a threshold;
based on determining that the remote LLM prompt confidence satisfies the threshold, generating a remote LLM prompt for the remote LLM executing remote from the user device;
transmitting, to the remote LLM, the remote LLM prompt;
receiving, at the digital assistant, from the remote LLM, response content providing the proactive assistance associated with the contextual event; and
providing, for output from the user device, presentation content based on the response content received from the remote LLM.
12 . The system of claim 11 , wherein the remote LLM prompt confidence comprises a probability generated by the local LLM.
13 . The system of claim 11 , wherein
the remote LLM prompt confidence comprises the initial presentation content interaction.
14 . The system of claim 11 , wherein the initial presentation content interaction comprises user consent for transmitting the remote LLM prompt to the remote LLM.
15 . The system of claim 11 , wherein the remote LLM prompt is based on output from the local LLM.
16 . The system of claim 15 , wherein the output comprises a summary of the contextual event.
17 . The system of claim 15 , wherein:
the contextual event comprises personal identification information associated with the user; and
generating the remote LLM prompt comprises redacting, using the output from the local LLM, the personal identification information.
18 . The system of claim 11 , wherein the contextual event comprises at least one of:
sensor data captured by a sensor of the user device; or
application-specific data generated by another application executing on the user device.
19 . The system of claim 11 , wherein:
the contextual event comprises non-textual data; and
the operations further comprise transforming the non-textual data into textual data.
20 . The system of claim 11 , wherein:
the operations further comprise batching a plurality of contextual events together; and
determining the remote LLM prompt confidence is further based on the batched plurality of contextual events.