RESOLVING AUTOMATED ASSISTANT REQUESTS THAT ARE BASED ON IMAGE(S) AND/OR OTHER SENSOR DATA
Methods, apparatus, and computer readable media are described related to causing processing of sensor data to be performed in response to determining a request related to an environmental object that is likely captured by the sensor data. Some implementations further relate to determining whether the request is resolvable based on the processing of the sensor data. When it is determined that the request is not resolvable, a prompt is determined and provided as user interface output, where the prompt provides guidance on further input that will enable the request to be resolved. In those implementations, the further input (e.g., additional sensor data and/or the user interface input) received in response to the prompt can then be utilized to resolve the request.
1 . A method implemented by one or more processors, the method comprising:
receiving a natural language input that is generated based on user input via an interface of a client device;
determining, based on processing the natural language input, that the natural language input indicates a request that is for an agent and that is related to an object in an environment with the client device;
in response to determining that the natural language input indicates the request that is for the agent and that is related to the object in the environment:
causing an image of the environment to be captured via a camera of the client device;
resolving one or more attributes of the object based on processing of the image;
generating an agent query, for the agent, that includes the one or more attributes of the object resolved based on processing of the image;
transmitting the agent query to the agent;
receiving, from the agent, responsive content that is responsive to the agent query; and
causing output, that reflects the responsive content, to be rendered at the client device.
2 . The method of claim 1 , further comprising:
selecting the agent from a plurality of available agents;
wherein transmitting the agent query to the agent is based on selecting the agent from the plurality of available agents.
3 . The method of claim 2 , wherein selecting the agent from the plurality of available agents is based on at least one of the one or more attributes of the object.
4 . The method of claim 3 , wherein transmitting the agent query to the agent comprises:
transmitting the query to the agent over one or more networks.
5 . The method of claim 1 , further comprising:
selecting the agent from a plurality of available agents;
wherein transmitting the query the agent is based on selecting the agent from the plurality of available agents.
6 . The method of claim 1 , further comprising:
determining that the request is not resolvable using the one or more attributes resolved based on processing of the image; and
in response to determining that the request is not resolvable:
providing, for presentation via the client device, a prompt; and
receiving, in response to the prompt, one or both of:
an additional image captured by the camera, and
voice input;
wherein generating the agent query is further based on one or both of:
the additional image received in response to the prompt, and
the voice input received in response to the prompt.
7 . The method of claim 6 , wherein the additional image is received in response to the prompt, and wherein generating the agent query is further based on the additional image.
8 . The method of claim 6 , wherein the voice input is received in response to the prompt, and wherein generating the agent query is further based on the voice input.
9 . The method of claim 6 , wherein the additional image and the voice input are received in response to the prompt, and wherein generating the agent query is further based on the additional image and the voice input.
10 . A system, comprising:
memory storing instructions;
one or more processors operable to execute the instructions to:
receive a natural language input that is generated based on user input via an interface of a client device;
determine, based on processing the natural language input, that the natural language input indicates a request that is for an agent and that is related to an object in an environment with the client device;
in response to determining that the natural language input indicates the request that is for the agent and that is related to the object in the environment:
cause an image of the environment to be captured via a camera of the client device;
resolve one or more attributes of the object based on processing of the image;
generate an agent query, for the agent, that includes the one or more attributes of the object resolved based on processing of the image;
transmit the agent query to the agent;
receive, from the agent, responsive content that is responsive to the agent query; and
cause output, that reflects the responsive content, to be rendered at the client device.
11 . The system of claim 10 , wherein one or more of the processors are further operable to execute the instructions to:
select the agent from a plurality of available agents;
wherein transmitting the agent query to the agent is based on selecting the agent from the plurality of available agents.
12 . The system of claim 11 , wherein in selecting the agent from the plurality of available agents one or more of the processors are to select the agent based on at least one of the one or more attributes of the object.
13 . The system of claim 12 , wherein in transmitting the agent query to the agent one or more of the processors are to transmit the query to the agent over one or more networks.
14 . The system of claim 10 , wherein one or more of the processors are further operable to execute the instructions to:
select the agent from a plurality of available agents;
wherein transmitting the query the agent is based on selecting the agent from the plurality of available agents.
15 . The system of claim 10 , wherein one or more of the processors are further operable to execute the instructions to:
determine that the request is not resolvable using the one or more attributes resolved based on processing of the image; and
in response to determining that the request is not resolvable:
provide, for presentation via the client device, a prompt; and
receive, in response to the prompt, one or both of:
an additional image captured by the camera, and
voice input;
wherein generating the agent query is further based on one or both of:
the additional image received in response to the prompt, and
the voice input received in response to the prompt.
16 . The system of claim 15 , wherein the additional image is received in response to the prompt, and wherein generating the agent query is further based on the additional image.
17 . The system of claim 15 , wherein the voice input is received in response to the prompt, and wherein generating the agent query is further based on the voice input.
18 . The system of claim 15 , wherein the additional image and the voice input are received in response to the prompt, and wherein generating the agent query is further based on the additional image and the voice input.