Command disambiguation based on environmental context
In one implementation, a method of changing a state of an object is performed at a device including an image sensor, one or more processors, and non-transitory memory. The method includes receiving a vocal command. The method includes obtaining, using the image sensor, an image of a physical environment. The method includes detecting, in the image of the physical environment, an object based on a visual model of the object stored in the non-transitory memory in association with an object identifier of the object. The method includes generating, based on the vocal command and detection of the object, an instruction including the object identifier of the object. The method includes effectuating the instruction to change a state of the object.
1 . A method comprising:
at an electronic device including an image sensor, one or more processors, and non-transitory memory:
receiving a vocal command;
obtaining, using the image sensor, an image of a physical environment;
detecting, in the image of the physical environment, an object based on a visual model of the object stored in the non-transitory memory in association with an object identifier of the object;
determining, based on the vocal command, a plurality of potential instructions;
selecting, based on the detection of the object, one of the plurality of potential instructions as an intended instruction; and
generating a data packet including the object identifier of the object to effectuate the intended instruction to change a state of the object.
2 . The method of claim 1 , wherein the object identifier of the object is further stored in association with an object type of the object and selecting the intended instruction is further based on the object type of the object.
3 . The method of claim 1 , wherein the object identifier of the object is further stored in association with a name of the object and selecting the intended instruction is further based on the name of the object.
4 . The method of claim 1 , wherein selecting the intended instruction is further based on determining that a gaze of a user is directed at a location of the detection of the object.
5 . The method of claim 1 , wherein selecting the intended instruction is further based on the state of the object.
6 . The method of claim 1 , further comprising:
receiving a second vocal command;
obtaining, using the image sensor, an image of a second physical environment;
detecting, in the image of the second physical environment, the object;
determining, based on the second vocal command, a second plurality of potential instructions;
selecting, based on the detection of the object, one of the second plurality of potential instructions as a second intended instruction; and
generating a second data packet including the object identifier of the object to effectuate the second intended instruction to change the state of the object.
7 . The method of claim 1 , further comprising:
obtaining a request to enroll the object;
obtaining, using the image sensor, one or more images of the object;
determining, based on the one or more images of the object, the visual model of the object; and
storing, in the non-transitory memory, the visual model of the object in association with the object identifier of the object.
8 . A device comprising:
an image sensor;
a non-transitory memory; and
one or more processors to:
receive a vocal command;
obtain, using the image sensor, an image of a physical environment;
detect, in the image of the physical environment, an object based on a visual model of the object stored in the non-transitory memory in association with an object identifier of the object;
determine, based on the vocal command, a plurality of potential instructions;
select, based on the detection of the object, one of the plurality of potential instructions as an intended instructions; and
generate a data packet including the object identifier of the object to effectuate the intended instruction to change a state of the object.
9 . The device of claim 8 , wherein the object identifier of the object is further stored in association with an object type of the object and the one or more processors are to select the intended instruction based on the object type of the object.
10 . The device of claim 8 , wherein the object identifier of the object is further stored in association with a name of the object and the one or more processors are to select the intended instruction based on the name of the object.
11 . The device of claim 8 , wherein the one or more processors are to select the intended instruction further based on determining that a gaze of a user is directed at a location of the detection of the object.
12 . The device of claim 8 , wherein the one or more processors are to select the intended instruction based on the state of the object.
13 . The device of claim 8 , wherein the one or more processors are further to:
receive a second vocal command;
obtain, using the image sensor, an image of a second physical environment;
detect, in the image of the second physical environment, the object;
determine, based on the second vocal command, a second plurality of potential instructions;
select, based on the detection of the object, one of the second plurality of potential instructions as a second intended instructions; and
generate a second data packet including the object identifier of the object to effectuate the second intended instruction to change the state of the object.
14 . The device of claim 8 , wherein the one or more processors are further to:
obtain a request to enroll the object;
obtain, using the image sensor, one or more images of the object;
determine, based on the one or more images of the object, the visual model of the object; and
store, in the non-transitory memory, the visual model of the object in association with the object identifier of the object.
15 . A non-transitory memory storing one or more programs, which, when executed by one or more processors of a device including an image sensor, cause the device to:
receive a vocal command;
obtain, using the image sensor, an image of a physical environment;
detect, in the image of the physical environment, an object based on a visual model of the object stored in the non-transitory memory in association with an object identifier of the object;
determine, based on the vocal command, a plurality of potential instructions;
select, based on the detection of the object, one of the plurality of potential instructions as an intended instructions; and
generate a data packet including the object identifier of the object to effectuate the intended instruction to change a state of the object.
16 . The non-transitory memory of claim 15 , wherein the object identifier of the object is further stored in association with an object type of the object and the programs, when executed, cause the device to select the intended instruction based on the object type of the object.
17 . The non-transitory memory of claim 15 , wherein the object identifier of the object is further stored in association with a name of the object and the programs, when executed, cause the device to select the intended instruction based on the name of the object.
18 . The non-transitory memory of claim 15 , wherein the programs, when executed, cause the device to select the intended instruction based on determining that a gaze of a user is directed at a location of the detection of the object.
19 . The non-transitory memory of claim 15 , wherein the programs, when executed, cause the device to select the intended instruction based on the state of the object.
20 . The non-transitory memory of claim 15 , wherein the programs, when executed, further cause the device to:
receive a second vocal command;
obtain, using the image sensor, an image of a second physical environment;
detect, in the image of the second physical environment, the object;
determine, based on the second vocal command, a second plurality of potential instructions;
select, based on the detection of the object, one of the second plurality of potential instructions as a second intended instruction; and
generate a second data packet including the object identifier of the object to effectuate the second intended instruction to change the state of the object.
21 . The non-transitory memory of claim 15 , wherein the programs, when executed, further cause the device to:
obtain a request to enroll the object;
obtain, using the image sensor, one or more images of the object;
determine, based on the one or more images of the object, the visual model of the object; and
store, in the non-transitory memory, the visual model of the object in association with the object identifier of the object.