Multimodal entity and coreference resolution for assistant systems
In one embodiment, a method includes receiving, at a client system, an audio input, where the audio input comprises a coreference to a target object, accessing visual data from one or more camera associated with the client system, where the visual data comprises images portraying one or more objects, resolving the coreference to the target object from among the one or more objects, resoling the target object to a specific entity, and providing, at the client system, a response to the audio input, where the response comprises information about the specific entity.
1 . A computer-implemented method, executed at a computing device comprising at least a processor and a non-transitory computer-readable memory device, the computer-implemented method comprising:
identifying, over a period of time, a plurality of objects portrayed in visual data, wherein the visual data is captured at a camera of a head-worn device worn by a user;
updating, over the period of time, a multimodal dialog state based on the plurality of objects in the visual data;
after the period of time and in response to obtaining an audio input, identifying a coreference to a target object based on the audio input, wherein the audio input is captured at a microphone of the head-worn device;
resolving the coreference to the target object from among the plurality of objects by identifying, based on the audio input and the multimodal dialog state, the target object from among the plurality of objects in the visual data; and
selecting an action to execute, at the head-worn device, based on the target object.
2 . The method of claim 1 , wherein the action comprises notifying the user of a relationship between the coreference and the target object.
3 . The method of claim 1 , wherein, the audio input is obtained after the period of time and while the user is not looking at the target object.
4 . The method of claim 1 , wherein the visual data is associated with a second sensor of the client system.
5 . The method of claim 1 , wherein the visual data is video data.
6 . The method of claim 1 , further comprising, after the period of time:
identifying, over another period of time, another plurality of objects portrayed in other visual data, wherein the other visual data is captured at the camera;
updating, over the other period of time, the multimodal dialog state based on the other plurality of objects in the other visual data.
7 . The method of claim 1 , further comprising:
analyzing the visual data to identify the plurality of objects portrayed in the visual data;
parsing the audio input to identify an intent of the audio input; and
updating the dialog state to include the intent of the audio input.
8 . The method of claim 7 , further comprising:
classifying the intent based on one or more pre-defined taxonomies of semantic intentions.
9 . The method of claim 1 , further comprising:
receiving one or more of a gesture and gaze information from the head-worn device; and
updating the dialog state to include one or more of the gesture and the gaze information.
10 . The method of claim 1 , wherein the resolving the coreference to the target object from among the plurality of objects is further based on additional information included in the dialog state.
11 . The method of claim 1 , wherein the action comprises one or more of a visual response and a speech response.
12 . The method of claim 1 , wherein one or more of the plurality of objects are virtual objects in a virtual reality environment.
13 . The method of claim 1 , wherein the resolving the coreference to the target object from among the plurality of objects includes identifying the target object from among the plurality of objects based on a position of the target object within a field of view associated with the visual data.
14 . One or more computer-readable non-transitory storage media comprising instructions executable by a processor to:
identify, over a period of time, a plurality of objects portrayed in visual data, wherein the visual data is captured at a camera of a head-worn device worn by a user;
update, over the period of time, a multimodal dialog state based on the plurality of objects in the visual data;
after the period of time and in response to obtaining an audio input, identify a coreference to a target object based on the audio input, wherein the audio input is captured at a microphone of the head-worn device;
resolve the coreference to the target object from among the plurality of objects by identifying, based on the audio input and the multimodal dialog state, the target object from among the plurality of objects in the visual data; and
select an action to execute, at the head-worn device, based on the target object.
15 . The media of claim 14 , wherein the action comprises notifying the user of a relationship between the coreference and the target object.
16 . The media of claim 14 , wherein the action comprises one or more of a visual response and a speech response.
17 . A processing system comprising:
one or more processors; and
a non-transitory memory coupled to the processors comprising instructions executable by the processors, the processors operable when executing the instructions to:
identify, over a period of time, a plurality of objects portrayed in visual data, wherein the visual data is captured at a camera of a head-worn device worn by a user;
update, over the period of time, a multimodal dialog state based on the plurality of objects in the visual data;
after the period of time and in response to obtaining an audio input, identify a coreference to a target object based on the audio input, wherein the audio input is captured at a microphone of the head-worn device;
resolve the coreference to the target object from among the plurality of objects by identifying, based on the audio input and the multimodal dialog state, the target object from among the plurality of objects in the visual data; and
select an action to execute, at the head-worn device, based on the target object.
18 . The processing system of claim 17 , wherein the action comprises notifying the user of a relationship between the coreference and the target object.
19 . The processing system of claim 17 , wherein the action comprises one or more of a visual response and a speech response.
20 . The processing system of claim 17 , wherein to resolve the coreference to the target object from among the plurality of objects, the processors are operable when executing the instructions to identify the target object from among the plurality of objects based on a position of the target object within a field of view associated with the visual data.