Systems and methods for efficient multimodal search refinement
Systems and methods of the present disclosure are directed to a computer-implemented method for multimodal search refinement. The method includes obtaining a visual search query from a user comprising one or more query images. The method includes providing a search interface for display to the user, the search interface comprising one or more result images responsive to the one or more query images and an interface element indicative of a request to the user to refine the visual search query. The method includes obtaining, from the user, textual data comprising a refinement to the visual search query. The method includes appending, by the computing system, the textual data to the visual search query to obtain a multimodal search query.
1 . A computer-implemented method for multimodal search refinement, the method comprising:
obtaining, by a computing system comprising one or more computing devices, a visual search query from a user comprising one or more query images;
causing, by the computing system, a search interface to be displayed for the user, the search interface comprising one or more result images responsive to the one or more query images and an interface element indicative of a request to the user to refine the visual search query;
obtaining, by the computing system from the user, textual data comprising a refinement to the visual search query;
determining, by the computing system, based on using a machine-learned model to process the one or more query images and the textual data, a multimodal search query, wherein the machine-learned model is trained to modify the one or more query images based on the textual data;
retrieving, by the computing system, one or more refined search results based on the multimodal search query, wherein a first refined search result of the one or more refined search results comprises an image associated with a particular web page;
processing, by the computing system, textual web content with the machine-learned model to obtain textual content descriptive of the first refined search result of the one or more refined search results, wherein the textual web content is extracted from the particular web page; and
causing, by the computing system, a refined search interface to be displayed for the user, the refined search interface comprising a refined search result interface element that comprises the first refined search result and the textual content descriptive of the first refined search result.
2 . The computer-implemented method of claim 1 , wherein, prior to obtaining the visual search query from the user, the method comprises:
causing, by the computing system, an interface for a virtual assistant application to be displayed to the user, wherein the interface for the virtual assistant application comprises an interface element indicative of a visual search feature of the virtual assistant application.
3 . The computer-implemented method of claim 1 , wherein the one or more refined search results further comprises a second refined search result, comprising:
a refined result image;
refined result video data;
an interface element comprising textual content responsive to the multimodal search query;
an interface element comprising a link to content responsive to the multimodal search query;
a commerce element comprising information descriptive of a product responsive to the multimodal search query; or
a multimedia interface element comprising textual content, one or more images, video data, a link to content responsive to the multimodal search query, and/or audio data.
4 . The computer-implemented method of claim 1 , wherein the refined search interface further comprises a textual input field for further refinement of the multimodal search query.
5 . The computer-implemented method of claim 1 , wherein the interface element indicative of the request to the user to refine the visual search query comprises a textual input field.
6 . The computer-implemented method of claim 1 , wherein the interface element indicative of the request to the user to refine the visual search query comprises a navigational element configured to navigate the user to a second interface.
7 . The computer-implemented method of claim 6 , wherein the second interface comprises a textual input field for input of the refinement to the visual search query; and
wherein obtaining the textual data comprising the refinement to the visual search query comprises obtaining, by the computing system from the user, textual data comprising a refinement to the visual search query via the textual input field.
8 . The computer-implemented method of claim 1 , wherein the interface element indicative of the request to the user to refine the visual search query comprises a voice interface element for collection of voice data comprising a spoken utterance from the user that is descriptive of the refinement to the visual search query.
9 . The computer-implemented method of claim 1 , wherein the one or more result images comprises a plurality of images that collectively form video data.
10 . The computer-implemented method of claim 1 , wherein obtaining the visual search query from the user comprising the one or more query images comprises:
obtaining, by the computing system, a first query image depicting a first object, wherein the first object comprises a plurality of different properties; and
wherein obtaining the textual data comprises:
obtaining, by the computing system from the user, textual data comprising the refinement to the visual search query, wherein the refinement is indicative of a value for a particular property of the plurality of different properties.
11 . The computer-implemented method of claim 10 , wherein appending the textual data to the visual search query to obtain the multimodal search query comprises:
determining, by the computing system, an intended property for the refinement from the plurality of different properties; and
generating, by the computing system, the multimodal search query, wherein the multimodal search query comprises a query for content associated with the object comprising the value for the particular property.
12 . A computing system for multimodal search refinement, comprising:
one or more processors;
one or more non-transitory computer-readable media that store instructions that, when executed by the one or more processors, cause the computing system to perform operations, the operations comprising:
obtaining a visual search query from a user comprising one or more query images;
providing a search interface for display to the user, the search interface comprising one or more result images responsive to the one or more query images and an interface element indicative of a request to the user to refine the visual search query;
obtaining textual data comprising a refinement to the visual search query;
determining, based on using a machine-learned model to process the one or more query images and the textual data, a multimodal search query, wherein the machine-learned model is trained to modify the one or more query images based on the textual data;
retrieving one or more refined search results based on the multimodal search query, wherein a first refined search result of the one or more refined search results comprises an image associated with a particular web page;
processing textual web content with the machine-learned model to obtain textual content descriptive of the first refined search result of the one or more refined search results, wherein the textual web content is extracted from the particular web page; and
causing a refined search interface to be displayed for the user, the refined search interface comprising a refined search result interface element that comprises the first refined search result and the textual content descriptive of the first refined search result.
13 . The computing system of claim 12 , wherein, prior to obtaining the visual search query from the user, the operations comprise:
causing an interface for a virtual assistant application to be displayed to the user, wherein the interface for the virtual assistant application comprises an interface element indicative of a visual search feature of the virtual assistant application.
14 . The computing system of claim 12 , wherein the one or more refined search results further comprises a second refined search result, comprising:
a refined result image;
refined result video data;
an interface element comprising textual content responsive to the multimodal search query;
an interface element comprising a link to content responsive to the multimodal search query;
a commerce element comprising information descriptive of a product responsive to the multimodal search query; or
a multimedia interface element comprising textual content, one or more images, video data, a link to content responsive to the multimodal search query, and/or audio data.
15 . The computing system of claim 12 , wherein the refined search interface further comprises a textual input field for further refinement of the multimodal search query.
16 . The computing system of claim 12 , wherein the interface element indicative of the request to the user to refine the visual search query comprises a textual input field.
17 . The computing system of claim 12 , wherein the interface element indicative of the request to the user to refine the visual search query comprises a navigational element configured to navigate the user to a second interface;
wherein the second interface comprises a textual input field for input of the refinement to the visual search query; and
wherein obtaining the textual data comprising the refinement to the visual search query comprises obtaining textual data comprising a refinement to the visual search query via the textual input field.
18 . One or more non-transitory computer-readable media that store instructions that, when executed by one or more processors, cause the computing system to perform operations, the operations comprising:
obtaining a visual search query from a user comprising one or more query images;
providing a search interface for display to the user, the search interface comprising one or more result images responsive to the one or more query images and an interface element indicative of a request to the user to refine the visual search query;
obtaining textual data comprising a refinement to the visual search query;
determining, based on using a machine-learned model to process the one or more query images and the textual data, a multimodal search query, wherein the machine-learned model is trained to modify the one or more query images based on the textual data;
retrieving one or more refined search results based on the multimodal search query, wherein a first refined search result of the one or more refined search results comprises an image associated with a particular web page;
processing textual web content and the image associated with the particular web page with the machine-learned model to obtain textual content descriptive of the first refined search result of the one or more refined search results, wherein the textual web content is extracted from the particular web page; and
causing a refined search interface to be displayed for the user, the refined search interface comprising a refined search result interface element that comprises the first refined search result and the textual content descriptive of the first refined search result.