IP Library › Granted Patent US 12,639,368
Granted Patent B2
US 12,639,368 · App. 18/314,663 · Granted May 26, 2026

Visual citations for information provided in response to multimodal queries

Inventors: Harshit Kharbanda (Pleasanton, CA); Jessica Lee (Brooklyn, NY); Christopher James Kelley (Orinda, CA); Belinda Luna Zeng (Cupertino, CA); Louis Wang (San Francisco, CA)
Assignee: GOOGLE LLC
G06F16/583G06V10/761
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,639,368
App. No.
18/314,663
Granted
May 26, 2026
Kind
B2
Abstract

Result images are retrieved based on a similarity to a query image. A set of textual inputs is processed with a machine-learned language model to obtain a language output comprising textual content, wherein the set of textual inputs comprises textual content from source documents that include the result images, and a prompt associated with the query image. The language output and the result images are provided to a user computing device. Information is received descriptive of an indication by a user that a first result image is visually dissimilar to the query image. Textual content associated with the source document that includes the first result image from the set of textual inputs is removed. The set of textual inputs is processed with the machine-learned language model to obtain a refined language output. The refined language output is provided to the user computing device.

Claims (79)

1 . A computer-implemented method, comprising:

retrieving, by a computing system comprising one or more computing devices, two or more result images based on a similarity between an intermediate representation of a query image and intermediate representations of the two or more result images;

processing, by the computing system, a set of textual inputs with a machine-learned language model to obtain a language output comprising textual content, wherein the set of textual inputs comprises textual content from source documents that include the two or more result images, and a prompt associated with the query image;

providing, by the computing system, the language output and the two or more result images to a user computing device for display within an interface of the user computing device;

receiving, by the computing system from the user computing device, information descriptive of an indication by a user of the user computing device that a first result image of the two or more result images is visually dissimilar to the query image;

removing, by the computing system, textual content associated with the source document that includes the first result image from the set of textual inputs;

processing, by the computing system, the set of textual inputs with the machine-learned language model to obtain a refined language output and a degree of certainty associated with the refined language output, wherein the processing the set of textual inputs with the machine-learned language model comprises obtaining attribution information comprising a link to access the source documents;

determining, by the computing system, based on the degree of certainty associated with the refined language output, a format of an interface element that displays the refined language output; and

providing, by the computing system, based on the format of the interface element, the refined language output to the user computing device for display within the interface of the user computing device.

2 . The computer-implemented method of claim 1 , wherein retrieving the two or more result images comprises:

processing, by the computing system, the query image with a machine-learned visual search model to obtain the intermediate representation of the query image; and

retrieving, by the computing system, the result image based on a degree of similarity between the intermediate representation of the query image and intermediate representations of the two or more of result images.

3 . The computer-implemented method of claim 2 , wherein prior to processing the query image, the method comprises obtaining, by the computing system, the query image from the user computing device.

4 . The computer-implemented method of claim 1 , wherein the attribution information comprises identifying information that identifies the source document.

5 . The computer-implemented method of claim 4 , wherein providing the language output and the two or more result images further comprises providing, by the computing system, the attribution information to the user computing device for display within the interface of the user computing device.

6 . The computer-implemented method of claim 4 , providing the language output and the two or more result images to the user computing device for display within the interface of the user computing device comprises:

providing interface data to the user computing device, wherein the interface data comprises instructions to generate:

(a) an interface element comprising the language output; and

(b) two or more selectable attribution elements respectively associated with the two or more result images, wherein each attribution element comprises a thumbnail of the associated result image and the attribution information for one or more source documents that include the associated result image.

7 . The computer-implemented method of claim 6 , wherein receiving the information descriptive of the indication by the user of the user computing device that the first result image of the two or more result images is visually dissimilar to the query image comprises:

receiving, from the user computing device, data indicative of selection of a first selectable attribution element of the two or more selectable attribution elements by the user of the user computing device, wherein the first selectable attribution element is associated with the first result image of the two or more result images.

8 . The computer-implemented method of claim 7 , wherein removing the textual content associated with the source document that includes the first result image from the set of textual inputs further comprises removing information associated with the source document that includes the first result image from the attribution information to obtain refined attribution information; and

wherein providing the refined language output to the user computing device further comprises providing the refined attribution information to the user computing device.

9 . The computer-implemented method of claim 6 , wherein the language output further comprises predictive information that predicts a portion of the language output as being most relevant to the prompt; and

wherein the interface data further comprises instructions to generate an emphasis element that highlights the portion of the language output.

10 . A user computing device, comprising:

one or more processors;

one or more non-transitory computer-readable media that collectively store a first set of instructions that, when executed by the one or more processors, cause the user computing device to perform operations, the operations comprising:

obtaining a query image;

obtaining textual data descriptive of a prompt;

providing the query image and the textual data descriptive of the prompt to a computing system associated with a visual search service;

responsive to providing the query image and the prompt, receiving, from the computing system, (a) two or more result images, (b) a language output from a machine-learned language model that processes the textual data descriptive of the prompt, and (c) a degree of certainty associated with the language output, wherein the language output is generated based on the prompt and textual content from source documents that include the two or more result images, and wherein the generating the language output comprises obtaining attribution information comprising a link to access the source documents;

determining, based on the degree of certainty associated with the language output, a format of an interface element that displays the language output; and

displaying, based on the format of the interface element, within an interface of an application executed by the user computing device:

(a) the interface element comprising the language output; and

(b) two or more selectable attribution elements respectively associated with the two or more result images, wherein each selectable attribution element comprises a thumbnail of the associated result image and the attribution information that identifies a source document that includes the associated result image.

11 . The user computing device of claim 10 , wherein each selectable attribution element comprises a first selectable portion and a second selectable portion.

12 . The user computing device of claim 11 , wherein the operations further comprise:

receiving, from a user via an input device associated with the user computing device, an input that selects the first selectable portion of a first selectable attribution element of the two or more selectable attribution elements.

13 . The user computing device of claim 12 , wherein receiving the input to the first selectable portion of the first selectable attribution element further comprises:

providing, to the computing system, information indicative of selection of the first selectable attribution element; and

responsive to providing the information, receiving, from the computing system, a refined language output, wherein the refined language output is generated based on the prompt and textual content from source documents that include the two or more result images other than a first result image associated with the first selectable attribution element.

14 . The user computing device of claim 13 , wherein the operations further comprise:

displaying, within the interface of the application executed by the user computing device:

(a) an interface element comprising the refined language output; and

(b) one or more selectable attribution elements, wherein the one or more selectable attribution elements comprises each of the two or more selectable attribution elements other than the first selectable attribution element.

15 . The user computing device of claim 12 , wherein the operations further comprise:

receiving, from the user via the input device associated with the user computing device, an input that selects the second selectable portion of the first selectable attribution element of the two or more selectable attribution elements; and

responsive to receiving the input that selects the second selectable portion of the first selectable attribution element, causing display of the source document identified by the attribution information included in the first selectable attribution element.

16 . The user computing device of claim 10 , wherein each of the source documents comprises:

one or more web pages of a web site;

an article;

a newspaper;

a book; or

a transcript.

17 . The user computing device of claim 10 , wherein obtaining the textual data descriptive of the prompt comprises:

obtaining a spoken utterance from the user via an audio capture device associated with the user computing device; and

determining the textual data descriptive of the prompt based at least in part on the spoken utterance.

18 . The user computing device of claim 10 , wherein obtaining the query image comprises:

obtaining an input indicative of a request to capture an image using an image capture device associated with the user computing device; and

responsive to obtaining the input, capturing the query image using the image capture device associated with the user computing device.

19 . One or more non-transitory computer-readable media that collectively store a first set of instructions that, when executed by one or more processors of a user computing device, cause the user computing device to perform operations, the operations comprising:

obtaining a query image;

obtaining textual data descriptive of a prompt;

providing the query image and the textual data descriptive of the prompt to a computing system associated with a visual search service;

responsive to providing the query image and the prompt, receiving, from the computing system, (a) two or more result images, (b) a language output from a machine-learned language model that processes the textual data descriptive of the prompt, and (c) a degree of certainty associated with the language output, wherein the language output is generated based on the prompt and textual content from source documents that include the two or more result images, and wherein the generating the language output comprises obtaining attribution information comprising a link to access the source documents;

determining, based on the degree of certainty associated with the language output, a format of an interface element that displays the language output;

displaying, based on the format of the interface element, within an interface of an application executed by the user computing device:

(a) the interface element comprising the language output; and

(b) two or more selectable attribution elements respectively associated with the two or more result images, wherein each selectable attribution element comprises a thumbnail of the associated result image and the attribution information that identifies a source document that includes the associated result image;

receiving, from a user via an input device associated with the user computing device, an input that selects a first selectable attribution element of the two or more selectable attribution elements;

responsive to receiving the input, providing, to the computing system, information indicative of selection of the first selectable attribution element; and

responsive to providing the information, receiving, from the computing system, a refined language output, wherein the refined language output is generated based on the prompt and textual content from source documents that include the two or more result images other than a first result image associated with the first selectable attribution element.

20 . The one or more non-transitory computer-readable media of claim 19 , wherein each of the source documents comprises:

one or more web pages of a web site;

an article;

a newspaper;

a book; or

a transcript.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded May 25, 2023
From: KHARBANDA, HARSHIT; LEE, JESSICA; KELLEY, CHRISTOPHER JAMES; ZENG, BELINDA LUNA; WANG, LOUIS
To: GOOGLE LLC
Reel/Frame 063758/0736 →
Continuity (1)
Related Publication 20240378237A1 · Nov 14, 2024
References Cited (9)
US 11004131B2 · Kale et al. · 2021 [cited by applicant]
US 11080324B2 · Penta et al. · 2021 [cited by applicant]
US 12175758B2 · Hosoya · 2024 [cited by applicant]
US 20180121768A1 · Lin · 2018 [cited by examiner]
US 20190361983A1 · Wang · 2019 [cited by examiner]
US 20210042662A1 · Pu · 2021 [cited by examiner]
US 20220335585A1 · Spivak · 2022 [cited by examiner]
JP 2012084179 · 2012 [cited by applicant]
JP 6700450 · 2020 [cited by applicant]