IP Library › Granted Patent US 12,033,620
Granted Patent B1
US 12,033,620 · App. 18/463,951 · Granted Jul 9, 2024

Systems and methods for analyzing text extracted from images and performing appropriate transformations on the extracted text

Inventors: Harshit Kharbanda (Pleasanton, CA); Jessica Lee (Brooklyn, NY); Christopher James Kelley (Orinda, CA); Fabian Roth (Zürich, CH); Dounia Berrada (Saratoga, CA); Samer Hassan Hassan (Saratoga, CA); Afroz Mohiuddin (Campbell, CA); Mikhail Khalman (San Francisco, CA); Ali Essam Ali Elqursh (San Jose, CA); Belinda Luna Zeng (Cupertino, CA)
Assignee: GOOGLE LLC
G10L15/183G06F16/5846G06V10/778G06V30/1456G06V30/153G10L15/22G10L15/30
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,033,620
App. No.
18/463,951
Granted
Jul 9, 2024
Kind
B1
Abstract

The present disclosure provides computer-implemented methods, systems, and devices for responding to requests associated with an image. A computing system obtains, wherein the image depicts a first set of textual content. The computing system determines one or more characteristics of the first set of textual content. The computing system determines a response type from a plurality of response types based on the one or more characteristics. The computing system generates a model input, wherein the model input comprises data descriptive of the first set of textual content and a prompt associated with the response type. The computing system provides providing the model input as an input to a machine-learned language model. The computing system receives a second set of text as an output of the machine-learned language model as a result of the machine-learned language model processing the model input. The computing system provides the second set of text for display to a user, wherein the second set of textual content is associated with the response type.

Claims (56)

1. A computing system, the system comprising:

one or more processors; and

one or more non-transitory computer-readable media that collectively store instructions that, when executed by the one or more processors, cause the computing system to perform operations, the operations comprising:

obtaining an image, wherein the image depicts a first set of textual content;

determining one or more characteristics of the first set of textual content, wherein the one or more characteristics of the first set of textual content includes a density of the first set of textual content;

determining a response type from a plurality of response types based on the one or more characteristics, wherein the plurality of response types includes a summarization response, an explanation response, and a query response wherein the determined response type is a summarization response, and wherein determining a response type from a plurality of response types based on the one or more characteristics further comprise:

determining the density for the first set of textual content within the image;

responsive to a determination that the density for the first set of textual content within the image satisfies a threshold, determining that the response type is a summarization response type; and

updating a user interface to include a summarize user interface element;

generating a model input, wherein the model input comprises data descriptive of the first set of textual content and a prompt associated with the response type;

providing the model input as an input to a machine-learned language model;

receiving a second set of text as an output of the machine-learned language model as a result of the machine-learned language model processing the model input; and

providing the second set of text for display to a user, wherein the second set of textual content is associated with the response type.

2. The computing system of claim 1 , wherein determining the density of the first set of textual content within the image further comprises:

determining an area of the image that includes the first set of textual content;

determining a total area of the image; and

determining a percentage of the image that includes the first set of textual content.

3. The computing system of claim 1 , wherein determining the density of the first set of textual content within the image further comprises:

determining a total number of words visible in the image.

4. The computing system of claim 3 , wherein determining the density of the first set of textual content within the image further comprises:

determining a number of words per pixel in the image.

5. The computing system of claim 1 , wherein the computing system generates the model input in response to a user request.

6. The computing system of claim 5 , wherein the user request is input by, a user selecting the summarize user interface element displayed proximate to the image in the user interface of a user computing device.

7. The computing system of claim 1 , wherein the second set of textual content has less textual content than the first set of textual content.

8. The computing system of claim 1 , wherein an optical character recognition process is used to generate text data representing the content of the first set of textual content from the image.

9. The computing system of claim 1 , wherein the machine-learned language model is a large language model.

10. The computing system of claim 1 , the operations further comprising:

updating the user interface to display the second set of textual content.

11. The computing system of claim 1 , wherein the input to the machine-learned model is multimodal.

12. The computing system of claim 1 , wherein the machine-learned language model is operated at a remote server system and the model input is transmitted to the remote server system and the second set of textual content is received from the remote server system.

13. A computer-implemented method for responding to queries about an image, the method comprising:

obtaining, by a computing system with one or more processors, an image, wherein the image depicts a first set of textual content;

determining, by the computing system, one or more characteristics of the first set of textual content, wherein the one or more characteristics of the first set of textual content includes a density of the first set of textual content;

determining, by the computing system, a response type from a plurality of response types based on the one or more characteristics, wherein the plurality of response types includes a summarization response, an explanation response, and a query response, wherein the determined response type is a summarization response, and wherein determining a response type from a plurality of response types based on the one or more characteristics further comprise:

determining the density for the first set of textual content within the image;

responsive to a determination that the density for the first set of textual content within the image satisfies a threshold, determining that the response type is a summarization response type; and

updating a user interface to include a summarize user interface element;

generating, by the computing system, a model input, wherein the model input comprises data descriptive of the first set of textual content and a prompt associated with the response type;

providing, by the computing system, the model input as an input to a machine-learned language model;

receiving, by the computing system, a second set of text as an output of the machine-learned language model as a result of the machine-learned language model processing the model input; and

providing, by the computing system, the second set of text for display to a user, wherein the second set of textual content is associated with the response type.

14. The computer-implemented method of claim 13 , wherein the image includes image content, and the model input includes data descriptive of the image content.

15. The computer-implemented method of claim 13 , wherein the first set of textual content and image content are used as context for responding to the query by the machine-learned language model.

16. The computer-implemented method of claim 13 , wherein the user query is received via voice communication.

17. The computer-implemented method of claim 13 , wherein the user query is received while the image is displayed on the display of a user computing device.

18. One or more non-transitory computer-readable media that collectively store instructions that, when executed by one or more computing devices, cause the one or more computing devices to perform operations, the operations comprising:

obtaining an image, wherein the image depicts a first set of textual content;

determining one or more characteristics of the first set of textual content, wherein the one or more characteristics of the first set of textual content includes a density of the first set of textual content;

determining a response type from a plurality of response types based on the one or more characteristics, wherein the plurality of response types includes a summarization response, an explanation response, and a query response, wherein the determined response type is a summarization response, and wherein determining a response type from a plurality of response types based on the one or more characteristics further comprise:

determining the density for the first set of textual content within the image;

responsive to a determination that the density for the first set of textual content within the image satisfies a threshold, determining that the response type is a summarization response type; and

updating a user interface to include a summarize user interface element;

generating a model input, wherein the model input comprises data descriptive of the first set of textual content and a prompt associated with the response type;

providing the model input as an input to a machine-learned language model;

receiving a second set of text as an output of the machine-learned language model as a result of the machine-learned language model processing the model input; and

providing the second set of text for display to a user, wherein the second set of textual content is associated with the response type.

Assignments (3)
CORRECTIVE ASSIGNMENT TO CORRECT THE CONFIRMATORY ASSIGNMENT OF WORLDWIDE RIGHTS TO REPLACE THE PRIOR FILING DATE PREVIOUSLY RECORDED ON REEL 66348 FRAME 395. ASSIGNOR(S) HEREBY CONFIRMS THE ASSIGNMENT. Recorded Jul 11, 2025
From: ZENG, BELINDA LUNA
To: GOOGLE LLC
Reel/Frame 071889/0966 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Feb 5, 2024
From: ZENG, BELINDA LUNA
To: GOOGLE LLC
Reel/Frame 066348/0395 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Sep 25, 2023
From: KHARBANDA, HARSHIT; LEE, JESSICA; KELLEY, CHRISTOPHER JAMES; ROTH, FABIAN; BERRADA, DOUNIA; HASSAN, SAMER HASSAN; MOHIUDDIN, AFROZ; KHALMAN, MIKHAIL; ELQURSH, ALI ESSAM ALI
To: GOOGLE LLC
Reel/Frame 065005/0918 →
Cited By (2)
US 12,431,131 US 12,525,231