IP Library Granted Patent US 10,740,400
Granted Patent B2
US 10,740,400 · App. 16/114,788 · Granted Aug 11, 2020

Image analysis for results of textual image queries

Inventors: Gokhan H. Bakir (Zürich, CH); Marcin Bortnik (Zürich, CH); Malte Nuhn (Zürich, CH); Kavin Karthik Ilangovan (Zürich, CH)
Assignee: Google LLC
G06F16/90332G06F16/5866G06F16/9038G06K9/6267G06T7/97G10L13/08G10L15/22G10L2015/221
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 10,740,400
App. No.
16/114,788
Granted
Aug 11, 2020
Kind
B2
Abstract

Methods, systems, and apparatus, including computer programs encoded on computer storage media, for analyzing images for generating query responses. One of the methods includes determining, using a textual query, an image category for images responsive to the textual query, and an output type that identifies a type of requested content; selecting, using data that associates a plurality of images with a corresponding category, a subset of the images that each belong to the image category, each image in the plurality of images belonging to one of the two or more categories; analyzing, using the textual query, data for the images in the subset of the images to determine images responsive to the textual query; determining a response to the textual query using the images responsive to the textual query; and providing, using the output type, the response to the textual query for presentation.

Claims (80)

1. A system comprising one or more computers and one or more storage devices on which are stored instructions that are operable, when executed by the one or more computers, to cause the one or more computers to perform operations comprising:

determining, using a textual query, an image category for images responsive to the textual query, and an output type that identifies a type of requested content;

selecting, using data that associates a plurality of images with a corresponding category, a subset of the images that each belong to the image category, each image in the plurality of images belonging to one of two or more categories;

analyzing, using the textual query, data for the images in the subset of the images to determine images responsive to the textual query;

determining a response to the textual query using the images responsive to the textual query; and

providing, using the output type, the response to the textual query for presentation.

2. The system of claim 1 , wherein providing, using the output type, the response to the textual query for presentation comprises:

generating instructions for an audible presentation of the data responsive to the textual query; and

providing the instructions to a speaker to cause the speaker to provide the audible presentation of the data responsive to the textual query.

3. The system of claim 1 , wherein:

determining the response to the textual query using the images responsive to the textual query comprises selecting, for each image responsive to the textual query using the output type, a portion of the image that depicts data responsive to the textual query; and

providing, using the output type, the response to the textual query for presentation comprises:

generating instructions for presentation of a user interface that emphasizes, for each image responsive to the textual query, the portion of the image that depicts the data responsive to the textual query; and

providing the instructions to a display to cause the display to present the user interface and at least one of the images responsive to the textual query.

4. The system of claim 3 , wherein selecting the subset of the images comprises selecting the subset of the images using the output type and the image category.

5. The system of claim 3 , wherein selecting the portion of the image that depicts data responsive to the textual query comprises:

determining a bounding box for the image that surrounds the data responsive to the textual query; and

selecting the portion of the image defined by the bounding box.

6. The system of claim 3 , wherein selecting the portion of the image that depicts data responsive to the textual query comprises cropping, for at least one of the images responsive to the textual query, the image to remove content that is not responsive to the textual query.

7. The system of claim 6 , wherein cropping the image to remove content that is not responsive to the textual query comprises cropping the image so that the data responsive to the textual query comprises a fixed size or a percent of the cropped image.

8. The system of claim 7 , the operations comprising:

determining the percent of the cropped image using context depicted in the image.

9. The system of claim 8 , wherein determining the percent of the cropped image using context depicted in the image comprises determine the percent of the cropped image using at least one of the data responsive to the query depicted in the image, text depicted in the image, or a boundary of an object depicted in the image.

10. The system of claim 3 , wherein generating the instructions for presentation of the user interface comprises:

determining an output format using a quantity of the images responsive to the textual query or the output type or both; and

generating the instructions for presentation of the user interface using the output format.

11. The system of claim 10 , wherein determining the output format comprises:

determining that a single image from an image database depicts data responsive to the textual query; and

in response to determining that a single image from the image database depicts data responsive to the textual query, selecting an output format that depicts, in the user interface, only data from the image.

12. The system of claim 10 , wherein:

determining the output format comprises:

determining that multiple images from the plurality of images depict data responsive to the textual query; and

in response to determining that multiple images from the plurality of images depict data responsive to the textual query, selecting a summary output format that depicts, in the user interface, a) a summary of the data responsive to the textual query from the multiple images and b) data from each of the multiple images; and

generating the instructions for presentation of the user interface using the output format comprises generating the instructions for presentation of the user interface that includes a) the summary of the data responsive to the textual query and b) the data from each of the multiple images.

13. The system of claim 12 , wherein the summary output format includes the summary above the data for each of the multiple images.

14. The system of claim 12 , wherein the summary comprises a list of the data responsive to the textual query from the multiple images.

15. The system of claim 12 , wherein the user interface comprises a navigation control that enables a user to scroll through presentation of the data from each of the multiple images.

16. The system of claim 3 , wherein providing the instructions to a display to cause the display to present the user interface and at least one of the images responsive to the textual query comprises providing the instructions to a display to cause the display to present an answer to the textual query in the user interface.

17. The system of claim 1 , comprising:

for each of two or more images in the plurality of images:

analyzing image data for the image using object recognition to determine an initial image category for the image from the two or more categories; and

determining whether the initial image category is included in a particular group of image categories;

for at least one image from the two or more images for which the initial image category is included in the particular group of image categories:

determining to use the initial image category as the image category for the image;

for at least one image from the two or more images for which the initial image category is not included in the particular group of image categories:

analyzing the image data for the image using text recognition to determine a second image category for the image from the two or more categories; and

determining the image category for the image using the initial image category and the second image category; and

storing, for each of the two or more images, data in a database that associates the image with the image category for the image.

18. The system of claim 17 , the operations comprising:

for each of the two or more images:

receiving the image data before the image data is stored in an image database; and

storing the image data in the image database, wherein analyzing the image data is responsive to receiving the image data.

19. The system of claim 1 , the operations comprising determining, using the textual query, one or more key phrases for the textual query, wherein analyzing, using the textual query, data for the images in the subset of the images to determine images responsive to the textual query comprises analyzing, using the one or more key phrases, data for the images in the subset of the images to determine images responsive to the textual query.

20. A non-transitory computer storage medium encoded with instructions that, when executed by one or more computers, cause the one or more computers to perform operations comprising:

for each of two or more images in a plurality of images:

analyzing image data for the image using object recognition to determine an initial image category for the image from two or more categories; and

determining whether the initial image category is included in a particular group of image categories;

for at least one image from the two or more images for which the initial image category is included in the particular group of image categories:

determining to use the initial image category as the image category for the image;

for at least one image from the two or more images for which the initial image category is not included in the particular group of image categories:

analyzing the image data for the image using text recognition to determine a second image category for the image from the two or more categories; and

determining the image category for the image using the initial image category and the second image category; and

storing, for each of the two or more images, data in a database that associates the image with the image category for the image.

21. The computer storage medium of claim 20 , the operations comprising:

for each of the two or more images:

receiving the image data before the image data is stored in an image database; and

storing the image data in the image database, wherein analyzing the image data is responsive to receiving the image data.

22. A computer-implemented method comprising:

determining, using a textual query, an image category for images responsive to the textual query, and an output type that identifies a type of requested content;

selecting, using data that associates a plurality of images with a corresponding category, a subset of the images that each belong to the image category, each image in the plurality of images belonging to one of two or more categories;

analyzing, using the textual query, data for the images in the subset of the images to determine images responsive to the textual query;

selecting, for each image responsive to the textual query using the output type, a portion of the image that depicts data responsive to the textual query;

generating instructions for an audible presentation of the data responsive to the textual query; and

providing the instructions to a speaker to cause the speaker to provide the audible presentation of the data responsive to the textual query.

23. The method of claim 22 , wherein:

generating the instructions comprises generating, for at least one of the images responsive to the textual query, instructions for an audible presentation of the data responsive to the textual query and that indicates a location of the portion of the image that depicts the data responsive to the query; and

providing the instructions comprises providing the instructions to the speaker to cause the speaker to provide, for the at least one of the images responsive to the textual query, the audible presentation of the data responsive to the textual query and that indicates a location of the portion of the image that depicts the data responsive to the query.

24. The method of claim 22 , comprising:

generating instructions for presentation of a user interface that emphasizes, for each image responsive to the textual query, the portion of the image that depicts the data responsive to the textual query; and

providing the instructions to a display to cause the display to present the user interface and at least one of the images responsive to the textual query.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Aug 28, 2018
From: BAKIR, GOKHAN H.; BORTNIK, MARCIN; NUHN, MALTE; ILANGOVAN, KAVIN KARTHIK
To: GOOGLE LLC
Reel/Frame 046728/0739 →
Continuity (1)
Related Publication 20200074014A1 · Mar 5, 2020