IP Library Granted Patent US 12,314,313
Granted Patent B2
US 12,314,313 · App. 18/746,969 · Granted May 27, 2025

Image query analysis

Inventors: Gokhan H. Bakir (Zürich, CH); Marcin Bortnik (Zürich, CH); Malte Nuhn (Zürich, CH); Kavin Karthik Ilangovan (Zürich, CH)
Assignee: GOOGLE LLC
G06F16/583G06F16/5866G06F16/90332G06F16/9038G06F18/24G06T7/97G10L13/08G10L15/22G10L2015/221
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,314,313
App. No.
18/746,969
Granted
May 27, 2025
Kind
B2
Abstract

Methods, systems, and apparatus, including computer programs encoded on computer storage media, for analyzing images for generating query responses. One of the methods includes determining, using a textual query, an image category for images responsive to the textual query, and an output type that identifies a type of requested content; selecting, using data that associates a plurality of images with a corresponding category, a subset of the images that each belong to the image category, each image in the plurality of images belonging to one of the two or more categories; analyzing, using the textual query, data for the images in the subset of the images to determine images responsive to the textual query; determining a response to the textual query using the images responsive to the textual query; and providing, using the output type, the response to the textual query for presentation.

Claims (45)

1. A computer-implemented method comprising:

accessing a spoken query received from a user by a microphone associated with a computing device;

determining, using a textual query derived from the spoken query, an image category responsive to the textual query, wherein the image category is determined from two or more image categories defined for a plurality of images obtained by the computing device;

selecting, using data that associates the plurality of images with a corresponding image category of the two or more image categories, a subset of the images that respectively belong to the image category;

analyzing, using the textual query, data for the images in the subset of the images to determine images responsive to the textual query;

selecting, for each image responsive to the textual query, a portion of the image that depicts data responsive to the textual query;

generating instructions for an audible presentation of the data responsive to the textual query; and

causing a speaker associated with the computing device to provide the audible presentation of the data responsive to the textual query.

2. The computer-implemented method of claim 1 , wherein the computing device comprises a mobile user computing device.

3. The computer-implemented method of claim 1 , further comprising converting the spoken query into the textual query.

4. The computer-implemented method of claim 1 , further comprising:

generating instructions for a visual presentation of the data responsive to the textual query; and

causing a display associated with the computing device to provide the visual presentation of the data responsive to the textual query.

5. The computer-implemented method of claim 1 , further comprising:

determining, using one or more of the spoken query or the textual query, an output type that identifies a type of content responsive to the spoken query.

6. The computer-implemented method of claim 5 , wherein the output type comprises at least one of an image, an annotated image, a total cost, or a textual summary.

7. The computer-implemented method of claim 5 , further comprising:

determining, using one or more of the spoken query or the textual query, one or more key phrases for the textual query; and

wherein analyzing, using the textual query, data for the images in the subset of the images to determine images responsive to the textual query comprises analyzing, using the one or more key phrases, data for the images in the subset of the images to determine images responsive to the textual query.

8. The computer-implemented method of claim 1 , further comprising:

determining a location of the portion of the image that depicts the data responsive to the textual query; and

wherein the audible presentation of the data responsive to the textual query indicates the location of the portion of the image that depicts the data responsive to the textual query.

9. The computer-implemented method of claim 1 , wherein the audible presentation of the data responsive to the textual query comprises a prompt indicating that a visual presentation of the data is available.

10. The computer-implemented method of claim 9 , further comprising:

causing a display associated with the computing device to provide the visual presentation of the data responsive to the textual query.

11. The computer-implemented method of claim 1 , wherein the subset of the images that respectively belong to the image category comprises receipts, and wherein the audible presentation of the data responsive to the textual query comprises text from the receipts.

12. The computer-implemented method of claim 1 , wherein the subset of images that respectively belong to the image category comprises restaurant menus, and wherein the audible presentation of the data responsive to the textual query comprises text from the restaurant menus.

13. The computer-implemented method of claim 1 , wherein the subset of the images that respectively belong to the image category comprises documents, and wherein the audible presentation of the data responsive to the textual query comprises text from the documents.

14. A computing system comprising:

one or more processors; and

a memory storing instructions for execution by the one or more processors to cause the computing system to perform operations comprising:

accessing a spoken query received from a user by a microphone associated with a computing device;

determining, using a textual query derived from the spoken query, an image category responsive to the textual query, wherein the image category is determined from two or more image categories defined for a plurality of images obtained by the computing device;

selecting, using data that associates the plurality of images with a corresponding image category of the two or more image categories, a subset of the images that respectively belong to the image category;

analyzing, using the textual query, data for the images in the subset of the images to determine images responsive to the textual query;

selecting, for each image responsive to the textual query, a portion of the image that depicts data responsive to the textual query;

generating instructions for an audible presentation of the data responsive to the textual query; and

causing a speaker associated with the computing device to provide the audible presentation of the data responsive to the textual query.

15. The computing system of claim 14 , wherein the audible presentation of the data responsive to the textual query comprises a prompt indicating that a visual presentation of the data is available.

16. The computing system of claim 15 , the operations further comprising:

causing a display associated with the computing device to provide the visual presentation of the data responsive to the textual query.

17. The computing system of claim 14 , wherein the subset of the images that respectively belong to the image category comprises receipts, and wherein the audible presentation of the data responsive to the textual query comprises text from the receipts.

18. The computing system of claim 14 , wherein the subset of images that respectively belong to the image category comprises restaurant menus, and wherein the audible presentation of the data responsive to the textual query comprises text from the restaurant menus.

19. The computing system of claim 14 , wherein the subset of the images that respectively belong to the image category comprises documents, and wherein the audible presentation of the data responsive to the textual query comprises text from the documents.

20. The computing system of claim 14 , wherein the computing device comprises a mobile user computing device.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Sep 7, 2024
From: BAKIR, GOKHAN H.; BORTNIK, MARCIN; NUHN, MALTE; ILANGOVAN, KAVIN KARTHIK
To: GOOGLE LLC
Reel/Frame 068519/0707 →
Continuity (4)
Continuation 18170091 · Feb 16, 2023
Continuation 16989294 · Aug 10, 2020
Continuation 16114788 · Aug 28, 2018
Related Publication 20240411805A1 · Dec 12, 2024
References Cited (52)
US 5659742A · Beattie et al. · 1997 [cited by applicant]
US 6182066B1 · Marques · 2001 [cited by applicant]
US 7986843B2 · Chaudhury et al. · 2011 [cited by applicant]
US 8391618B1 · Chuang et al. · 2013 [cited by applicant]
US 8897579B2 · Chaudhury et al. · 2014 [cited by applicant]
US 8909625B1 · Stewenius · 2014 [cited by examiner]
US 8983939B1 · Wang et al. · 2015 [cited by applicant]
US 9317534B2 · Brandt · 2016 [cited by applicant]
US 9372920B2 · Bengio et al. · 2016 [cited by applicant]
US 10459971B2 · Fu et al. · 2019 [cited by applicant]
US 10929461B2 · Marriott et al. · 2021 [cited by applicant]
US 20020038299A1 · Zernik · 2002 [cited by examiner]
US 20030123737A1 · Mojsilovic et al. · 2003 [cited by applicant]
US 20070244925A1 · Albouze · 2007 [cited by applicant]
US 20090060351A1 · Li et al. · 2009 [cited by applicant]
US 20120054658A1 · Chuat et al. · 2012 [cited by applicant]
US 20120117051A1 · Liu · 2012 [cited by examiner]
US 20140310255A1 · Cardell et al. · 2014 [cited by applicant]
US 20140358900A1 · Payne et al. · 2014 [cited by applicant]
US 20150161176A1 · Majkowska et al. · 2015 [cited by applicant]
US 20150193528A1 · Bengio et al. · 2015 [cited by applicant]
US 20150213057A1 · Brucher et al. · 2015 [cited by applicant]
US 20150370833A1 · Fey et al. · 2015 [cited by applicant]
US 20160342895A1 · Gao et al. · 2016 [cited by applicant]
US 20170132528A1 · Aslan et al. · 2017 [cited by applicant]
US 20170193011A1 · Kale · 2017 [cited by examiner]
US 20170220907A1 · Liu et al. · 2017 [cited by applicant]
US 20170242875A1 · Jiang et al. · 2017 [cited by applicant]
US 20170255648A1 · Dube et al. · 2017 [cited by applicant]
US 20180232399A1 · Goyal et al. · 2018 [cited by applicant]
US 20190087695A1 · Inoue · 2019 [cited by applicant]
WO WO2017116691 · 2017 [cited by applicant]
Xie, Xing, et al., “Mobile Search with Multimodal Queries”, Proceedings of the IEEE, vol. 96, Issue 4, Apr. 2008, pp. 589-601. [cited by examiner]
Schalkwyk, Johan, et al., “Google Search by Voice: A Case Study”, © 2010, 35 pages. [cited by examiner]
Doungpaisan, Pafan, et al., “Query by Example of Speaker Audio Signals using Power Spectrum and MFCCs”, International Journal of Electrical and Computer Engineering, vol. 7, No. 6, Dec. 2017, pp. 3369-3384. [cited by examiner]
Wang, Ye-Yi, et al., “An Introduction to Voice Search”, IEEE Signal Processing Magazine, vol. 25, Issue 3, May 2008, pp. 29-38. [cited by examiner]
9TO5google, “Google App 7.26 Preps Emails on Home, Reveals Assistant Reservations, Smart Displays UI, more [APK Insight],” https://9to5google.com/2018/04/05/google-app-7-26-apk-insight-teardown, retrieved on Jun. 12, 20… [cited by applicant]
Bank of America, “Erica is Here and Ready to Help”, https://promo.bankofamerica.com/erica/faqs/, retrieved on Jun. 4, 2018, 8 pages. [cited by applicant]
Cai et al., “Automatic Query Type Classification for Web Image Retrieval”, 2007 International Conference on Multimedia and Ubiquitous Engineering, Seoul, South Korea, Apr. 26-28, 2007, pp. 1021-1026. [cited by applicant]
Fergus et al., “Learning Object Categories from Internet Image Searches”, Institute of Electrical and Electronics Engineers, vol. 98, No. 8, Aug. 2010, pp. 1453-1466. [cited by applicant]
Hubdoc, “How it Works,” https://www.hubdoc.com/how-it-works, retrieved on Aug. 27, 2018, 8 pages. [cited by applicant]
Iskandar et al., “Content-based Image Retrieval Using Image Regions as Query Examples”, Nineteenth Australasian Database Conference, Gold Coast, Australia, Dec. 2-4, 2007, 9 pages. [cited by applicant]
Koskela et al., “Use of Image Subset Features in Image Retrieval with Self-Organizing Maps.”, Conference on Image and Video Retrieval 2004, LNCS 3115, Springer-Verlag, Berlin, Germany, 2004, pp. 508-516. [cited by applicant]
Mukherjea et al., “AMORE: A World Wide Web Image Retrieval Engine.”, World Wide Web, vol. 2, 1999, pp. 115-132. [cited by applicant]
Noce, Lucia, et al., “Embedded Textual Content for Document Image Classification with Convolutional Neural Networks”, 2016 Association for Computing Machinery Symposium on Document Engineering, Vienna, Austria, Sep. 13-… [cited by applicant]
Ogle et al., “Chabot: Retrieval from a Relational Database of Images.”, Computer, vol. 28, Issue 9, Sep. 1995, pp. 40-48. [cited by applicant]
Pocket-Lint, “What is Google Lens, How Does it Work, and Which Devices Have It?”, https://www.pocket-lint.com/apps/news/google/141075-whatOis-google-lens-and-how-does-it-work-and-which-devices-have-it, retrieved on Jun.… [cited by applicant]
Srihari et al., “Intelligent Indexing and Semantic Retrieval of Multimodal Documents”, Information Retrieval, vol. 2, May 2000, pp. 245-275. [cited by applicant]
Vijayanarasimhan et al., “Keywords to Visual Categories: Multiple-Instance Learning for Weakly Supervised Object Categorization.”, 2008 Conference on Computer Vision and Pattern Recognition, Anchorage, Alaska, United St… [cited by applicant]
Wang et al., “Annotating Images by Mining Image Search Results.”, Institute of Electrical and Electronics Engineers Transactions on Pattern Analysis and Machine Intelligence, vol. 30, No. 11, Nov. 2008, pp. 1919-1932. [cited by applicant]
Wang et al., “IGroup: Presenting Web Image Search Results in Semantic Clusters.”, Conference on Human Factors in Computing 2007—Web Usability, San Jose, California, United States, Apr. 28-May 3, 2007, pp. 587-596. [cited by applicant]
Wang et al., “Web Image Re-Ranking Using Query-Specific Semantic Signatures.”, Institute of Electrical and Electronics Engineers Transactions on Pattern Analysis and Machine Intelligence, vol. 36, No. 4, Apr. 2014, pp. … [cited by applicant]