IP Library › Granted Patent US 12,561,364
Granted Patent B2
US 12,561,364 · App. 18/158,294 · Granted Feb 24, 2026

Fool-proofing product identification

Inventors: Daniel V. Klein (Pittsburgh, PA); Igor Dos Santos Ramos (Round Rock, TX)
Assignee: Google LLC
G06F16/5854G06F16/532G06F16/5846
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,561,364
App. No.
18/158,294
Granted
Feb 24, 2026
Kind
B2
Abstract

A method includes receiving, from an image capture device in communication with the data processing hardware, image data for an area of interest of a user. The method also includes receiving a query from the user referring to one or more objects detected within the image data and requesting a digital assistant to discern insights associated with the one or more objects referred to by the query. The method also includes processing the query and the image data to: identify, based on context data extracted from the image data, the one or more objects referred to by the query; and determine the insights associated with the identified one or more objects for the digital assistant to discern. The method also includes generating, for output from a user device associated with the user, content indicating the discerned insights associated with the identified one or more objects.

Claims (112)

1 . A computer-implemented method that when executed on data processing hardware causes the data processing hardware to perform operations comprising:

executing a training process to train a content generator on a plurality of training samples to teach the content generator to learn to generate personalized graphical content when particular objects associated with user-defined insight are uniquely identified in image data, each training sample comprising:

a corresponding ground-truth label uniquely identifying one of the particular objects;

a corresponding user-defined insight associated with the particular object; and

data representing the personalized graphical content to generate that is graphically representative of the corresponding user-defined insight associated with the particular object;

receiving, from an image capture device in communication with the data processing hardware, image data for an area of interest of a user;

receiving a query from the user referring to one or more objects detected within the image data and requesting a digital assistant to discern insights associated with the one or more objects referred to by the query, wherein the query refers to, but does not explicitly identify, the one or more objects associated with the insights the digital assistant is requested to discern;

processing the query and the image data to:

identify, based on context data extracted from the image data, the one or more objects referred to by the query;

determine that one of the identified one or more objects corresponds to the particular object of one of the plurality of training samples; and

determine the insights associated with the identified one or more objects for the digital assistant to discern; and

generating, for output from a user device associated with the user, using the trained content generator, the personalized graphical content that is graphically representative of the corresponding user-defined insight associated with the particular object.

2 . The computer-implemented method of claim 1 , wherein:

the context data extracted from the image data comprises a hand of the user recognized within the image data; and

processing the query and the image data to identify the one or more objects comprises processing the query and the image data to identify the one or more objects based on a proximity of the hand of the user recognized within the image data to at least one of the one or more objects detected within the image data.

3 . The computer-implemented method of claim 1 , wherein:

the context data extracted from the image data comprises a point of focus of the image capture device; and

processing the query and the image data to identify the one or more objects referred to by the query comprises processing the query and the image data to identify the one or more objects based on locations of the one or more objects detected within the image data relative to the point of focus of the image capture device.

4 . The computer-implemented method of claim 1 , wherein processing the query and the image data to identify the one or more objects associated with the insights comprises:

performing query interpretation on the received query to identify one or more terms conveying a descriptor of the one or more objects referred to by the query;

extracting visual features from the received image data to obtain object recognition results;

determining an association between one or more of the object recognition results and the descriptor of the one or more objects; and

identifying the one or more objects referred to by the query based on the association between one or more of the object recognition results and the descriptor of the one or more objects.

5 . The computer-implemented method of claim 4 , wherein the operations further comprise:

extracting textual features from the received image data; and

combining the textual features extracted from the received image data with the visual features extracted from the received image data to obtain the object recognition results.

6 . The computer-implemented method of claim 4 , wherein the descriptor conveyed by the one or more terms identified by performing the query interpretation on the received query comprises at least one of:

an object category associated with the one or more objects;

a physical trait associated with the one or more objects; or

a location of the one or more objects relative to reference object in the field of view of the image data.

7 . The computer-implemented method of claim 1 , wherein processing the query and the image data to determine the insights associated with the identified one or more objects for the digital assistant to discern comprises performing query interpretation on the received query to identify a type of the insight for the digital assistant to discern.

8 . The computer-implemented method of claim 7 , wherein the type of insight identified for the digital assistant to discern comprises at least one of:

an insight to uniquely identify a single object;

an insight to identify multiple related objects;

an insight to obtain additional information about an object;

an insight to provide personalized information about an object;

an insight to distinguish between two or more objects; or

an insight to enhance available information.

9 . The computer-implemented method of claim 1 , wherein the operations further comprise, after processing the query and the image data to identify the one or more objects and determine the insights associated with the identified one or more objects for the digital assistant to discern:

performing one or more operations to discern the insights associated with the identified one or more objects; and

generating, for output from a user device associated with the user, content indicating the discerned insights associated with the identified one or more objects based on the one or more operations performed to discern the insights.

10 . The computer-implemented method of claim 9 , wherein performing the one or more operations to discern the insights comprise at least one of:

extracting, from the image data, textual features containing detailed product information associated with at least one of the identified one or more objects;

extracting, from the image data, textual features containing an object identifier that uniquely identifies at least one of the identified one or more objects;

retrieving search results containing product information associated with at least one of the identified one or more objects;

retrieving textual data containing product information associated with at least one of the identified one or more objects, the textual data uploaded by a merchant;

retrieving personal information associated with at least one of the identified one or more objects; or

retrieving custom information associated with at least one of the identified one or more objects.

11 . The computer-implemented method of claim 1 , wherein the operations further comprise:

generating, for output from a user device associated with the user, content indicating the discerned insights associated with the identified one or more objects; and

generating graphical content that indicates the discerned insights wherein the graphical content is superimposed in a graphical user interface displayed on a screen of the user device.

12 . The computer-implemented method of claim 1 , wherein the operations further comprise:

generating, for output from a user device associated with the user, content indicating the discerned insights associated with the identified one or more objects; and

generating audible content indicating the discerned insights, wherein the audible content is audibly output from the user device.

13 . The computer-implemented method of claim 1 , wherein executing the training process to train the content generator on the plurality of training samples further comprises training a visual feature recognizer on the plurality of training samples to teach the visual feature recognizer to learn to uniquely identify particular objects, each training sample further comprising corresponding training image data representing the particular object.

14 . A system comprising:

data processing hardware; and

memory hardware in communication with the data processing hardware, the memory hardware storing instructions that when executed on the data processing hardware cause the data processing hardware to perform operations comprising:

executing a training process to train a content generator on a plurality of training samples to teach the content generator to learn to generate personalized graphical content when particular objects associated with user-defined insight are uniquely identified in image data, each training sample comprising:

a corresponding ground-truth label uniquely identifying one of the particular objects;

a corresponding user-defined insight associated with the particular object; and

data representing the personalized graphical content to generate that is graphically representative of the corresponding user-defined insight associated with the particular object;

receiving, from an image capture device in communication with the data processing hardware, image data for an area of interest of a user;

receiving a query from the user referring to one or more objects detected within the image data and requesting a digital assistant to discern insights associated with the one or more objects referred to by the query, wherein the query refers to, but does not explicitly identify, the one or more objects associated with the insights the digital assistant is requested to discern;

processing the query and the image data to:

identify, based on context data extracted from the image data, the one or more objects referred to by the query;

determine that one of the identified one or more objects corresponds to the particular object of one of the plurality of training samples; and

determine the insights associated with the identified one or more objects for the digital assistant to discern; and

generating, for output from a user device associated with the user, using the trained content generator, the personalized graphical content that is graphically representative of the corresponding user-defined insight associated with the particular object.

15 . The system of claim 14 , wherein:

the context data extracted from the image data comprises a hand of the user recognized within the image data; and

processing the query and the image data to identify the one or more objects comprises processing the query and the image data to identify the one or more objects based on a proximity of the hand of the user recognized within the image data to at least one of the one or more objects detected within the image data.

16 . The system of claim 14 , wherein:

the context data extracted from the image data comprises a point of focus of the image capture device; and

processing the query and the image data to identify the one or more objects referred to by the query comprises processing the query and the image data to identify the one or more objects based on locations of the one or more objects detected within the image data relative to the point of focus of the image capture device.

17 . The system of claim 14 , wherein processing the query and the image data to identify the one or more objects associated with the insights comprises:

performing query interpretation on the received query to identify one or more terms conveying a descriptor of the one or more objects referred to by the query;

extracting visual features from the received image data to obtain object recognition results;

determining an association between one or more of the object recognition results and the descriptor of the one or more objects; and

identifying the one or more objects referred to by the query based on the association between one or more of the object recognition results and the descriptor of the one or more objects.

18 . The system of claim 17 , wherein the operations further comprise:

extracting textual features from the received image data; and

combining the textual features extracted from the received image data with the visual features extracted from the received image data to obtain the object recognition results.

19 . The system of claim 17 , wherein the descriptor conveyed by the one or more terms identified by performing the query interpretation on the received query comprises at least one of:

an object category associated with the one or more objects;

a physical trait associated with the one or more objects; or

a location of the one or more objects relative to reference object in the field of view of the image data.

20 . The system of claim 14 , wherein processing the query and the image data to determine the insights associated with the identified one or more objects for the digital assistant to discern comprises performing query interpretation on the received query to identify a type of the insight for the digital assistant to discern.

21 . The system of claim 20 , wherein the type of insight identified for the digital assistant to discern comprises at least one of:

an insight to uniquely identify a single object;

an insight to identify multiple related objects;

an insight to obtain additional information about an object;

an insight to provide personalized information about an object;

an insight to distinguish between two or more objects; or

an insight to enhance available information.

22 . The system of claim 14 , wherein the operations further comprise, after processing the query and the image data to identify the one or more objects and determine the insights associated with the identified one or more objects for the digital assistant to discern:

performing one or more operations to discern the insights associated with the identified one or more objects; and

generating, for output from a user device associated with the user, content indicating the discerned insights associated with the identified one or more objects based on the one or more operations performed to discern the insights.

23 . The system of claim 22 , wherein performing the one or more operations to discern the insights comprise at least one of:

extracting, from the image data, textual features containing detailed product information associated with at least one of the identified one or more objects;

extracting, from the image data, textual features containing an object identifier that uniquely identifies at least one of the identified one or more objects;

retrieving search results containing product information associated with at least one of the identified one or more objects;

retrieving textual data containing product information associated with at least one of the identified one or more objects, the textual data uploaded by a merchant;

retrieving personal information associated with at least one of the identified one or more objects; or

retrieving custom information associated with at least one of the identified one or more objects.

24 . The system of claim 14 , wherein the operations further comprise:

generating, for output from a user device associated with the user, content indicating the discerned insights associated with the identified one or more objects; and

generating graphical content that indicates the discerned insights wherein the graphical content is superimposed in a graphical user interface displayed on a screen of the user device.

25 . The system of claim 14 , wherein the operations further comprise:

generating, for output from a user device associated with the user, content indicating the discerned insights associated with the identified one or more objects; and

generating audible content indicating the discerned insights, wherein the audible content is audibly output from the user device.

26 . The system of claim 14 , wherein executing the training process to train the content generator on the plurality of training samples further comprises training a visual feature recognizer on the plurality of training samples to teach the visual feature recognizer to learn to uniquely identify particular objects, each training sample further comprising corresponding training image data representing the particular object.

Assignments (2)
CORRECTIVE ASSIGNMENT TO CORRECT THE ADD THE SECOND INVENTOR PREVIOUSLY RECORDED AT REEL: 062457 FRAME: 0036. ASSIGNOR(S) HEREBY CONFIRMS THE ASSIGNMEN. Recorded Jul 1, 2023
From: KLEIN, DANIEL V; RAMOS, IGOR DOS SANTOS
To: GOOGLE LLC
Reel/Frame 064190/0958 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jan 23, 2023
From: KLEIN, DANIEL V
To: GOOGLE LLC
Reel/Frame 062457/0036 →
Continuity (2)
Provisional Application 63267141 · Jan 25, 2022
Related Publication 20230237091A1 · Jul 27, 2023
References Cited (15)
US 20090109243A1 · Kraft et al. · 2009 [cited by applicant]
US 20160091964A1 · Iyer et al. · 2016 [cited by applicant]
US 20180267994A1 · Badr et al. · 2018 [cited by applicant]
US 20180336009A1 · Yoganandan · 2018 [cited by examiner]
US 20190025909A1 · Mittal et al. · 2019 [cited by applicant]
US 20190027147A1 · Diamant · 2019 [cited by examiner]
US 20200021740A1 · Badr et al. · 2020 [cited by applicant]
US 20200302510A1 · Chachek et al. · 2020 [cited by applicant]
US 20210118442A1 · Poddar · 2021 [cited by examiner]
US 20220156312A1 · Nagpal · 2022 [cited by examiner]
JP 2019056956A · 2019 [cited by applicant]
WO 2020033747A1 · 2020 [cited by applicant]
Chen, Chao—“On-device Supermarket Product Recognition”, <https://al.googleblog.com/2020/07/on-device-supermarket-product.html> vice-supermarket-product.html <https://ai.googleblog.com/2020/07/on-device-supermarket-produ… [cited by applicant]
International Search Report and Written Opinion for the related Application No. PCT/US2023/061102, dated Apr. 25, 2023, 53 pages. [cited by applicant]
Japanese Office Action for the related Application No. 2024-544493 dated Sep. 30, 2025. [cited by applicant]