IP Library Granted Patent US 11,615,136
Granted Patent B2
US 11,615,136 · App. 17/157,022 · Granted Mar 28, 2023

System and method of identifying visual objects

Inventors: David Petrou (Brooklyn, NY); Matthew J. Bridges (New Providence, NJ); Shailesh Nalawadi (Morgan Hill, CA); Hartwig Adam (Marina del Rey, CA); Matthew R. Casey (San Francisco, CA); Hartmut Neven (Malibu, CA); Andrew Harp (New York, NY)
Assignee: GOOGLE LLC
G06F16/5838G06F3/048G06F16/50G06F16/5846G06F16/9535G06K9/6271G06V10/10G06V10/56G06V10/96G06V20/20G06V20/63G06V30/142H04N5/225
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,615,136
App. No.
17/157,022
Granted
Mar 28, 2023
Kind
B2
Abstract

A system and method of identifying objects is provided. In one aspect, the system and method includes a hand-held device with a display, camera and processor. As the camera captures images and displays them on the display, the processor compares the information retrieved in connection with one image with information retrieved in connection with subsequent images. The processor uses the result of such comparison to determine the object that is likely to be of greatest interest to the user. The display simultaneously displays the images the images as they are captured, the location of the object in an image, and information retrieved for the object.

Claims (62)

1. A computer-implemented method for mobile recognition and annotation of objects, comprising:

obtaining, by a user device comprising one or more processors, a plurality of image frames from an image sensor of the user device, wherein a subset of the plurality of image frames depicts one or more objects;

determining, by the user device, a primary object of user interest from the one or more objects based at least in part on at least one of a location of the primary object of user interest within the subset of image frames or a number of image frames included in the subset of image frames;

determining, by the user device, that the primary object of user interest comprises an unknown object;

providing, by the user device, one or more image frames of the subset of image frames to an object recognition system;

in response to providing the one or more image frames, obtaining, by the user device from the object recognition system, annotation data descriptive of the primary object of user interest; and

displaying, by the user device on a display device associated with the user device, a user interface element based at least in part on the annotation data.

2. The computer-implemented method of claim 1 , wherein the plurality of image frames comprises video capture data.

3. The computer-implemented method of claim 1 , wherein determining, by the user device, that the primary object of user interest comprises the unknown object comprises:

processing, by the user device, the one or more image frames with a device object recognition process to obtain device recognition data associated with the primary object of user interest; and

determining, by the user device based at least in part on the device recognition data, that the primary object of user interest comprises the unknown object.

4. The computer-implemented method of claim 3 , wherein the device object recognition process comprises a lightweight representation of the object recognition process.

5. The computer-implemented method of claim 3 , wherein the device object recognition process comprises a machine-learned recognition model.

6. The computer-implemented method of claim 1 , wherein the annotation data describes one or more characteristics of the primary object of user interest, and wherein the one or more aspects comprise at least one of:

an identity of the primary object of user interest;

one or more entities associated with the primary object of user interest;

purchase information associated with the primary object of user interest;

one or more separate images depicting the primary object of interest; or

search result data for the primary object of interest.

7. The computer-implemented method of claim 1 , wherein the one or more objects comprise a plurality of faces.

8. The computer-implemented method of claim 1 , wherein the user device comprises the object recognition system.

9. The computer-implemented method of claim 1 , wherein the user interface element comprises at least one of:

a visual indication of the annotation data;

an augmented reality object corresponding to the primary object of user interest; or

a user interface element configured to facilitate purchase of the primary object of user interest.

10. A user device comprising:

one or more processors; and

one or more non-transitory computer-readable media comprising instructions that when executed by the one or more processors cause the one or more processors to perform operations comprising:

obtaining a plurality of image frames from an image sensor of the user device, wherein a subset of the plurality of image frames depicts one or more objects;

determining a primary object of user interest from the one or more objects based at least in part on at least one of a location of the primary object of user interest within the subset of image frames or a number of image frames included in the subset of image frames;

determining that the primary object of user interest comprises an unknown object;

providing one or more image frames of the subset of image frames to an object recognition system;

in response to providing the one or more image frames, obtaining, from the object recognition system, annotation data descriptive of the primary object of user interest; and

displaying, on a display device associated with the user device, a user interface element based at least in part on the annotation data.

11. The user device of claim 10 , wherein the plurality of image frames comprises video capture data.

12. The user device of claim 10 , wherein determining that the primary object of user interest comprises the unknown object comprises:

processing the one or more image frames with a device object recognition process to obtain device recognition data associated with the primary object of user interest; and

determining, based at least in part on the device recognition data, that the primary object of user interest comprises the unknown object.

13. The user device of claim 12 , wherein the device object recognition process comprises a lightweight representation of the object recognition process.

14. The user device of claim 12 , wherein the device object recognition process comprises a machine-learned recognition model.

15. The user device of claim 10 , wherein the annotation data describes one or more characteristics of the primary object of user interest, and wherein the one or more aspects comprise at least one of:

an identity of the primary object of user interest;

one or more entities associated with the primary object of user interest;

purchase information associated with the primary object of user interest;

one or more separate images depicting the primary object of interest; or

search result data for the primary object of interest.

16. The user device of claim 10 , wherein the one or more objects comprise one or more faces.

17. The user device of claim 10 , wherein the user interface element comprises at least one of:

a visual indication of the annotation data;

an augmented reality object corresponding to the primary object of user interest; or

a user interface element configured to facilitate purchase of the primary object of user interest.

18. One or more non-transitory computer-readable media comprising instructions that when executed by one or more processors cause the one or more processors to perform operations comprising:

obtaining a plurality of image frames from an image sensor of a user device, wherein a subset of the plurality of image frames depicts one or more objects;

determining a primary object of user interest from the one or more objects based at least in part on at least one of a location of the primary object of user interest within the subset of image frames or a number of image frames included in the subset of image frames;

determining that the primary object of user interest comprises an unknown object;

providing one or more image frames of the subset of image frames to an object recognition system;

in response to providing the one or more image frames, obtaining, from the object recognition system, annotation data descriptive of the primary object of user interest; and

displaying, on a display device associated with the user device, a user interface element based at least in part on the annotation data.

19. The one or more non-transitory computer-readable media of claim 18 , wherein the plurality of image frames comprises video capture data.

20. The one or more non-transitory computer-readable media of claim 18 , wherein determining that the primary object of user interest comprises the unknown object comprises:

processing the one or more image frames with a device object recognition process to obtain device recognition data associated with the primary object of user interest; and

determining, based at least in part on the device recognition data, that the primary object of user interest comprises the unknown object.

Assignments (2)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jan 26, 2021
From: PETROU, DAVID; BRIDGES, MATTHEW; NALAWADI, SHAILESH; ADAM, HARTWIG; CASEY, MATTHEW R.; NEVEN, HARTMUT; HARP, ANDREW
To: GOOGLE INC.
Reel/Frame 055026/0352 →
CHANGE OF NAME Recorded Jan 26, 2021
From: GOOGLE INC.
To: GOOGLE LLC
Reel/Frame 055107/0497 →
Continuity (8)
Continuation 16744998 · Jan 16, 2020
Continuation 16563375 · Sep 6, 2019
Continuation 16243660 · Jan 9, 2019
Continuation 15247542 · Aug 25, 2016
Continuation 14541437 · Nov 14, 2014
Continuation 13693665 · Dec 4, 2012
Provisional Application 61567611 · Dec 6, 2011
Related Publication 20210141827A1 · May 13, 2021