IP Library Granted Patent US 12,347,154
Granted Patent B2
US 12,347,154 · App. 18/601,650 · Granted Jul 1, 2025

Multi-angle object recognition

Inventor: Ibrahim Badr (New York, NY)
Assignee: GOOGLE LLC
G06V10/235G06V10/16G06V10/24G06V10/776G06V10/98H04N23/61H04N23/631H04N23/64H04N23/661H04N5/265H04N23/698
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,347,154
App. No.
18/601,650
Granted
Jul 1, 2025
Kind
B2
Abstract

Methods, systems, and apparatus for controlling smart devices are described. In one aspect a method includes capturing, by a camera on a user device, a plurality of successive images for display in an application environment of an application executing on the user device, performing an object recognition process on the images, the object recognition process including determining that a plurality of images, each depicting a particular object, are required to perform object recognition on the particular object, and in response to the determination, generating a user interface element that indicates a camera operation to be performed, the camera option capturing two or more images, determining that a user, in response to the user interface element, has caused the indicated camera operation to be performed to capture the two or more images, and in response, determining whether a particular object is positively identified from the plurality of images.

Claims (58)

1. A computer-implemented method, comprising:

obtaining, by a user computing device comprising one or more processors, first image data captured in a first camera operation, wherein the first image data includes an object;

in response to a first object recognition process performed with respect to the first image data indicating a failure to recognize the object in the first image data, providing, for presentation to a user of the user computing device, a first user interface element indicating a second camera operation to be performed;

in response to the second camera operation being performed, obtaining, by the user computing device, second image data captured in the second camera operation, wherein the second image data includes the object; and

in response to a second object recognition process performed with respect to the second image data indicating a successful recognition of the object in the second image data, providing, for presentation to the user of the user computing device, a second user interface element indicating the successful recognition of the object.

2. The computer-implemented method of claim 1 , wherein

in the first camera operation the first image data is captured at a first frequency, and

in the second camera operation the second image data is captured at a second frequency, the second frequency being greater than the first frequency.

3. The computer-implemented method of claim 1 , wherein:

in the first camera operation an image of the object is captured from a first position, and

the first user interface element indicating the second camera operation to be performed includes an indication to perform the second camera operation by capturing one or more images of the object from one or more positions different than the first position.

4. The computer-implemented method of claim 1 , wherein:

in the first camera operation an image of the object is captured at a first zoom level, and

the first user interface element indicating the second camera operation to be performed includes an indication to perform the second camera operation by capturing one or more images of the object at one or more zoom levels different than the first zoom level.

5. The computer-implemented method of claim 1 , wherein:

in the first camera operation an image of a first side of the object is captured from a first position, and

the first user interface element indicating the second camera operation to be performed includes an indication to perform the second camera operation by adjusting a position of the object to capture one or more images of a second side of the object from the first position.

6. The computer-implemented method of claim 1 , wherein the object comprises a machine-readable code, a person, a vehicle, or a landmark.

7. The computer-implemented method of claim 1 , further comprising determining the first user interface element based on a type of the object.

8. The computer-implemented method of claim 1 , wherein the first user interface element is provided for presentation to the user on a user interface of the user computing device in real-time while the object is provided for presentation to the user on the user interface of the user computing device.

9. The computer-implemented method of claim 8 , wherein

the first user interface element comprises one or more of an icon, an animation, a video, an audio indication, or a descriptive text, and

the first user interface element is overlaid on the object on the user interface of the user computing device.

10. The computer-implemented method of claim 1 , further comprising determining the second camera operation to be performed and the first user interface element to be presented, based on determining images needed for the successful recognition of the object.

11. The computer-implemented method of claim 1 , further comprising:

transmitting, by the user computing device, the first image data to a server computing system; and

receiving, from the server computing system, a first indication indicating the failure to recognize the object in the first image data based on the first object recognition process.

12. The computer-implemented method of claim 11 , further comprising:

transmitting, by the user computing device, the second image data to the server computing system; and

receiving, from the server computing system, a second indication indicating the successful recognition of the object in the second image data based on the second object recognition process.

13. The computer-implemented method of claim 12 , wherein the second image data is transmitted to the server computing system at a second frequency which is greater than a first frequency at which the first image data is transmitted to the server computing system.

14. The computer-implemented method of claim 12 , further comprising suspending transmission of image data to the server computing system between completion of the first camera operation and performance of the second camera operation.

15. The computer-implemented method of claim 1 , wherein the second user interface element comprises one or more of an identification of the object or a hyperlink to another computing resource which provides content associated with the object.

16. The computer-implemented method of claim 1 , further comprising generating a composite image based on the second image data and the first image data, wherein the second object recognition process is performed with respect to the composite image.

17. The computer-implemented method of claim 1 , further comprising generating a panoramic image based on the second image data and the first image data, wherein the second object recognition process is performed with respect to the panoramic image.

18. The computer-implemented method of claim 1 , further comprising performing, by the user computing device, the second object recognition process with respect to the second image data,

wherein the second image data includes a plurality of image frames and each of the plurality of image frames includes the object,

wherein performing, by the user computing device, the second object recognition process with respect to the second image data comprises:

assigning a respective weight to the object in each of the plurality of image frames,

determining a weighted average for the object based on the respective weight assigned to the object in each of the plurality of image frames, and

determining the successful recognition of the object in the second image data when the weighted average is greater than a weighted average threshold value.

19. A user computing device, comprising:

a camera;

one or more processors; and

one or more non-transitory computer-readable media configured to store instructions that, when executed by the one or more processors, cause the user computing device to perform operations, the operations comprising:

capturing, by the camera in a first camera operation, a first plurality of images each including an object,

in response to a first object recognition process performed with respect to the first plurality of images indicating a failure to recognize the object in the first plurality of images, providing, for presentation to a user of the user computing device, a first user interface element indicating a second camera operation to be performed,

capturing, by the camera in the second camera operation, a second plurality of images each including the object, and

in response to a second object recognition process performed with respect to the second plurality of images indicating a successful recognition of the object in the second plurality of images, providing, for presentation to the user of the user computing device, a second user interface element indicating the successful recognition of the object.

20. A server computing system, comprising:

one or more processors; and

one or more non-transitory computer-readable media configured to store instructions that, when executed by the one or more processors, cause the server computing system to perform operations, the operations comprising:

receiving a first plurality of images captured in a first camera operation, wherein each of the first plurality of images include an object,

performing a first object recognition process with respect to the first plurality of images and providing a first indication to a user computing device indicating a failure to recognize the object in the first plurality of images,

in response to the failure to recognize the object in the first plurality of images, providing, for presentation by the user computing device, a first user interface element indicating a second camera operation to be performed,

receiving a second plurality of images captured in the second camera operation, wherein each of the second plurality of images includes the object,

performing a second object recognition process with respect to the second plurality of images and providing a second indication to the user computing device indicating a successful recognition of the object in the second plurality of images, and

in response to the successful recognition of the object in the second plurality of images, providing, for presentation by the user computing device, a second user interface element indicating the successful recognition of the object.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Apr 4, 2024
From: BADR, IBRAHIM
To: GOOGLE LLC
Reel/Frame 066999/0702 →
Continuity (5)
Continuation 18182737 · Mar 13, 2023
Continuation 17839104 · Jun 13, 2022
Continuation 17062983 · Oct 5, 2020
Continuation 16058575 · Aug 8, 2018
Related Publication 20240346796A1 · Oct 17, 2024
References Cited (24)
US 9836483B1 · Hickman et al. · 2017 [cited by applicant]
US 10269155B1 · Brailovskiy et al. · 2019 [cited by applicant]
US 10540390B1 · Angel et al. · 2020 [cited by applicant]
US 20050189411A1 · Ostrowski et al. · 2005 [cited by applicant]
US 20100253787A1 · Grant · 2010 [cited by applicant]
US 20120086792A1 · Akbarzadeh et al. · 2012 [cited by applicant]
US 20120327174A1 · Hines et al. · 2012 [cited by applicant]
US 20150035947A1 · Skyberg · 2015 [cited by applicant]
US 20150364037A1 · Lee · 2015 [cited by examiner]
US 20160227106A1 · Adachi · 2016 [cited by examiner]
US 20170142320A1 · Ito et al. · 2017 [cited by applicant]
US 20210297589A1 · Tadano · 2021 [cited by examiner]
CN 104854616 · 2015 [cited by applicant]
CN 108756504 · 2018 [cited by applicant]
CN 105025208 · 2018 [cited by applicant]
CN 108320401 · 2021 [cited by applicant]
CN 113923301A · 2022 [cited by examiner]
KR 20160089222 · 2016 [cited by applicant]
WO WO2017080294 · 2017 [cited by applicant]
WO WO2018110002A1 · 2018 [cited by examiner]
International Preliminary Report on Patentability for Application No. PCT/US2019/045083, mailed on Feb. 18, 2021, 8 pages. [cited by applicant]
International Search Report and Written Opinion for PCT/US2019/045083, dated Oct. 23, 2019, 13 pages. [cited by applicant]
Lowe, “Object Recognition from Local Scale-Invariant Features”, 7 [cited by applicant]
Chinese Search Report Corresponding to Application No. 2019800255603 on Jun. 25, 2024. [cited by applicant]