IP Library Granted Patent US 12,299,949
Granted Patent B1
US 12,299,949 · App. 18/202,508 · Granted May 13, 2025

Utilizing sensor data for automated user identification

Inventors: Zheng Tang (Redmond, WA); Prithviraj Banerjee (Redmond, WA); Manoj Aggarwal (Seattle, WA); Gerard Guy Medioni (Los Angeles, CA)
Assignee: Amazon Technologies, Inc.
G06V10/25G06F18/217G06F18/22G06V40/1335G06V40/1347G06V40/1365
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,299,949
App. No.
18/202,508
Granted
May 13, 2025
Kind
B1
Abstract

This disclosure describes a user-recognition system that may perform one or more verification methods upon identifying a previous image that matches a current image of a palm of a user. For instance, the user-recognition system may perform the verification method(s) as part of the recognition method (e.g., after recognizing a matching image), in response to an audit process, in response to a request to re-analyze the image data (e.g., because a user indicates that he or she was not associated with a particular purchase or shopping session), and/or the like.

Claims (120)

1. A system comprising:

one or more processors; and

one or more computer-readable media storing computer-executable instructions that, when executed, cause the one or more processors to perform acts comprising:

receiving first image data;

inputting the first image data into a trained model to determine first coordinates associated with a first portion of interest of the first image data, the trained model configured to identify one or more visually salient portions of user palms;

determining first feature data based at least in part on one or more pixel values associated with the first coordinates;

inputting second image data into the trained model to determine second coordinates associated with a second portion of interest of second image data, the second image data representing a palm of a user;

determining second feature data based at least in part on one or more pixel values associated with the second coordinates;

determining that the second coordinates are within a threshold distance of the first coordinates;

generating data indicating a similarity between the first feature data and the second feature data at least partly in response to determining that the second coordinates are within the threshold distance of the first coordinates; and

determining, using the data, that the first image data represents the palm of the user.

2. The system as recited in claim 1 , wherein the data comprises first data and the one or more computer-readable media further store computer-executable instructions that, when executed, cause the one or more processors to perform acts comprising outputting second data indicating the first portion of interest at the first coordinates of the first image data and the second portion of interest at the second coordinates of the second image data.

3. The system as recited in claim 1 , wherein the data comprises first data, and the one or more computer-readable media further store computer-executable instructions that, when executed, cause the one or more processors to perform an act comprising:

determining third coordinates associated with a third portion of interest of the second image data;

determining third feature data associated based least in part on one or more pixel values associated with the third coordinates;

generating second data indicating a similarity between the first feature data and the third feature data; and

determining, based at least in part on the first data and the second data, that the similarity between the first feature data and the second feature data is greater than the similarity between the first feature data and the third feature data;

and wherein the determining that the first image data represents the palm of the user comprises determining, using the first data, that the first image data represents the palm of the user based at least in part on the determining that the similarity between the first feature data and the second feature data is greater than the similarity between the first feature data and the third feature data.

4. The system as recited in claim 3 , wherein the one or more computer-readable media further store computer-executable instructions that, when executed, cause the one or more processors to perform an act comprising:

determining fourth coordinates associated with a fourth portion of interest of the first image data;

determining fourth feature data associated based least in part on one or more pixel values associated with the fourth coordinates;

generating third data indicating a similarity between the second feature data and the fourth feature data; and

determining, based at least in part on the first data and the third data, that the similarity between the first feature data and the second feature data is greater than the similarity between the second feature data and the fourth feature data;

and wherein the determining that the first image data represents the palm of the user comprises determining, using the first data, that the first image data represents the palm of the user based at least in part on the determining that the similarity between the first feature data and the second feature data is greater than the similarity between the second feature data and the fourth feature data.

5. The system as recited in claim 1 , wherein the data comprises first data, and the one or more computer-readable media further store computer-executable instructions that, when executed, cause the one or more processors to perform an act comprising:

determining third coordinates associated with a third portion of interest of the first image data;

determining third feature data based at least in part on one or more pixel values associated with the third coordinates;

determining fourth coordinates associated with a fourth portion of interest of the second image data;

determining fourth feature data based at least in part on one or more pixel values associated with the fourth coordinates; and

generating second data indicating a similarity between the third feature data and the fourth feature data; and

and wherein the determining that the first image data represents the palm of the user comprises determining, using the first data and the second data, that the first image data represents the palm of the user.

6. The system as recited in claim 1 , wherein the one or more computer-readable media further store computer-executable instructions that, when executed, cause the one or more processors to perform an act comprising:

determining a first confidence value associated with the first feature data;

determining that the first confidence value is greater than a threshold value;

determining third coordinates associated with a third portion of interest of the first image data;

determining third feature data based at least in part on one or more pixel values associated with the third coordinates;

determining a second confidence value associated with the third feature data;

determining that the second confidence value is less than the threshold value; and

determining to refrain from generating data indicating a similarity between the third feature data and feature data associated with the second image data based at least in part on determining that the second confidence value is less than the threshold value.

7. The system as recited in claim 1 , wherein:

the first portion of interest comprises a first pixel of the first image data and at one or more pixels adjacent to the first pixel; and

the second portion of interest comprises a second pixel of the second image data and at one or more pixels adjacent to the second pixel.

8. A method comprising:

receiving first image data;

inputting the first image data into a trained model to determine first coordinates associated with a first portion of interest of the first image data, the trained model configured to identify one or more visually salient portions of user palms;

determining first feature data based at least in part on one or more pixel values associated with the first coordinates;

inputting second image data into the trained model to determine second coordinates associated with a second portion of interest of second image data, the second image data representing a palm of a user;

determining second feature data based at least in part on one or more pixel values associated with the second coordinates;

determining that the second coordinates are within a threshold distance of the first coordinates;

generating data indicating a similarity between the first feature data and the second feature data at least partly in response to determining that the second coordinates are within the threshold distance of the first coordinates; and

determining, using the data, that the first image data represents the palm of the user.

9. The method as recited in claim 8 , wherein the data comprises first data and further comprising and outputting second data indicating the first portion of interest at the first coordinates of the first image data and the second portion of interest at the second coordinates of the second image data.

10. The method as recited in claim 8 , wherein the data comprises first data, and further comprising:

determining third coordinates associated with a third portion of interest of the second image data;

determining third feature data associated based least in part on one or more pixel values associated with the third coordinates;

generating second data indicating a similarity between the first feature data and the third feature data; and

determining, based at least in part on the first data and the second data, that the similarity between the first feature data and the second feature data is greater than the similarity between the first feature data and the third feature data;

and wherein the determining that the first image data represents the palm of the user comprises determining, using the first data, that the first image data represents the palm of the user based at least in part on the determining that the similarity between the first feature data and the second feature data is greater than the similarity between the first feature data and the third feature data.

11. The method as recited in claim 10 , further comprising:

determining fourth coordinates associated with a fourth portion of interest of the first image data;

determining fourth feature data associated based least in part on one or more pixel values associated with the fourth coordinates;

generating third data indicating a similarity between the second feature data and the fourth feature data; and

determining, based at least in part on the first data and the third data, that the similarity between the first feature data and the second feature data is greater than the similarity between the second feature data and the fourth feature data;

and wherein the determining that the first image data represents the palm of the user comprises determining, using the first data, that the first image data represents the palm of the user based at least in part on the determining that the similarity between the first feature data and the second feature data is greater than the similarity between the second feature data and the fourth feature data.

12. The method as recited in claim 8 , wherein the data comprises first data, and further comprising:

determining third coordinates associated with a third portion of interest of the first image data;

determining third feature data based at least in part on one or more pixel values associated with the third coordinates;

determining fourth coordinates associated with a fourth portion of interest of the second image data;

determining fourth feature data based at least in part on one or more pixel values associated with the fourth coordinates; and

generating second data indicating a similarity between the third feature data and the fourth feature data; and

and wherein the determining that the first image data represents the palm of the user comprises determining, using the first data and the second data, that the first image data represents the palm of the user.

13. The method as recited in claim 8 , further comprising:

determining a first confidence value associated with the first feature data;

determining that the first confidence value is greater than a threshold value;

determining third coordinates associated with a third portion of interest of the first image data;

determining third feature data based at least in part on one or more pixel values associated with the third coordinates;

determining a second confidence value associated with the third feature data;

determining that the second confidence value is less than the threshold value; and

determining to refrain from generating data indicating a similarity between the third feature data and feature data associated with the second image data based at least in part on determining that the second confidence value is less than the threshold value.

14. The method as recited in claim 8 , wherein:

the first portion of interest comprises a first pixel of the first image data and at one or more pixels adjacent to the first pixel; and

the second portion of interest comprises a second pixel of the second image data

and at one or more pixels adjacent to the second pixel.

15. One or more non-transitory computer-readable media storing computer-executable instructions that, when executed, cause one or more processors to perform acts comprising:

receiving first image data;

inputting the first image data into a trained model to determine first coordinates associated with a first portion of interest of the first image data, the trained model configured to identify one or more visually salient portions of user palms;

determining first feature data based at least in part on one or more pixel values associated with the first coordinates;

inputting second image data into the trained model to determine second coordinates associated with a second portion of interest of second image data, the second image data representing a palm of a user;

determining second feature data based at least in part on one or more pixel values associated with the second coordinates;

determining that the second coordinates are within a threshold distance of the first coordinates;

generating data indicating a similarity between the first feature data and the second feature data at least partly in response to determining that the second coordinates are within the threshold distance of the first coordinates; and

determining, using the data, that the first image data represents the palm of the user.

16. The one or more non-transitory computer-readable media as recited in claim 15 , wherein the data comprises first data and further storing computer-executable instructions that, when executed, cause the one or more processors to perform acts comprising outputting second data indicating the first portion of interest at the first coordinates of the first image data and the second portion of interest at the second coordinates of the second image data.

17. The one or more non-transitory computer-readable media as recited in claim 15 , wherein the data comprises first data, and further storing computer-executable instructions that, when executed, cause the one or more processors to perform an act comprising:

determining third coordinates associated with a third portion of interest of the second image data;

determining third feature data associated based least in part on one or more pixel values associated with the third coordinates;

generating second data indicating a similarity between the first feature data and the third feature data; and

determining, based at least in part on the first data and the second data, that the similarity between the first feature data and the second feature data is greater than the similarity between the first feature data and the third feature data;

and wherein the determining that the first image data represents the palm of the user comprises determining, using the first data, that the first image data represents the palm of the user based at least in part on the determining that the similarity between the first feature data and the second feature data is greater than the similarity between the first feature data and the third feature data.

18. The one or more non-transitory computer-readable media as recited in claim 17 , further storing computer-executable instructions that, when executed, cause the one or more processors to perform an act comprising:

determining fourth coordinates associated with a fourth portion of interest of the first image data;

determining fourth feature data associated based least in part on one or more pixel values associated with the fourth coordinates;

generating third data indicating a similarity between the second feature data and the fourth feature data; and

determining, based at least in part on the first data and the third data, that the similarity between the first feature data and the second feature data is greater than the similarity between the second feature data and the fourth feature data;

and wherein the determining that the first image data represents the palm of the user comprises determining, using the first data, that the first image data represents the palm of the user based at least in part on the determining that the similarity between the first feature data and the second feature data is greater than the similarity between the second feature data and the fourth feature data.

19. The one or more non-transitory computer-readable media as recited in claim 15 , wherein the data comprises first data, and further storing computer-executable instructions that, when executed, cause the one or more processors to perform an act comprising:

determining third coordinates associated with a third portion of interest of the first image data;

determining third feature data based at least in part on one or more pixel values associated with the third coordinates;

determining fourth coordinates associated with a fourth portion of interest of the second image data;

determining fourth feature data based at least in part on one or more pixel values associated with the fourth coordinates; and

generating second data indicating a similarity between the third feature data and the fourth feature data; and

and wherein the determining that the first image data represents the palm of the user comprises determining, using the first data and the second data, that the first image data represents the palm of the user.

20. The one or more non-transitory computer-readable media as recited in claim 15 , further storing computer-executable instructions that, when executed, cause the one or more processors to perform an act comprising:

determining a first confidence value associated with the first feature data;

determining that the first confidence value is greater than a threshold value;

determining third coordinates associated with a third portion of interest of the first image data;

determining third feature data based at least in part on one or more pixel values associated with the third coordinates;

determining a second confidence value associated with the third feature data;

determining that the second confidence value is less than the threshold value; and

determining to refrain from generating data indicating a similarity between the third feature data and feature data associated with the second image data based at least in part on determining that the second confidence value is less than the threshold value.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded May 26, 2023
From: TANG, ZHENG; BANERJEE, PRITHVIRAJ; AGGARWAL, MANOJ; MEDIONI, GERARD GUY
To: AMAZON TECHNOLOGIES, INC.
Reel/Frame 063773/0528 →
Continuity (1)
Continuation 17209845 · Mar 23, 2021
References Cited (9)
US 6052474A · Nakayama · 2000 [cited by examiner]
US 9117106B2 · Dedeoglu et al. · 2015 [cited by applicant]
US 9235928B2 · Medioni et al. · 2016 [cited by applicant]
US 9473747B2 · Kobres et al. · 2016 [cited by applicant]
US 10127438B1 · Fisher et al. · 2018 [cited by applicant]
US 10133933B1 · Fisher et al. · 2018 [cited by applicant]
US 10728242B2 · LeCun · 2020 [cited by examiner]
US 20130284806A1 · Margalit · 2013 [cited by applicant]
Office Action for U.S. Appl. No. 17/209,845, mailed on Sep. 22, 2022, Zheng Tang, “Utilizing Sensor Data for Automated User Identification”, 26 pages. [cited by applicant]
Cited By (2)
US 12,567,174 US 12,573,088