IP Library Granted Patent US 12,001,997
Granted Patent B2
US 12,001,997 · App. 18/350,053 · Granted Jun 4, 2024

Systems and methods for machine vision based object recognition

Inventors: Ujjval Patel (Stamford, CT); Xiaodan Du (Stamford, CT); Lucas McDonald (Stamford, CT)
Assignee: SYNCHRONY BANK
G06Q10/0833G06N3/04G06N3/08G06Q20/12G06Q20/3224G06V10/764G06V10/82G06V20/52G06V40/172G06V40/23
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,001,997
App. No.
18/350,053
Filed
Jul 11, 2023
Granted
Jun 4, 2024
Kind
B2
Art Unit
2645
USPC
382/103
Abstract

The present disclosure is related to object recognition and tracking using multi-camera driven machine vision. In one aspect, a method includes capturing, via a multi-camera system, a plurality of images of a user, each of the plurality of images representing the user from a unique angle; identifying, using the plurality of images, the user; detecting, throughout a facility, an item selected by the user; creating a visual model of the item to track movement of the item throughout the facility; determining, using the visual model, whether the item is selected for purchase; and detecting that the user is leaving the facility; and processing a transaction for the item when the item is selected for purchase and when the user has left the facility.

Claims (59)

1. A computer-implemented method comprising:

receiving one or more images captured by cameras associated with a facility;

applying a machine-learning model to the one or more images to generate an output indicative of whether an item located within the facility has been selected;

detecting a selection of the item based on the generated output;

associating the selected item with a user;

generating a three-dimensional representation of the selected item, wherein the three-dimensional representation is identified from the one or more images;

tracking in real-time a movement of the selected item throughout the facility, wherein tracking includes determining a location of the three-dimensional representation relative to the facility;

determining that the user has left the facility based on the tracked movement of the selected item; and

processing a transaction for the selected item after determining that that the user has left the facility.

2. The computer-implemented method of claim 1 , wherein generating the three-dimensional representation includes modifying the three- dimensional representation into another representation associated with a different dimensional space.

3. The computer-implemented method of claim 1 , wherein the machine-learning model includes a convolutional neural network.

4. The computer-implemented method of claim 1 , wherein the transaction is processed when the tracked movement indicates that the selected item is exiting the facility.

5. The computer-implemented method of claim 1 , further comprising:

associating the selected item with a user;

determining that the user has left the facility, wherein the tracked movement indicates that the selected item continues to remain within the facility; and

determining a deselection of the selected item based on the indication.

6. The computer-implemented method of claim 1 , wherein the machine-learning model is trained using sample images that depict items associated with the facility.

7. The computer-implemented method of claim 1 , wherein generating the three-dimensional representation includes calibrating the one or more images to create a consensus between images that represent the selected item.

8. The computer-implemented method of claim 1 , wherein the machine-learning model is trained based on a loss determined between the generated output and a label provided for the one or more images.

9. A system, comprising:

one or more processors; and

memory storing thereon instructions that, as a result of being executed by the one or more processors, cause the system to perform operations comprising:

receiving one or more images captured by cameras associated with a facility;

applying a machine-learning model to the one or more images to generate an output indicative of whether an item located within the facility has been selected;

detecting a selection of the item based on the generated output;

associating the selected item with a user;

generating a three-dimensional representation of the selected item, wherein the three-dimensional representation is identified from the one or more images;

tracking in real-time a movement of the selected item throughout the facility, wherein tracking includes determining a location of the three-dimensional representation relative to the facility;

determining that the user has left the facility based on the tracked movement of the selected item; and

processing a transaction for the selected item after determining that that the user has left the facility.

10. The system of claim 9 , wherein generating the three-dimensional representation includes modifying the three-dimensional representation into another representation associated with a different dimensional space.

11. The system of claim 9 , wherein the machine-learning model includes a convolutional neural network.

12. The system of claim 9 , wherein the transaction is processed when the tracked movement indicates that the selected item is exiting the facility.

13. The system of claim 9 , wherein the instructions further cause the system to perform operations comprising:

associating the selected item with a user;

determining that the user has left the facility, wherein the tracked movement indicates that the selected item continues to remain within the facility; and

determining a deselection of the selected item based on the indication.

14. The system of claim 9 , wherein the machine-learning model is trained using sample images that depict items associated with the facility.

15. The system of claim 9 , wherein generating the three-dimensional representation includes calibrating the one or more images to create a consensus between images that represent the selected item.

16. The system of claim 9 , wherein the machine-learning model is trained based on a loss determined between the generated output and a label provided for the one or more images.

17. A non-transitory, computer-readable storage medium storing thereon executable instructions that, as a result of being executed by one or more processors of a computer system, cause the computer system to perform operations comprising:

receiving one or more images captured by cameras associated with a facility;

applying a machine-learning model to the one or more images to generate an output indicative of whether an item located within the facility has been selected;

detecting a selection of the item based on the generated output;

associating the selected item with a user;

generating a three-dimensional representation of the selected item, wherein the three-dimensional representation is identified from the one or more images;

tracking in real-time a movement of the selected item throughout the facility, wherein tracking includes determining a location of the three-dimensional representation relative to the facility;

determining that the user has left the facility based on the tracked movement of the selected item; and

processing a transaction for the selected item after determining that that the user has left the facility.

18. The non-transitory, computer-readable storage medium of claim 17 , wherein generating the three-dimensional representation includes modifying the three-dimensional representation into another representation associated with a different dimensional space.

19. The non-transitory, computer-readable storage medium of claim 17 , wherein the machine-learning model includes a convolutional neural network.

20. The non-transitory, computer-readable storage medium of claim 17 , wherein the transaction is processed when the tracked movement indicates that the selected item is exiting the facility.

21. The non-transitory, computer-readable storage medium of claim 17 , wherein the instructions further cause the computer system to perform operations comprising:

associating the selected item with a user;

determining that the user has left the facility, wherein the tracked movement indicates that the selected item continues to remain within the facility; and

determining a deselection of the selected item based on the indication.

22. The non-transitory, computer-readable storage medium of claim 17 , wherein the machine-learning model is trained using sample images that depict items associated with the facility.

23. The non-transitory, computer-readable storage medium of claim 17 , wherein generating the three-dimensional representation includes calibrating the one or more images to create a consensus between images that represent the selected item.

24. The non-transitory, computer-readable storage medium of claim 17 , wherein the machine-learning model is trained based on a loss determined between the generated output and a label provided for the one or more images.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jul 11, 2023
From: PATEL, UJJVAL; DU, XIAODAN; SCHMIDT, ARNOLD
To: SYNCHRONY BANK
Reel/Frame 064208/0613 →
Continuity (4)
Continuation 17378193 · Jul 16, 2021
Continuation 17156207 · Jan 22, 2021
Provisional Application 62965367 · Jan 24, 2020
Related Publication 20240037492A1 · Feb 1, 2024