IP Library Granted Patent US 11,361,468
Granted Patent B2
US 11,361,468 · App. 17/357,867 · Granted Jun 14, 2022

Systems and methods for automated recalibration of sensors for autonomous checkout

Inventors: Nagasrikanth Kallakuri (San Francisco, CA); Tushar Dadlani (Pleasanton, CA); Dhananjay Singh (San Francisco, CA); Daniel L. Fischetti (San Francisco, CA)
Assignee: STANDARD COGNITION, CORP.
G06T7/80G06K9/00214G06K9/00771G06K9/6215G06K9/6232G06K9/6257G06K9/6265G06N3/04G06N3/08G06Q10/087G06T2207/30196
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,361,468
App. No.
17/357,867
Granted
Jun 14, 2022
Kind
B2
Abstract

Systems and techniques are provided for recalibrating cameras in a real space for tracking puts and takes of items by subjects. The method includes first processing one or more selected images selected from a plurality of sequences of images received from a plurality of cameras calibrated using a set of calibration images that were used to calibrate the cameras previously. The first processing includes a process step to extract a plurality of feature descriptors from the images. The first processing also includes a process step to match one or more feature descriptors as extracted from the selected images with feature descriptors extracted from the set of calibration images that were used to calibrate the cameras previously. The feature descriptors correspond to points located at displays or structures that remain substantially immobile.

Claims (62)

1. A method for recalibrating cameras in a real space for tracking puts and takes of items by subjects, the method including:

first processing one or more images selected from a plurality of sequences of images received from a plurality of cameras, in which images in the plurality of sequences of images have respective fields of view in the real space, to:

extract from the images, a plurality of feature descriptors;

match one or more feature descriptors as extracted from the selected images with feature descriptors extracted from a set of calibration images;

calculate based upon feature descriptors as matched, transformation information between the selected images and the set of calibration images;

compare the transformation information as calculated with a threshold; and

update calibration of a camera with the transformation information whenever the transformation information for the camera meets or exceeds the threshold; and

wherein feature descriptors corresponding to points located at displays or structures that remain substantially immobile are extracted using a trained neural network classifier.

2. The method of claim 1 , wherein the trained neural network classifier has been trained using a synthetic shapes dataset created by a second neural network.

3. The method of claim 2 , wherein the second neural network has been trained using a plurality of synthetic shapes having no ambiguity in interest point locations, wherein the synthetic shapes comprise three-dimensional models created automatically, and a plurality of viewpoints generated for the three-dimensional models for matching features; and

wherein three-dimensional models are finetuned by data collected from like real space environments having matching features annotated between different images captured from different viewpoints.

4. The method of claim 1 , wherein feature descriptors corresponding to points located at displays or structures that remain substantially immobile are extracted using a scale invariant feature transform.

5. The method of claim 1 , further including second processing sequences of images of the plurality of sequences of images, to track puts and takes of items by subjects within respective fields of view in the real space; and

wherein first processing and second processing occur substantially contemporaneously, thereby enabling cameras to be calibrated without clearing subjects from the real space or interrupting tracking puts and takes of items by subjects.

6. The method of claim 5 , wherein second processing at least one sequence of images of the plurality of sequences of images to track a take or put event, further includes, detecting the take or put event using a trained neural network.

7. The method of claim 6 , wherein second processing to track puts and takes of items by subjects includes tracking inventory caches involved in an exchange that move over time having locations in three dimensions.

8. The method of claim 7 , wherein locations of the inventory caches include locations corresponding to hands of identified subjects, and wherein processing the sequences of images includes using an image recognition engine to detect an inventory item in hands of a subject identified in the exchange as detected.

9. The method of claim 5 , wherein second processing at least one sequence of images of the plurality of sequences of images to track a take or put event, further including, detecting the take or put event using a trained random forest.

10. The method of claim 1 , further including storing the transformation information and images used to calibrate the cameras in a database.

11. The method of claim 1 , wherein the transformation information is determined relative to an origin point that is selected as a reference point for calibration.

12. The method of claim 1 , wherein the threshold includes a first threshold of at least a 1 centimeter change in camera translation value.

13. The method of claim 1 , wherein the threshold includes a second threshold of at least a 1 degree change in camera rotation value.

14. A system including one or more processors and memory accessible by the processors, the memory loaded with computer instructions recalibrating cameras in a real space for tracking puts and takes of items by subjects between inventory caches which can act as at least one of sources and sinks of inventory items in exchanges of inventory items, which computer instructions, when executed on the processors, implement actions comprising:

first processing one or more images selected from a plurality of sequences of images received from a plurality of cameras, in which images in the plurality of sequences of images have respective fields of view in the real space, to:

extract from the images, a plurality of feature descriptors;

match one or more feature descriptors as extracted from the selected images with feature descriptors extracted from a set of calibration images;

calculate based upon feature descriptors as matched, transformation information between the selected images and the set of calibration images;

compare the transformation information as calculated with a threshold; and

update calibration of a camera with the transformation information whenever the transformation information for the camera meets or exceeds the threshold; and

wherein feature descriptors corresponding to points located at displays or structures that remain substantially immobile are extracted using a trained neural network classifier.

15. The system of claim 14 , wherein the trained neural network classifier has been trained using a synthetic shapes dataset created by a second neural network.

16. The system of claim 15 , wherein the second neural network has been trained using a plurality of synthetic shapes having no ambiguity in interest point locations, wherein the synthetic shapes comprise three-dimensional models created automatically, and a plurality of viewpoints generated for the three-dimensional models for matching features; and

wherein three-dimensional models are finetuned by data collected from like real space environments having matching features annotated between different images captured from different viewpoints.

17. The system of claim 14 , wherein feature descriptors corresponding to points located at displays or structures that remain substantially immobile are extracted using a scale invariant feature transform.

18. The system of claim 14 , further including second processing sequences of images of the plurality of sequences of images, to track puts and takes of items by subjects within respective fields of view in the real space; and

wherein first processing and second processing occur substantially contemporaneously, thereby enabling cameras to be calibrated without clearing subjects from the real space or interrupting tracking puts and takes of items by subjects.

19. The system of claim 18 , wherein second processing at least one sequence of images of the plurality of sequences of images to track a take or put event, further includes, detecting the take or put event using a trained neural network.

20. The system of claim 19 , wherein second processing to track puts and takes of items by subjects includes tracking inventory caches involved in an exchange which move over time having locations in three dimensions.

21. The system of claim 20 , wherein locations of the inventory caches include locations corresponding to hands of identified subjects, and wherein processing the sequences of images includes using an image recognition engine to detect an inventory item in hands of a subject identified in the exchange as detected.

22. The system of claim 18 , wherein second processing at least one sequence of images of the plurality of sequences of images to track a take or put event, further including, detecting the take or put event using a trained random forest.

23. The system of claim 14 , further including storing the transformation information and images used to calibrate the cameras in a database.

24. The system of claim 14 , wherein the transformation information is determined relative to an origin point that is selected as a reference point for calibration.

25. A non-transitory computer readable storage medium impressed with computer program instructions to recalibrating cameras in a real space for tracking puts and takes of items by subjects between inventory caches which can act as at least one of sources and sinks of inventory items in exchanges of inventory items, which computer program instructions when executed implement a method comprising:

first processing one or more images selected from a plurality of sequences of images received from a plurality of cameras calibrated using a set of calibration images that were used to calibrate the cameras previously, in which images in the plurality of sequences of images have respective fields of view in the real space, to:

extract from the images, a plurality of feature descriptors;

match one or more feature descriptors as extracted from the selected images with feature descriptors extracted from the set of calibration images;

calculate based upon feature descriptors as matched, transformation information between the selected images and the set of calibration images;

compare the transformation information as calculated with a threshold; and

update calibration of a camera with the transformation information whenever the transformation information for the camera meets or exceeds the threshold; and

second processing sequences of images of the plurality of sequences of images, to track puts and takes of items by subjects within respective fields of view in the real space;

wherein first processing and second processing occur substantially contemporaneously, thereby enabling cameras to be calibrated without clearing subjects from the real space or interrupting tracking puts and takes of items by subjects.

26. A method for recalibrating cameras in a real space for tracking puts and takes of items by subjects, the method including:

first processing one or more images selected from a plurality of sequences of images received from a plurality of cameras, in which images in the plurality of sequences of images have respective fields of view in the real space, to:

extract from the images, a plurality of feature descriptors;

match one or more extracted feature descriptors as extracted from the selected images with feature descriptors extracted from a set of calibration images;

calculate based upon feature descriptors as matched, transformation information between the selected images and the set of calibration images;

compare the transformation information as calculated with a threshold; and

update calibration of a camera with the transformation information whenever the transformation information for the camera meets or exceeds the threshold; and

second processing sequences of images of the plurality of sequences of images, to track puts and takes of items by subjects within respective fields of view in the real space; and

wherein first processing and second processing occur substantially contemporaneously, thereby enabling cameras to be calibrated without clearing subjects from the real space or interrupting tracking puts and takes of items by subjects.

27. A system including one or more processors and memory accessible by the processors, the memory loaded with computer instructions recalibrating cameras in a real space for tracking puts and takes of items by subjects between inventory caches which can act as at least one of sources and sinks of inventory items in exchanges of inventory items, which computer instructions, when executed on the processors, implement a method according to claim 26 .

28. A non-transitory computer readable storage medium impressed with computer program instructions to recalibrating cameras in a real space for tracking puts and takes of items by subjects between inventory caches which can act as at least one of sources and sinks of inventory items in exchanges of inventory items, which computer program instructions when executed implement a method according to claim 26 .

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jun 25, 2021
From: KALLAKURI, NAGASRIKANTH; DADLANI, TUSHAR; SINGH, DHANANJAY; FISCHETTI, DANIEL L.
To: STANDARD COGNITION, CORP.
Reel/Frame 056672/0845 →
Continuity (2)
Provisional Application 63045007 · Jun 26, 2020
Related Publication 20210407131A1 · Dec 30, 2021
Cited By (5)
US 12,206,979 US 12,231,818 US 12,288,294 US 12,373,971 US 12,737,921