IP Library › Granted Patent US 11,587,243
Granted Patent B2
US 11,587,243 · App. 17/204,788 · Granted Feb 21, 2023

System and method for position tracking using edge computing

Inventors: Jon Andrew Crain (Grapevine, TX); Sailesh Bharathwaaj Krishnamurthy (Irving, TX); Kyle Dalal (Coppell, TX); Shahmeer Ali Mirza (Celina, TX)
Assignee: 7-ELEVEN, INC.
G06T7/292G06Q30/0201G06T7/60G06T7/70H04N5/247G06T2207/10016G06T2207/30196G06T2207/30232
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,587,243
App. No.
17/204,788
Granted
Feb 21, 2023
Kind
B2
Abstract

A tracking system includes a camera subsystem that includes cameras that capture vide of a space. Each camera is coupled with a camera client that determines local coordinates of people in the captured video. The camera clients generate frames that include color frames and depth frames labeled with an identifier number of the camera and their corresponding timestamps. The camera clients generate tracks that include metadata describing historical people detections, tracking identifications, timestamps, and the identifier number of the camera. The camera clients send the frames and tracks to cluster servers that maintain the frames and tracks such that they are retrievable using their corresponding labels. A camera server queries the cluster servers to receive the frames and tracks using their corresponding labels. The camera server determines the physical positions of people in the space based on the determined local coordinates.

Claims (109)

1. A system comprising:

an array of cameras positioned above a space, wherein:

each camera of the array of cameras is operatively coupled with a respective camera client from an array of camera clients; and

each camera of the array of cameras is configured to capture a video of a portion of the space, the space containing a person;

the array of camera clients operably coupled with the array of cameras; wherein:

a first camera client of the array of camera clients is operably coupled with a first camera and configured to:

receive a first plurality of frames of a first video from the first camera, wherein at least one frame of the first plurality of frames shows the person within the space;

generate a timestamp when each frame from the first plurality of frames is received by the first camera client;

for a first frame of the first plurality of frames, label the first frame with one or more first labels comprising an identifier number of the first camera and the timestamp associated with the first frame, such that the first frame is retrievable using the one or more first labels from a first server from among a plurality of cluster servers;

send the first plurality of frames labeled with the one or more first labels to the first server;

generate a first plurality of tracks by performing a local position tracking of the person, wherein at least one track of the first plurality of tracks indicates a location of the person in the at least one frame of the first plurality of frames;

for a first track of the first plurality of tracks, label the first track with one or more second labels comprising the identifier number of the first camera, the timestamp associated with the first track, a particular historical detection associated with the person detected in the first track, and a particular tracking identification detected in the first track, such that the first track is retrievable using the one or more second labels from a second server from among the plurality of cluster servers; and

send the first plurality of tracks labeled with the one or more second labels to the second server, wherein:

the second server is separate from the first server; and

the array of the camera clients is separate from the plurality of cluster servers.

2. The system of claim 1 , wherein:

a second camera client of the array of camera clients is operably coupled with a second camera and separate from the first camera client, the second camera client configured to:

receive a second plurality of frames of a second video from the second camera, wherein at least one frame of the second plurality of frames shows the person within the space;

generate a timestamp when each frame of the second plurality of frames is received by the second camera client;

send the second plurality of frames labeled with one or more corresponding timestamps and an identifier number of the second camera to the first server;

generate a second plurality of tracks by performing a local position tracking of the person; wherein at least one track of the second plurality of tracks indicates a location of the person in the least one frame of the second plurality of frames; and

send the second plurality of tracks labeled with one or more corresponding timestamps and an identifier number of the second camera to the second server.

3. The system of claim 2 , further comprising:

the first server, operably coupled with at least one of the first camera client and the second camera client, the first server configured to:

receive at least one frame from the first and second plurality of frames; and

store the at least one frame such that the at least one frame is retrievable;

the second server, operably coupled with at least one of the first camera client and the second camera client, the second server configured to:

receive at least one track from the first and second plurality of tracks; and

store the at least one track such the at least one track is retrievable.

4. The system of claim 1 , wherein:

the first plurality of frames comprises a first plurality of color frames and a first plurality of depth frames;

the first plurality of color frames corresponds to visual colors of objects in the space; and

the first plurality of depth frames corresponds to distances of objects in the space from the first camera.

5. The system of claim 4 , wherein generating the first plurality of tracks comprises performing the local position tracking of the person in the first plurality of depth frames, and wherein:

for a first depth frame of the first plurality of depth frames, generating a first track of the first plurality of tracks comprises:

detecting a first contour associated with the person;

determining, based on pixel coordinates of the first contour, a first bounding area around the person shown in the first depth frame;

determining, based on the first bounding area, first coordinates of the person in the first depth frame; and

associating a first tracking identification to the person, wherein:

the first tracking identification is linked to historical detections associated with the person, and

the historical detections associated with the person comprise at least one of a contour, a bounding area, and a segmentation mask associated with the person;

for a second depth frame of the first plurality of depth frames, generating a second track of the first plurality of tracks comprises:

detecting a second contour associated with the person;

determining, based on pixel coordinates of the second contour, a second bounding area around the person shown in the second depth frame;

determining, based on the second bounding area, second coordinates of the person in the second depth frame;

determining whether the second bounding area corresponds to the first bounding area; and

in response to determining that the second bounding area corresponds to the first bounding area, associating the first tracking identification to the person.

6. The system of claim 5 , wherein at least one track from the first plurality of tracks is further labeled with the historical detections associated with the person and the first tracking identification associated with the person.

7. A method comprising:

receiving, by a first camera client, a first plurality of frames of a first video from a first camera, wherein at least one frame of the first plurality of frames shows a person within a space;

generating, by the first camera client, a timestamp when each frame from the first plurality of frames is received from the first camera;

for a first frame of the first plurality of frames, labeling the first frame with one or more first labels comprising an identifier number of the first camera and the timestamp associated with the first frame, such that the first frame is retrievable using the one or more first labels from a first server from among a plurality of cluster servers;

sending, by the first camera client, the first plurality of frames labeled with the one or more first labels to the first server;

generating, by the first camera client, a first plurality of tracks by performing a local position tracking of the person, wherein at least one track of the first plurality of tracks indicates a location of the person in the at least one frame of the first plurality of frames;

for a first track of the first plurality of tracks, labeling the first track with one or more second labels comprising the identifier number of the first camera, the timestamp associated with the first track, a particular historical detection associated with the person detected in the first track, and a particular tracking identification detected in the first track, such that the first track is retrievable using the one or more second labels from a second server from among the plurality of cluster servers; and

sending, by the first camera client, the first plurality of tracks labeled with one or more corresponding timestamps and the identifier number of the first camera to the second server from among the plurality of cluster servers, wherein:

the second server is separate from the first server; and

the array of the camera clients is separate from the plurality of cluster servers.

8. The method of claim 7 , further comprising:

receiving, by a second camera client, a second plurality of frames of a second video from a second camera, wherein at least one frame of the second plurality of frames shows the person within the space;

generating, by the second camera client, a timestamp when each frame from the second plurality of frames is received from the second camera;

sending, by the second camera client, the second plurality of frames labeled with one or more corresponding timestamps and an identifier number of the second camera to the first server;

generating, by the second camera client, a second plurality of tracks by performing a local position tracking of the person; wherein at least one track of the second plurality of tracks indicates a location of the person in the at least one frame of the second plurality of frames; and

sending, by the second camera client, the first plurality of tracks labeled with one or more corresponding timestamps and the identifier number of the second camera to the second server.

9. The method of claim 8 , further comprising:

receiving, by the first server, at least one frame from the first and second plurality of frames;

storing, by the first server, the at least one frame such that the at least one frame is retrievable;

receiving, by the second server, at least one track from the first and second plurality of tracks; and

storing, by the second server, the at least one track such that the at least one track is retrievable.

10. The method of claim 9 , wherein:

the first plurality of frames comprises a first plurality of color frames and a first plurality of depth frames;

the first plurality of color frames corresponds to visual colors of objects in the space; and

the first plurality of depth frames corresponds to distances of objects in the space from the first camera.

11. The method of claim 10 , wherein generating the first plurality of tracks comprises performing the local position tracking of the person in the first plurality of depth frames, wherein:

for a first depth frame of the first plurality of depth frames, generating a first track of the first plurality of tracks comprises:

detecting a first contour associated with the person;

determining, based on pixel coordinates of the first contour, a first bounding area around the person shown in the first depth frame;

determining, based on the first bounding area, first coordinates of the person in the first depth frame; and

associating a first tracking identification to the person, wherein:

the first tracking identification is linked to historical detections associated with the person, and

the historical detections associated with the person comprise at least one of a contour, a bounding area, and a segmentation mask associated with the person;

for a second depth frame of the first plurality of depth frames, generating a second track of the first plurality of tracks comprises:

detecting a second contour associated with the person;

determining, based on pixel coordinates of the second contour, a second bounding area around the person shown in the second depth frame;

determining, based on the second bounding area, second coordinates of the person in the second depth frame;

determining whether the second bounding area corresponds to the first bounding area; and

in response to determining that the second bounding area corresponds to the first bounding area, associating the first tracking identification to the person.

12. The method of claim 11 , wherein at least one track from the first plurality of tracks is further labeled with the historical detections associated with the person and the first tracking identification associated with the person.

13. A non-transitory computer-readable medium that stores instructions, wherein when executed by a processor causes the processor to:

receive, by a first camera client, a first plurality of frames of a first video from a first camera, wherein at least one frame of the first plurality of frames shows a person within a space;

generate, by the first camera client, a timestamp when each frame from the first plurality of frames is received from the first camera;

for a first frame of the first plurality of frames, label the first frame with one or more first labels comprising an identifier number of the first camera and the timestamp associated with the first frame, such that the first frame is retrievable using the one or more first labels from a first server from among a plurality of cluster servers;

send, by the first camera client, the first plurality of frames labeled with the one or more first labels to the first server;

generate, by the first camera client, a first plurality of tracks by performing a local position tracking of the person, wherein at least one track of the first plurality of tracks indicates a location of the person in the at least one frame of the first plurality of frames;

for a first track of the first plurality of tracks, label the first track with one or more second labels comprising the identifier number of the first camera, the timestamp associated with the first track, a particular historical detection associated with the person detected in the first track, and a particular tracking identification detected in the first track, such that the first track is retrievable using the one or more second labels from a second server from among the plurality of cluster servers; and

send, by the first camera client, the first plurality of tracks labeled with the one or more second labels to the second server, wherein:

the second server is separate from the first server; and

the array of the camera clients is separate from the plurality of cluster servers.

14. The non-transitory computer-readable medium of claim 13 , wherein the instructions when executed by the processor, further cause the processor to:

receive, by a second camera client, a second plurality of frames of a second video from a second camera, wherein at least one frame of the second plurality of frames shows the person within the space;

generate, by the second camera client, a timestamp when each frame from the second plurality of frames is received from the second camera;

send, by the second camera client, the second plurality of frames labeled with one or more corresponding timestamps and an identifier number of the second camera to the first server;

generate, by the second camera client, a second plurality of tracks by performing a local position tracking of the person; wherein at least one track of the second plurality of tracks indicates a location of the person in the at least one frame of the second plurality of frames; and

send, by the second camera client, the first plurality of tracks labeled with one or more corresponding timestamps and the identifier number of the second camera to the second server.

15. The non-transitory computer-readable medium of claim 14 , wherein the instructions when executed by the processor, further cause the processor to:

receive, by the first server, at least one frame from the first and second plurality of frames;

store, by the first server, the at least one frame such that the at least one frame is retrievable;

receive, by the second server, at least one track from the first and second plurality of tracks; and

store, by the second server, the at least one track such that the at least one track is retrievable.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Mar 17, 2021
From: CRAIN, JON ANDREW; KRISHNAMURTHY, SAILESH BHARATHWAAJ; DALAL, KYLE; MIRZA, SHAHMEER ALI
To: 7-ELEVEN, INC.
Reel/Frame 055629/0808 →
Continuity (48)
Continuation 17104323 · Nov 25, 2020
Continuation In Part 16663633 · Oct 25, 2019
Continuation In Part 17204788
Continuation In Part 16663415 · Oct 25, 2019
Continuation In Part 17018146 · Sep 11, 2020
Division 16663415 · Oct 25, 2019
Division 17204788
Continuation In Part 16991947 · Aug 12, 2020
Continuation 16663669 · Oct 25, 2019
Continuation 17204788
Continuation In Part 16941787 · Jul 29, 2020
Continuation 16663432 · Oct 25, 2019
Continuation 17204788
Continuation In Part 16941825 · Jul 29, 2020
Division 16663432 · Oct 25, 2019
Division 17204788
Continuation In Part 16663710 · Oct 25, 2019
Continuation In Part 16663766 · Oct 25, 2019
Continuation In Part 16663451 · Oct 25, 2019
Continuation In Part 16663794 · Oct 25, 2019
Continuation In Part 16663822 · Oct 25, 2019
Continuation In Part 16941415 · Jul 28, 2020
Continuation 16794057 · Feb 18, 2020
Continuation 16663472 · Oct 25, 2019
Continuation 17204788
Continuation In Part 16663856 · Oct 25, 2019
Continuation In Part 16664160 · Oct 25, 2019
Continuation In Part 17071262 · Oct 15, 2020
Continuation 16857990 · Apr 24, 2020
Continuation 16793998 · Feb 18, 2020
Continuation 16663500 · Oct 25, 2019
Continuation 17204788
Continuation In Part 16857990 · Apr 24, 2020
Continuation 16793998 · Feb 18, 2020
Continuation 16663500 · Oct 25, 2019
Continuation 17204788
Continuation In Part 16664219 · Oct 25, 2019
Continuation In Part 16664269 · Oct 25, 2019
Continuation In Part 16664332 · Oct 25, 2019
Continuation In Part 16884363 · May 27, 2020
Continuation In Part 16664391 · Oct 25, 2019
Continuation In Part 16664426 · Oct 25, 2019
Continuation In Part 16884434 · May 27, 2020
Continuation 16663533 · Oct 5, 2019
Continuation 17204788
Continuation In Part 16663901 · Oct 25, 2019
Continuation In Part 16663948 · Oct 25, 2019
Related Publication 20210201510A1 · Jul 1, 2021