IP Library Granted Patent US 12,573,237
Granted Patent B1
US 12,573,237 · App. 18/344,519 · Granted Mar 10, 2026

Detecting events by actors using dynamically cropped images

Inventors: Tian Lan (Seattle, WA); Hui Liang (Issaquah, WA); Samuel Nathan Hallman (Marina del Rey, CA); Vijaya Naga Jyoth Sumanth Chennupati (Bothell, WA); Chuhang Zou (Seattle, WA); Hao Pan (Seattle, WA); Suyu Sang (San Jose, CA)
Assignee: Amazon Technologies, Inc.
G06V40/28G06T7/10G06T7/20G06T7/70G06T2207/20132G06T2207/30196
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,573,237
App. No.
18/344,519
Granted
Mar 10, 2026
Kind
B1
Abstract

Systems within materials handling facilities or retail establishments are programmed to receive images from cameras, process clips of the images to generate sets of features representing product spaces and actors depicted within such images, and to classify the clips as depicting or not depicting a shopping event. The images are dynamically cropped to reduce amounts of data that must be processed to in order to determine whether the images depict shopping events. The images are cropped by calculating a center point based on positions of points on product spaces and detected overlaps of hands and the product spaces. Features of clips determined to depict a shopping event are combined into a sequence and transferred, along with classifications of such clips, to a multi-camera system that generates a shopping hypothesis based on such features.

Claims (131)

1 . A system comprising:

a camera comprising at least one processor and at least one memory component, wherein the camera has a field of view including at least a portion of a fixture having a plurality of product spaces within the field of view; and

a computer system in communication with at least the camera, wherein the computer system has positions of at least one point corresponding to each of the plurality of product spaces stored thereon, and wherein the computer system is programmed with one or more sets of instructions that, when executed by the computer system, cause the computer system to execute a method comprising:

receiving a plurality of images from the camera, wherein each of the plurality of images was captured over a period of time;

determining positions of at least one hand of an actor over at least a portion of the period of time;

identifying a first point of an image plane of the camera corresponding to a portion of a first product space;

identifying a second point of the image plane corresponding to a portion of a second product space;

calculating, for each of the plurality of images, a first score representative of an overlap between the positions of the at least one hand and the first point;

calculating, for each of the plurality of images, a second score representative of an overlap between the positions of the at least one hand and the second point;

selecting a crop center for a first clip of images based at least in part on the first scores calculated for the images of the first clip, the second scores calculated for the images of the first clip, the first point and the second point, wherein each of the images of the first clip is one of the plurality of images;

cropping each of the first clip of images by a crop window about the crop center;

providing the cropped first clip of images as inputs to a model, wherein the model comprises:

a feature encoder having a convolutional neural network backbone, an action encoder, a region encoder and a hand encoder;

a feature queue; and

a sequence head comprising an action head, an item head, a hand head and a quantity head;

receiving outputs from the model in response to the inputs;

determining that the actor executed at least one of a taking event, a return event or an event that is neither the taking event nor the return event with an item associated with the product space based at least in part on the outputs received from the model in response to the inputs; and

storing an indication that the actor executed the at least one of the taking event, the return event or the event that is neither the taking event nor the return event in association with the actor to an external system in communication with the camera over one or more networks.

2 . The system of claim 1 , wherein each of the first clip of images comprises a plurality of pixels, and

wherein the method further comprises:

prior to providing the cropped first clip of images as inputs to the model,

stacking each of the cropped first clip of images with channels representing hands and items,

wherein each of the plurality of pixels is represented by:

a plurality of color channels representing a color of one of the plurality of pixels;

a channel indicating whether the one of the plurality of pixels depicts a portion of a hand; and

a channel indicating whether the one of the plurality of pixels depicts an item within the portion of the hand.

3 . The system of claim 1 , wherein an area of each of the plurality of images is approximately twelve times greater than an area of the crop window.

4 . A method comprising:

capturing at least a first plurality of images by a first camera having a first field of view, wherein at least a portion of a fixture comprising a first product space and a second product space is within the first field of view;

determining positions of at least one hand of a first actor over at least a first period of time;

selecting a first center point for cropping at least some of the first plurality of images captured over the first period of time, wherein the first center point is selected based at least in part on a first position corresponding to the first product space, a second position corresponding to the second product space and the positions of the at least one hand of the first actor over the first period of time;

generating a first clip of images, wherein each of the images of the first clip is one of the first plurality of images captured over the first period of time cropped by a window about the first center point;

generating a first hypothesis based at least in part on the first clip of images, wherein the first hypothesis identifies:

a first event, wherein the first event is one of a taking event, a return event or neither a taking event nor a return event;

the first actor; and

a first item associated with one of the first product space or the second product space; and

associating at least a first quantity of the first item with the first actor based at least in part on the first hypothesis.

5 . The method of claim 4 , wherein generating the first hypothesis comprises:

generating a first set of features from the first clip of images, wherein each of the first set of features is one of an action feature, a region feature or a hand feature;

determining a first classification of an event type depicted in the first clip of images; and

generating the first hypothesis based at least in part on the first set of features and the first classification.

6 . The method of claim 5 , further comprising:

determining positions of the at least one hand of the first actor over at least a second period of time;

generating a second clip of images, wherein each of the images of the second clip is one of the first plurality of images captured over the second period of time cropped by the window about the first center point;

generating a second set of features from the second clip of images, wherein each of the second set of features is one of an action feature, a region feature or a hand feature;

determining a second classification of an event type depicted in the second clip of images;

determining that the first classification is consistent with the second classification; and

in response to determining that the first classification is consistent with the second classification,

generating a sequence of features comprising the first set of features and the second set of features,

wherein the first hypothesis is generated based at least in part on the sequence of features, the first classification and the second classification.

7 . The method of claim 6 , wherein determining the first classification of at least the first clip of images comprises:

determining a first plurality of scores based at least in part on the first clip of images, wherein each of the first plurality of scores represents one of a probability that the first clip of images depicts a taking event, a probability that the first clip of images depicts a return event, or a probability that the first clip of images does not depict a taking event or a return event; and

determining the first classification based at least in part on a greatest one of the first plurality of scores,

wherein determining the second classification of at least the second clip of images comprises:

determining a second plurality of scores based at least in part on the second clip of images, wherein each of the second plurality of scores represents one of a probability that the second clip of images depicts a taking event, a probability that the second clip of images depicts a return event, or a probability that the second clip of images does not depict a taking event or a return event; and

determining the second classification based at least in part on a greatest one of the second plurality of scores, and

wherein generating the sequence of features comprises:

concatenating at least the first set of features and the second set of features,

wherein the first hypothesis is generated based at least in part on the concatenated first set of features and second set of features.

8 . The method of claim 5 , wherein generating the first set of features from the first clip of images comprises:

providing each of the images of the first clip as inputs to an encoder comprising a convolutional neural network backbone; and

receiving outputs from the encoder in response to the inputs, wherein each of the outputs is a feature map corresponding to one of the images of the first clip, and

wherein each of the first set of features is generated based at least in part on one of the feature maps.

9 . The method of claim 8 , wherein the encoder further comprises a hand attention module, and

wherein one of the first set of features is an action feature generated based at least in part on the feature maps corresponding to the images of the first clip and the positions of the at least one hand of the first actor over the first period of time.

10 . The method of claim 8 , further comprising:

identifying portions of the first plurality of images depicting at least the first product space and the second product space,

wherein one of the first set of features is a region feature generated based at least in part on the feature maps corresponding to the images of the first clip and the portions of the first plurality of images depicting at least the first product space and the second product space, and

wherein the region feature represents an action performed at one of the first product space or the second product space.

11 . The method of claim 8 , further comprising:

generating a trajectory of the at least one hand based at least in part on the positions of the at least one hand over the first period of time,

wherein one of the first set of features is a hand feature generated based at least in part on the feature maps corresponding to the images of the first clip and the trajectory of the at least one hand.

12 . The method of claim 4 , wherein selecting the first center point comprises:

calculating first scores for each of the first plurality of images captured over the first period of time based at least in part on the positions of the at least one hand of the first actor over the first period of time, wherein each of the first scores is representative of an overlap of the at least one hand and the first product space;

calculating second scores for each of the first plurality of images captured over the first period of time based at least in part on the positions of the at least one hand of the first actor over the first period of time, wherein each of the second scores is representative of an overlap of the at least one hand and the second product space; and

determining the first center point based at least in part on the first scores, the second scores, the first position and the second position.

13 . The method of claim 4 , wherein each of the first plurality of images is six hundred forty pixels by four hundred eighty pixels, and

wherein the window is one hundred sixty pixels by one hundred sixty pixels.

14 . The method of claim 4 , wherein each of the first clip of images comprises a plurality of pixels, and

wherein the method further comprises:

stacking each of the first clip of images with channels representing hands and items,

wherein each of the plurality of pixels is represented by:

a plurality of color channels representing a color of one of the plurality of pixels;

a channel indicating whether the one of the plurality of pixels depicts a portion of a hand; and

a channel indicating whether the one of the plurality of pixels depicts an item within the portion of the hand.

15 . The method of claim 5 , wherein generating the first set of features comprises:

providing each of the first clip of images as inputs to a model comprising:

a feature encoder having a convolutional neural network backbone, an action encoder, a region encoder and a hand encoder;

a feature queue configured to define a sequence of clips based at least in part on features received from the feature encoder; and

a sequence head comprising an action head, an item head, an actor head and a quantity head, wherein the sequence head is configured to generate a hypothesis based on a sequence of clips defined by the feature queue.

16 . The method of claim 4 , further comprising:

determining positions of at least one hand of a second actor over a second period of time;

selecting a second center point for cropping at least some of the first plurality of images captured over the second period of time, wherein the second center point is selected based at least in part on the first position, the second position and the positions of the at least one hand of the second actor over the second period of time;

generating a second clip of images, wherein each of the images of the second clip is one of the first plurality of images captured over the second period of time cropped by a window about the second center point;

generating a second hypothesis based at least in part on the second clip of images, wherein the second hypothesis identifies:

a second event, wherein the second event is one of a taking event, a return event or neither a taking event nor a return event;

the second actor; and

a second item associated with one of the first product space or the second product space; and

associating at least a second quantity of the second item with the second actor based at least in part on the second hypothesis.

17 . The method of claim 4 , further comprising:

capturing at least a second plurality of images by a second camera having a second field of view, wherein at least a portion of the fixture is within the second field of view;

determining positions of the at least one hand of the first actor over at least a second period of time;

selecting a second center point for cropping at least some of the second plurality of images captured over the second period of time, wherein the second center point is selected based at least in part on the first position, the second position and the positions of the at least one hand of the first actor over the second period of time;

generating a second clip of images, wherein each of the images of the second clip is one of the second plurality of images captured over the second period of time cropped by a window about the second center point; and

generating a second hypothesis based at least in part on the second clip of images, wherein the second hypothesis identifies:

the first event or a second event, wherein the second event is one of a taking event, a return event or neither a taking event nor a return event;

the first actor; and

the first item or a second item associated with one of the first product space or the second product space,

wherein the first quantity is associated with the first actor based at least in part on the first hypothesis and the second hypothesis.

18 . The method of claim 17 , further comprising:

determining, by a computer system, that the first actor executed one of the first event or the second based at least in part on the first hypothesis and the second hypothesis.

19 . The method of claim 4 , further comprising:

determining positions of the at least one hand of the first actor over a second period of time;

selecting a second center point for cropping at least some of the first plurality of images captured over the second period of time, wherein the second center point is selected based at least in part on the first position, the second position and the positions of the at least one hand of the first actor over the second period of time; and

generating a second clip of images, wherein each of the images of the second clip is one of the first plurality of images captured over the second period of time cropped by a window about the second center point;

wherein the first hypothesis is generated based at least in part on the first clip of images and the second clip of images.

20 . A computer system comprising at least one processor and at least one data store,

wherein the computer system is in communication with a plurality of cameras, and

wherein the computer system is programmed with one or more sets of instructions that, when executed by the at least one processor, cause the computer system to execute a method comprising:

generating a first sequence of features based at least in part on a first set of images,

wherein each of the images of the first set is cropped from a second set of images about a cropping window defined from a crop center,

wherein each of the images of the second set is captured by a first camera of the plurality of cameras,

wherein each of the images of the first set is a multi-channel image including, for each pixel of such images, a plurality of channels corresponding to color values, a channel corresponding to a mask for a hand, and a channel corresponding to a mask for a product, and

wherein each of the first set of images has been classified as depicting at least one event at one product space by a first model configured to generate a feature map based at least in part on a sequence of features derived from sets of images and a hypothesis of a type of an event and a location of the event based at least in part on a feature map and a first plurality of positional embeddings, and wherein each of the first plurality of positional embeddings corresponds to one of a plurality of product spaces;

generating a second sequence of features based at least in part on a third set of images,

wherein each of the images of the third set is cropped from a fourth set of images about a cropping window defined from a crop center,

wherein each of the images of the fourth set is captured by a second camera,

wherein each of the images of the third set is a multi-channel image including, for each pixel of such images, a plurality of channels corresponding to color values, a channel corresponding to a mask for a hand, and a channel corresponding to a mask for a product, and

wherein each of the third set of images has been classified as depicting at least one event at one product space by a second model configured to generate a feature map based at least in part on a sequence of features derived from sets of images and a hypothesis of a type of an event and a location of the event based at least in part on a feature map and a second plurality of positional embeddings, wherein each of the second plurality of positional embeddings corresponds to one of a plurality of product spaces;

determining that an actor executed at least one of a taking event, a return event or an event that is neither the taking event nor the return event with an item associated with a product space based at least in part on the first sequence of features and the second sequence of features; and

storing information regarding the at least one of the taking event, the return event or the event that is neither the taking event nor the return event in association with the actor in the at least one data store.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jun 29, 2023
From: LAN, TIAN; LIANG, HUI; HALLMAN, SAMUEL NATHAN; CHENNUPATI, VIJAYA NAGA JYOTH SUMANTH; ZOU, CHUHANG; PAN, HAO; SANG, SUYU
To: AMAZON TECHNOLOGIES, INC.
Reel/Frame 064118/0001 →
References Cited (205)
US 6154559A · Beardsley · 2000 [cited by applicant]
US 7050624B2 · Dialameh et al. · 2006 [cited by applicant]
US 7225980B2 · Ku et al. · 2007 [cited by applicant]
US 7949568B2 · Fano et al. · 2011 [cited by applicant]
US 8009863B1 · Sharma et al. · 2011 [cited by applicant]
US 8009864B2 · Linaker et al. · 2011 [cited by applicant]
US 8175925B1 · Rouaix · 2012 [cited by applicant]
US 8189855B2 · Opalach et al. · 2012 [cited by applicant]
US 8285060B2 · Cobb et al. · 2012 [cited by applicant]
US 8369622B1 · Hsu et al. · 2013 [cited by applicant]
US 8423431B1 · Rouaix et al. · 2013 [cited by applicant]
US RE44225E · Aviv · 2013 [cited by applicant]
US 8577705B1 · Baboo et al. · 2013 [cited by applicant]
US 8630924B2 · Groenevelt et al. · 2014 [cited by applicant]
US 8688598B1 · Shakes et al. · 2014 [cited by applicant]
US 8943441B1 · Patrick et al. · 2015 [cited by applicant]
US 9158974B1 · Laska et al. · 2015 [cited by applicant]
US 9160979B1 · Ulmer · 2015 [cited by applicant]
US 9208675B2 · Xu et al. · 2015 [cited by applicant]
US 9336456B2 · Delean · 2016 [cited by applicant]
US 9449233B2 · Taylor · 2016 [cited by applicant]
US 9473747B2 · Kobres et al. · 2016 [cited by applicant]
US 9536177B2 · Chalasani et al. · 2017 [cited by applicant]
US 9582891B2 · Geiger et al. · 2017 [cited by applicant]
US 9727838B2 · Campbell · 2017 [cited by applicant]
US 9846840B1 · Lin et al. · 2017 [cited by applicant]
US 9881221B2 · Bala et al. · 2018 [cited by applicant]
US 9898677B1 · Anđjelkovićet al. · 2018 [cited by applicant]
US 9911290B1 · Zalewski et al. · 2018 [cited by applicant]
US 10055853B1 · Fisher et al. · 2018 [cited by applicant]
US 10133933B1 · Fisher et al. · 2018 [cited by applicant]
US 10147210B1 · Desai et al. · 2018 [cited by applicant]
US 10192415B2 · Heitz et al. · 2019 [cited by applicant]
US 10318917B1 · Goldstein et al. · 2019 [cited by applicant]
US 10354262B1 · Hershey et al. · 2019 [cited by applicant]
US 10438164B1 · Xiong et al. · 2019 [cited by applicant]
US 10474992B2 · Fisher et al. · 2019 [cited by applicant]
US 10510219B1 · Zalewski et al. · 2019 [cited by applicant]
US 10535146B1 · Buibas et al. · 2020 [cited by applicant]
US 10635844B1 · Roose et al. · 2020 [cited by applicant]
US 10699421B1 · Cherevatsky et al. · 2020 [cited by applicant]
US 10839203B1 · Guigues et al. · 2020 [cited by applicant]
US 11030442B1 · Bergamo · 2021 [cited by examiner]
US 11087273B1 · Bergamo · 2021 [cited by applicant]
US 11195146B2 · Fisher et al. · 2021 [cited by applicant]
US 11232294B1 · Banerjee et al. · 2022 [cited by applicant]
US 11270260B2 · Fisher et al. · 2022 [cited by applicant]
US 11284041B1 · Bergamo · 2022 [cited by examiner]
US 11367083B1 · Saurabh et al. · 2022 [cited by applicant]
US 11468681B1 · Kumar · 2022 [cited by examiner]
US 11468698B1 · Kim et al. · 2022 [cited by applicant]
US 11482045B1 · Kim et al. · 2022 [cited by applicant]
US 11538186B2 · Fisher et al. · 2022 [cited by applicant]
US 11734949B1 · Kviatkovsky et al. · 2023 [cited by applicant]
US 12131539B1 · Broaddus · 2024 [cited by examiner]
US 12283201B2 · Adato et al. · 2025 [cited by applicant]
US 20030002712A1 · Steenburgh et al. · 2003 [cited by applicant]
US 20030002717A1 · Hamid · 2003 [cited by applicant]
US 20030107649A1 · Flickner et al. · 2003 [cited by applicant]
US 20030128337A1 · Jaynes et al. · 2003 [cited by applicant]
US 20040181467A1 · Raiyani et al. · 2004 [cited by applicant]
US 20050251347A1 · Perona et al. · 2005 [cited by applicant]
US 20060018516A1 · Masoud et al. · 2006 [cited by applicant]
US 20060061583A1 · Spooner et al. · 2006 [cited by applicant]
US 20060222206A1 · Garoutte · 2006 [cited by applicant]
US 20070092133A1 · Luo · 2007 [cited by applicant]
US 20070156625A1 · Visel · 2007 [cited by applicant]
US 20070182818A1 · Buehler · 2007 [cited by applicant]
US 20070242066A1 · Rosenthal · 2007 [cited by applicant]
US 20070276776A1 · Sagher et al. · 2007 [cited by applicant]
US 20080055087A1 · Horii et al. · 2008 [cited by applicant]
US 20080077511A1 · Zimmerman · 2008 [cited by applicant]
US 20080109114A1 · Orita et al. · 2008 [cited by applicant]
US 20080137989A1 · Ng et al. · 2008 [cited by applicant]
US 20080159634A1 · Sharma et al. · 2008 [cited by applicant]
US 20080166019A1 · Lee · 2008 [cited by applicant]
US 20080193010A1 · Eaton et al. · 2008 [cited by applicant]
US 20080195315A1 · Hu et al. · 2008 [cited by applicant]
US 20090060352A1 · Distante et al. · 2009 [cited by applicant]
US 20090083815A1 · McMaster et al. · 2009 [cited by applicant]
US 20090121017A1 · Cato et al. · 2009 [cited by applicant]
US 20090132371A1 · Strietzel et al. · 2009 [cited by applicant]
US 20090210367A1 · Armstrong et al. · 2009 [cited by applicant]
US 20090245573A1 · Saptharishi et al. · 2009 [cited by applicant]
US 20090276705A1 · Ozdemir et al. · 2009 [cited by applicant]
US 20100002082A1 · Buehler et al. · 2010 [cited by applicant]
US 20100033574A1 · Ran et al. · 2010 [cited by applicant]
US 20110011936A1 · Morandi et al. · 2011 [cited by applicant]
US 20110205022A1 · Cavallaro et al. · 2011 [cited by applicant]
US 20120106800A1 · Khan et al. · 2012 [cited by applicant]
US 20120148103A1 · Hampel et al. · 2012 [cited by applicant]
US 20120159290A1 · Pulsipher et al. · 2012 [cited by applicant]
US 20120257789A1 · Lee et al. · 2012 [cited by applicant]
US 20120284132A1 · Kim et al. · 2012 [cited by applicant]
US 20120327220A1 · Ma · 2012 [cited by applicant]
US 20130076898A1 · Philippe et al. · 2013 [cited by applicant]
US 20130095961A1 · Marty et al. · 2013 [cited by applicant]
US 20130156260A1 · Craig · 2013 [cited by applicant]
US 20130253700A1 · Carson et al. · 2013 [cited by applicant]
US 20130322767A1 · Chao et al. · 2013 [cited by applicant]
US 20140139633A1 · Wang et al. · 2014 [cited by applicant]
US 20140139655A1 · Mimar · 2014 [cited by applicant]
US 20140259056A1 · Grusd · 2014 [cited by applicant]
US 20140279294A1 · Field-Darragh et al. · 2014 [cited by applicant]
US 20140282162A1 · Fein et al. · 2014 [cited by applicant]
US 20140334675A1 · Chu et al. · 2014 [cited by applicant]
US 20140362195A1 · Ng-Thow-Hing et al. · 2014 [cited by applicant]
US 20140362223A1 · LaCroix et al. · 2014 [cited by applicant]
US 20140379296A1 · Nathan et al. · 2014 [cited by applicant]
US 20150019391A1 · Kumar et al. · 2015 [cited by applicant]
US 20150039458A1 · Reid · 2015 [cited by applicant]
US 20150073907A1 · Purves et al. · 2015 [cited by applicant]
US 20150131851A1 · Bernal et al. · 2015 [cited by applicant]
US 20150199824A1 · Kim et al. · 2015 [cited by applicant]
US 20150206188A1 · Tanigawa et al. · 2015 [cited by applicant]
US 20150262116A1 · Katircioglu et al. · 2015 [cited by applicant]
US 20150269143A1 · Park et al. · 2015 [cited by applicant]
US 20150294483A1 · Wells et al. · 2015 [cited by applicant]
US 20160003636A1 · Ng-Thow-Hing et al. · 2016 [cited by applicant]
US 20160012465A1 · Sharp · 2016 [cited by applicant]
US 20160059412A1 · Oleynik · 2016 [cited by applicant]
US 20160125245A1 · Saitwal et al. · 2016 [cited by applicant]
US 20160127641A1 · Gove · 2016 [cited by applicant]
US 20160292881A1 · Bose et al. · 2016 [cited by applicant]
US 20160307335A1 · Perry et al. · 2016 [cited by applicant]
US 20170116473A1 · Sashida et al. · 2017 [cited by applicant]
US 20170206669A1 · Saleemi et al. · 2017 [cited by applicant]
US 20170262994A1 · Kudriashov et al. · 2017 [cited by applicant]
US 20170278255A1 · Shingu et al. · 2017 [cited by applicant]
US 20170309136A1 · Schoner · 2017 [cited by applicant]
US 20170323376A1 · Glaser et al. · 2017 [cited by applicant]
US 20170345165A1 · Stanhill et al. · 2017 [cited by applicant]
US 20170352234A1 · Awaysheh et al. · 2017 [cited by applicant]
US 20170353661A1 · Kawamura · 2017 [cited by applicant]
US 20180025175A1 · Kato · 2018 [cited by applicant]
US 20180070056A1 · DeAngelis et al. · 2018 [cited by applicant]
US 20180084242A1 · Rublee et al. · 2018 [cited by applicant]
US 20180164103A1 · Hill · 2018 [cited by applicant]
US 20180165728A1 · McDonald et al. · 2018 [cited by applicant]
US 20180218515A1 · Terekhov et al. · 2018 [cited by applicant]
US 20180315329A1 · D'Amato et al. · 2018 [cited by applicant]
US 20180343442A1 · Yoshikawa et al. · 2018 [cited by applicant]
US 20190043003A1 · Fisher et al. · 2019 [cited by applicant]
US 20190073627A1 · Nakdimon et al. · 2019 [cited by applicant]
US 20190102044A1 · Wang et al. · 2019 [cited by applicant]
US 20190156273A1 · Fisher et al. · 2019 [cited by applicant]
US 20190156274A1 · Fisher et al. · 2019 [cited by applicant]
US 20190156277A1 · Fisher et al. · 2019 [cited by applicant]
US 20190158801A1 · Matsubayashi · 2019 [cited by applicant]
US 20190236531A1 · Adato et al. · 2019 [cited by applicant]
US 20190315329A1 · Adamski et al. · 2019 [cited by applicant]
US 20200005490A1 · Paik et al. · 2020 [cited by applicant]
US 20200043086A1 · Sorensen · 2020 [cited by applicant]
US 20200090484A1 · Chen et al. · 2020 [cited by applicant]
US 20200134701A1 · Zucker et al. · 2020 [cited by applicant]
US 20200279382A1 · Zhang et al. · 2020 [cited by applicant]
US 20200320287A1 · Porikli et al. · 2020 [cited by applicant]
US 20200381111A1 · Huang et al. · 2020 [cited by applicant]
US 20210019914A1 · Lipchin et al. · 2021 [cited by applicant]
US 20210027485A1 · Zhang · 2021 [cited by applicant]
US 20210124936A1 · Mirza et al. · 2021 [cited by applicant]
US 20210182922A1 · Zheng et al. · 2021 [cited by applicant]
US 20210287013A1 · Carter et al. · 2021 [cited by applicant]
US 20210350555A1 · Fischetti et al. · 2021 [cited by applicant]
US 20210398097A1 · Wu · 2021 [cited by applicant]
US 20220028230A1 · Srinivasan et al. · 2022 [cited by applicant]
US 20220101007A1 · Kadav et al. · 2022 [cited by applicant]
US 20230185386A1 · Bosworth · 2023 [cited by examiner]
US 20230298351A1 · Lee et al. · 2023 [cited by applicant]
US 20250131805A1 · Yamayoshi · 2025 [cited by examiner]
CN 104778690B · 2017 [cited by applicant]
CN 111626681A · 2020 [cited by applicant]
EP 1574986B1 · 2008 [cited by applicant]
JP 2013196199A · 2013 [cited by applicant]
JP 201489626A · 2014 [cited by applicant]
JP 2018207336A · 2018 [cited by applicant]
JP 2019018743A · 2019 [cited by applicant]
JP 2019096996A · 2019 [cited by applicant]
KR 20170006097A · 2017 [cited by applicant]
WO 0021021A1 · 2000 [cited by applicant]
WO 02059836A2 · 2002 [cited by applicant]
WO 2017151241A2 · 2017 [cited by applicant]
Zhang, Z., “A Flexible New Technique for Camera Calibration,” Technical Report MSR-TR-98-71, Dec. 2, 1998, Microsoft Research, Microsoft Corporation, URL: https://www.microsoft.com/en-us/research/wp-content/uploads/2016… [cited by applicant]
Abhaya Asthana et al., “An Indoor Wireless System for Personalized Shopping Assistance”, Proceedings of IEEE Workshop on Mobile Computing Systems and Applications, 1994, pp. 69-74, Publisher: IEEE Computer Society Press. [cited by applicant]
Black, J. et al., “Multi View Image Surveillance and Tracking,” IEEE Proceedings of the Workshop on Motion and Video Computing, 2002, https://www.researchgate.net/publication/4004539_Multi_view_image_surveillance_and_tr… [cited by applicant]
Ciplak G, Telceken S., “Moving Object Tracking Within Surveillance Video Sequences Based on EDContours,” 2015 9th International Conference on Electrical and Electronics Engineering (ELECO), Nov. 26, 2015 (pp. 720-723). … [cited by applicant]
Cristian Pop, “Introduction to the BodyCom Technology”, Microchip AN1391, May 2, 2011, pp. 1-24, vol. AN1391, No. DS01391A, Publisher: 2011 Microchip Technology Inc. [cited by applicant]
Fuentes et al., “People tracking in surveillance applications,” Proceedings 2nd IEEE Int. Workshop on PETS, Kauai, Hawaii, USA, Dec. 9, 2001, 6 pages. [cited by applicant]
Grinciunaite, A., et al., “Human Pose Estimation in Space and Time Using 3D CNN,” ECCV Workshop on Brave New Ideas for Motion Representations in Videos, Oct. 19, 2016, URL: https://arxiv.org/pdf/1609.00036.pdf, 7 pages. [cited by applicant]
Harville, M., “Stereo Person Tracking with Adaptive Plan-View Templates of Height and Occupancy Statistics,” Image and Vision Computing, vol. 22, Issue 2, Feb. 1, 2004, https://www.researchgate.net/publication/223214495… [cited by applicant]
He, K., et al., “Identity Mappings in Deep Residual Networks,” ECCV 2016 Camera-Ready, URL: https://arxiv.org/pdf/1603.05027.pdf, Jul. 25, 2016, 15 pages. [cited by applicant]
Huang, K. S. et al. “Driver's View and Vehicle Surround Estimation Using Omnidirectional Video Stream,” IEEE IV2003 Intelligent Vehicles Symposium. Proceedings (Cal. No.03TH8683), Jun. 9-11, 2003, http://cvrr.ucsd.edu/V… [cited by applicant]
Lee, K. and Kacorri, H., (May 2019), “Hands Holding Clues for Object Recognition in Teachable Machines”, In Proceedings of the 2019 CHI Conference on Human Factors in Computing Systems (pp. 1-12). [cited by applicant]
Liu, C., et al. “Accelerating Vanishing Point-Based Line Sampling Scheme for Real-Time People Localization”, IEEE Transactions on Circuits and Systems for Video Technology. vol 27. No. Mar. 3, 2017 (Year: 2017). [cited by applicant]
Longuet-Higgins, H.C., “A Computer Algorithm for Reconstructing a Scene from Two Projections,” Nature 293, Sep. 10, 1981, https://cseweb.ucsd.edu/classes/fa01/cse291/hclh/SceneReconstruction.pdf, pp. 133-135. [cited by applicant]
Manocha et al., “Object Tracking Techniques for Video Tracking: A Survey,” The International Journal of Engineering and Science (IJES), vol. 3, Issue 6, pp. 25-29, 2014. [cited by applicant]
Phalke K, Hegadi R., “Pixel Based Object Tracking,” 2015 2nd International Conference on Signal Processing and Integrated Networks (SPIN), Feb. 19, 2015 (pp. 575-578). IEEE. [cited by applicant]
Redmon, J., et al., “You Only Look Once: Unified, Real-Time Object Detection,” University of Washington, Allen Institute for AI, Facebook AI Research, URL: https://arxiv.org/pdf/1506.02640.pdf, May 9, 2016, 10 pages. [cited by applicant]
Redmon, Joseph and Ali Farhadi, “YOLO9000: Better, Faster, Stronger,” URL: https://arxiv.org/pdf/1612.08242.pdf, Dec. 25, 2016, 9 pages. [cited by applicant]
Rossi, M. and Bozzoli, E. A., “Tracking and Counting Moving People,” IEEE Int'l Conf. on Image Processing, ICIP-94, Nov. 13-16, 1994, http://citeseerx.ist.psu.edu/viewdoc/download;jsessionid=463D09F419FA5595DBF9DEF30D7E… [cited by applicant]
Sikdar A, Zheng YF, Xuan D., “Robust Object Tracking in the X-Z Domain,” 2016 IEEE International Conference on Multisensor Fusion and Integration for Intelligent Systems (MFI), Sep. 19, 2016 (pp. 499-504). IEEE. [cited by applicant]
Toshev, Alexander and Christian Szegedy, “DeepPose: Human Pose Estimation via Deep Neural Networks,” IEEE Conference on Computer Vision and Pattern Recognition, Aug. 20, 2014, URL: https://arxiv.org/pdf/1312.4659.pdf, 9… [cited by applicant]
Vincze, M., “Robust Tracking of Ellipses at Frame Rate,” Pattern Recognition, vol. 34, Issue 2, Feb. 2001, pp. 487-498. [cited by applicant]
Zhang, Z., “A Flexible New Technique for Camera Calibration,” IEEE Transactions on Pattern Analysis and Machine Intelligence, vol. 22, No. 11, Nov. 2000, 5 pages. [cited by applicant]
Zhang, Z., “A Flexible New Technique for Camera Calibration,” Technical Report MSR-TR-98-71, Microsoft Research, Microsoft Corporation, microsoft.com/en-us/research/wp-content/uploads/2016/02/tr98-71.pdf, 22 pages. [cited by applicant]