IP Library › Granted Patent US 12,592,096
Granted Patent B1
US 12,592,096 · App. 17/038,490 · Granted Mar 31, 2026

Modeling and detecting shopping events using visual images and machine learning

Inventors: Michael Dillon (Seattle, WA); Gerard Guy Medioni (Seattle, WA); Ali Rahimi (Berkeley, CA)
Assignee: Amazon Technologies, Inc.
G06V40/107G06F18/28G06V20/52G06V40/28H04N7/18
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,592,096
App. No.
17/038,490
Granted
Mar 31, 2026
Kind
B1
Abstract

Cameras installed at a facility capture images over periods of time and detect features such as locations of hands, heads or other body parts within the images. Time series of data regarding such features are provided to machine learning systems to determine whether an event such as a taking or a return of the item is depicted within the images, or whether no such event is detected. The time series are provided to the machine learning system in an iterative manner, with time series associated with a set of baseline features provided to the machine learning system first. If an event is not detected and associated with an actor based on the baseline features, time series associated with supplemental features are successively provided to the machine learning system, along with the time series associated with the baseline features, until the event is detected and associated with an actor.

Claims (146)

1 . A method comprising:

capturing at least a first plurality of images by a first imaging device having a first field of view, wherein the first plurality of images are captured over a first period of time;

detecting, by the first imaging device, locations of at least a first body part of a first actor depicted within at least a subset of the first plurality of images;

detecting, by the first imaging device, locations of at least a second body part of a second actor depicted within the subset of the first plurality of images;

determining, by the first imaging device, first values of data corresponding to at least a first feature of the first actor based at least in part on at least one of the locations of at least the first body part of the first actor depicted within at least the subset of the first plurality of images;

determining, by the first imaging device, second values of data corresponding to at least a second feature of the second actor based at least in part on at least one of the locations of at least the second body part of the second actor depicted within at least the subset of the first plurality of images;

constructing, by the first imaging device, a first time series of data, wherein the first time series of data comprises the first values of data corresponding to the first feature and statuses of the first body part of the first actor at times at which each of the subset of the first plurality of images was captured;

constructing, by the first imaging device, a second time series of data, wherein the second time series of data comprises the second values of data corresponding to the second feature and statuses of the second body part of the second actor at the times at which each of the subset of the first plurality of images was captured;

providing at least a portion of the first time series of data as a first input to a first machine learning system, wherein the first machine learning system is trained to determine whether at least one actor is associated with at least one of a plurality of events based at least in part on a feature of the at least one actor;

receiving at least a first output from the first machine learning system in response to the first input;

providing at least a portion of the second time series of data as a second input to the first machine learning system;

receiving at least a second output from the first machine learning system in response to the second input; and

determining, based at least in part on the first output and the second output, that the first actor is associated with a first event of the plurality of events,

wherein the first event comprises one of:

placing at least one item on a storage unit within the first field of view; or

removing at least one item from the storage unit.

2 . The method of claim 1 , wherein receiving at least the first output from the first machine learning system in response to the first input comprises:

receiving at least the first output and a third output from the first machine learning system in response to the first input, and

wherein the method further comprises:

identifying the at least one item associated with the first event based at least in part on the third output.

3 . The method of claim 2 , wherein identifying the least one item associated with the first event based at least in part on the third output comprises:

identifying a location on the storage unit based at least in part on the third output; and

identifying at least one of an item or a type of item provided on the location based at least in part on planogram data for the storage unit, wherein the planogram data identifies locations of items on the storage unit,

wherein the at least one item comprises the item or one of the type of the item provided on the location.

4 . The method of claim 1 , further comprising:

prior to capturing at least the first plurality of images,

training the first machine learning system, wherein training the first machine learning system comprises:

identifying at least a second plurality of images captured by one of the first imaging device or a second imaging device having a second field of view, wherein the second plurality of images are captured over a second period of time;

determining that a third actor depicted within at least some of the second plurality of images is associated with a second event that occurred within at least one of the first field of view or the second field of view during the second period of time, wherein the second event is one of the plurality of events;

detecting locations of at least a third body part of the third actor depicted within at least a subset of the second plurality of images;

determining third values of data corresponding to at least a third feature of the third actor based at least in part on at least one of the locations of at least the third body part of the third actor depicted within at least the subset of the second plurality of images;

constructing a third time series of data, wherein the third time series of data comprises the third values of data corresponding to the third feature and statuses of the third body part of the third actor at times at which each of at least the subset of the second plurality of images was captured;

providing at least a portion of the third time series of data as a third input to the first machine learning system; and

receiving at least a third output from the first machine learning system in response to the third input,

wherein the first machine learning system is trained based at least in part on the third input and the third output.

5 . The method of claim 1 , wherein the first feature is one of:

a trajectory of a hand of the first actor over the first period of time, wherein the first body part is the hand of the first actor;

a trajectory of a head of the first actor over the first period of time, wherein the first body part is the head of the first actor;

a trajectory of a shoulder of the first actor over the first period of time, wherein the first body part is the shoulder of the first actor; or

distances between a location of the first body part and a location of at least one item over the first period of time.

6 . The method of claim 1 , wherein each of the first values of data is one of:

a location of a hand of the first actor depicted within one of the images of the subset of the first plurality of images at a time when the one of the images was captured;

a location of a head of the first actor depicted within one of the images of the subset of the first plurality of images at a time when the one of the images was captured;

a location of a shoulder of the first actor depicted within one of the images of the subset of the first plurality of images at a time when the one of the images was captured;

a location of at least one item on the storage unit; a location of one hand of the first actor depicted within one of the images of the subset of the first plurality of images at a time when the one of the images was captured; or

an identifier of at least one item within the hand of the first actor depicted within one of the images of the subset of the first plurality of images at a time when the one of the images was captured.

7 . The method of claim 1 , wherein the first time series of data further comprises third values of data corresponding to a third feature of the first actor and the times at which each of the subset of the first plurality of images was captured, and

wherein providing at least the portion of the first time series of data as the first input to the first machine learning system comprises:

providing the first third values of data corresponding to the first third feature and the times at which each of the subset of the first plurality of images was captured as a third input to the first machine learning system;

receiving at least a third output from the first machine learning system in response to the third input; and

determining, based at least in part on the third output, that the first actor may not be associated with the at least one of the plurality of events based on the first third values of data.

8 . The method of claim 7 , wherein each of the first values of data is a location of one hand of the first actor depicted within one of the subset of the first plurality of images,

wherein each of the third values of data is a status of the one hand of the first actor depicted within the at least some of the first plurality of images, and

wherein the status is one of empty or full.

9 . The method of claim 1 , wherein the first time series of data further comprises third values of data corresponding to a third feature of the first actor and third fourth values corresponding to a third fourth feature of the first actor, and

wherein the method further comprises:

prior to providing at least the portion of the first time series of data as the first input to the first machine learning system,

providing the first values of data corresponding to the first feature and the times at which each of the subset of the first plurality of images was captured as a third input to the first machine learning system;

receiving at least a third output from the first machine learning system in response to the third input;

determining, based at least in part on the third output, that the first actor may not be associated with the at least one of the plurality of events based on the first values of data;

in response to determining that the first actor may not be associated with the at least one of the plurality of events based on the first values of data,

providing the first values of data corresponding to the first feature, the third values of data corresponding to the third feature, and the times at which each of the subset of the first plurality of images was captured as a fourth input to the first machine learning system;

receiving at least a fourth output from the first machine learning system in response to the fourth input; and

determining, based at least in part on the fourth output, that the first actor may not be associated with the at least one of the plurality of events based on the first values of data and the third values of data,

wherein the first time series of data is provided as the first input to the first machine learning system in response to determining that the first actor may not be associated with the at least one of the plurality of events based on the first values of data and the third values of data.

10 . The method of claim 1 , wherein the first machine learning system operates on at least one of:

at least one computer system in communication with the first imaging device;

a first processor unit provided on the first imaging device, wherein the first imaging device is in communication with a second imaging device; or

a second processor unit of the second imaging device.

11 . The method of claim 1 , wherein detecting the locations of at least the first body part depicted within at least the subset of the first plurality of images comprises:

providing each of the first plurality of images as inputs to an algorithm configured to detect at least the first body part within an image; and

receiving outputs from the algorithm, wherein each of the outputs is received in response to one of the inputs,

wherein each of the locations of at least the first body part is detected based at least in part on one of the outputs.

12 . The method of claim 1 , wherein the first body part is one of:

a hand of the first actor;

a head of the first actor; or

a portion of an arm of the first actor.

13 . The method of claim 1 , wherein determining that the first actor is associated with the first event comprises at least one of:

determining that the first actor has extended at least the first body part toward the storage unit based at least in part on the first output; or

determining that the first actor has extended at least the first body part into one of a bag, a basket, a cart or a pocket based at least in part on the first output.

14 . A system comprising:

a camera having a field of view including at least a portion of at least one surface for accommodating one or more items, wherein the camera comprises a processor unit and an optical sensor; and

a computer system in communication with the camera, wherein the computer system is configured to execute a machine learning system trained to determine whether at least one actor is associated with at least one of a plurality of events based at least in part on a feature of the at least one actor detected in at least one image,

wherein the processor unit is programmed with one or more sets of instructions that, when executed by the processor unit, cause the camera to execute a first method comprising:

capturing a plurality of images over a period of time;

detecting locations of at least a first body part of a first actor depicted within at least a subset of the plurality of images;

detecting locations of at least a second body part of a second actor depicted within the subset of the plurality of images;

determining first values of data corresponding to at least a first feature of the first actor based at least in part on at least one of the locations of at least the first body part of the first actor depicted within at least the subset of the plurality of images;

determining second values of data corresponding to at least a second feature of the second actor based at least in part on at least one of the locations of at least the second body part of the second actor depicted within at least the subset of the first plurality of images;

constructing, by the first imaging device, a first time series of data, wherein the first time series of data comprises the first values of data corresponding to the first feature and statuses of the first body part of the first actor at times at which each of the subset of the plurality of images was captured;

constructing, by the first imaging device, a second time series of data, wherein the second time series of data comprises the second values of data corresponding to the second feature and statuses of the second body part of the second actor at the times at which each of the subset of the plurality of images was captured; and

transmitting at least the first time series of data and the second time series of data to the computer system,

wherein the computer system is programmed with one or more sets of instructions that, when executed by the processor unit, cause the computer system to execute a second method comprising:

providing at least a portion of the first time series of data as a first input to the machine learning system;

receiving at least a first output from the machine learning system in response to the first input;

providing at least a portion of the second time series of data as a second input to the machine learning system;

receiving at least a second output from the machine learning system in response to the second input; and

determining, based at least in part on the first output and the second output, that the first actor is associated with one of the plurality of events,

wherein the one of the plurality of events comprises one of:

placing at least one item on the at least one surface; or

removing at least one item from the at least one surface.

15 . The camera of claim 14 , wherein the first feature is one of:

a trajectory of a hand of the first actor over the period of time, wherein the first body part is the hand of the first actor;

a trajectory of a head of the first actor over the period of time, wherein the first body part is the head of the first actor; or

a trajectory of a shoulder of the first actor over the period of time, wherein the first body part is the shoulder of the first actor, and

wherein the second feature is one of:

a trajectory of a hand of the second actor over the period of time, wherein the second body part is the hand of the second actor;

a trajectory of a head of the second actor over the period of time, wherein the second body part is the head of the second actor; or

a trajectory of a shoulder of the second actor over the period of time, wherein the second body part is the shoulder of the second actor.

16 . The system of claim 14 , wherein each of the first values of data is a location of the first body part depicted within one of the subset of the plurality of images,

wherein each one of the statuses of the first body part is one of empty or full,

wherein each one of the second values of data is a location of the second body part depicted within one of the plurality of images, and

wherein each one of the statuses of the second body part is one of empty or full.

17 . The system of claim 14 , wherein determining that the first actor is associated with the one of the plurality of events comprises at least one of:

determining that the first actor has extended at least the first body part toward the at least one surface based at least in part on the first output; or

determining that the first actor has extended at least the first body part into one of a bag, a basket, a cart or a pocket based at least in part on the first output.

18 . A camera having a field of view including at least a portion of at least one surface for accommodating one or more items, wherein the camera comprises a processor unit and an optical sensor,

wherein the processor unit is configured to execute a machine learning system trained to determine whether at least one actor is associated with at least one of a plurality of events based at least in part on a feature of the at least one actor detected in at least one image,

wherein the processor unit is programmed with one or more sets of instructions that, when executed by the processor unit, cause the camera to execute a method comprising:

capturing a plurality of images over a period of time;

detecting locations of at least a first body part of a first actor depicted within at least a subset of the plurality of images;

detecting locations of at least a second body part of a second actor depicted within the subset of the plurality of images;

determining first values of data corresponding to at least a first feature of the first actor based at least in part on at least one of the locations of at least the first body part of the first actor depicted within at least the subset of the plurality of images;

determining second values of data corresponding to at least a second feature of the second actor based at least in part on at least one of the locations of at least the second body part of the second actor depicted within at least the subset of the plurality of images;

constructing, by the first imaging device, a first time series of data, wherein the first time series of data comprises the first values of data corresponding to the first feature and statuses of the first body part of the first actor at times at which each of the subset of the plurality of images was captured;

constructing, by the first imaging device, a second time series of data, wherein the second time series of data comprises the second values of data corresponding to the second feature and statuses of the second body part of the second actor at the times at which each of the subset of the first plurality of images was captured;

providing at least a portion of the first time series of data as a first input to the machine learning system;

receiving at least a first output from the machine learning system in response to the first input;

providing at least a portion of the second time series of data as a second input to the machine learning system;

receiving at least a second output from the machine learning system in response to the second input; and

determining, based at least in part on the first output and the second output, that the first actor is associated with one of the plurality of events,

wherein the one of the plurality of events comprises one of:

placing at least one item on the at least one surface; or

removing at least one item from the at least one surface.

19 . The camera of claim 18 , wherein the first feature is one of:

a trajectory of a hand of the first actor over the period of time, wherein the first body part is the hand of the first actor;

a trajectory of a head of the first actor over the period of time, wherein the first body part is the head of the first actor; or

a trajectory of a shoulder of the first actor over the period of time, wherein the first body part is the shoulder of the first actor, and

wherein the second feature is one of:

a trajectory of a hand of the second actor over the period of time, wherein the second body part is the hand of the second actor;

a trajectory of a head of the second actor over the period of time, wherein the second body part is the head of the second actor; or

a trajectory of a shoulder of the second actor over the period of time, wherein the second body part is the shoulder of the second actor.

20 . The camera of claim 18 , wherein each of the first values of data is a location of the first body part depicted within one of the subset of the plurality of images,

wherein each one of the statuses of the first body part is one of empty or full,

wherein each one of the second values of data is a location of the second body part depicted within one of the plurality of images, and

wherein each one of the statuses of the second body part is one of empty or full.

References Cited (183)
US 6154559A · Beardsley · 2000 [cited by applicant]
US 7050624B2 · Dialameh et al. · 2006 [cited by applicant]
US 7225980B2 · Ku et al. · 2007 [cited by applicant]
US 7949568B2 · Fano et al. · 2011 [cited by applicant]
US 8009863B1 · Sharma et al. · 2011 [cited by applicant]
US 8009864B2 · Linaker et al. · 2011 [cited by applicant]
US 8175925B1 · Rouaix · 2012 [cited by applicant]
US 8189855B2 · Opalach et al. · 2012 [cited by applicant]
US 8285060B2 · Cobb et al. · 2012 [cited by applicant]
US 8369622B1 · Hsu et al. · 2013 [cited by applicant]
US 8423431B1 · Rouaix et al. · 2013 [cited by applicant]
US RE44225E · Aviv · 2013 [cited by applicant]
US 8577705B1 · Baboo et al. · 2013 [cited by applicant]
US 8630924B2 · Groenevelt et al. · 2014 [cited by applicant]
US 8688598B1 · Shakes et al. · 2014 [cited by applicant]
US 8943441B1 · Patrick et al. · 2015 [cited by applicant]
US 9158974B1 · Laska et al. · 2015 [cited by applicant]
US 9160979B1 · Ulmer · 2015 [cited by applicant]
US 9208675B2 · Xu et al. · 2015 [cited by applicant]
US 9336456B2 · Delean · 2016 [cited by applicant]
US 9449233B2 · Taylor · 2016 [cited by applicant]
US 9473747B2 · Kobres et al. · 2016 [cited by applicant]
US 9536177B2 · Chalasani et al. · 2017 [cited by applicant]
US 9582891B2 · Geiger et al. · 2017 [cited by applicant]
US 9727838B2 · Campbell · 2017 [cited by applicant]
US 9846840B1 · Lin et al. · 2017 [cited by applicant]
US 9881221B2 · Bala et al. · 2018 [cited by applicant]
US 9898677B1 · Andjelković et al. · 2018 [cited by applicant]
US 9911290B1 · Zalewski · 2018 [cited by examiner]
US 10055853B1 · Fisher et al. · 2018 [cited by applicant]
US 10133933B1 · Fisher et al. · 2018 [cited by applicant]
US 10147210B1 · Desai · 2018 [cited by examiner]
US 10192415B2 · Heitz et al. · 2019 [cited by applicant]
US 10318917B1 · Goldstein et al. · 2019 [cited by applicant]
US 10354262B1 · Hershey · 2019 [cited by examiner]
US 10438277B1 · Jiang et al. · 2019 [cited by applicant]
US 10474992B2 · Fisher et al. · 2019 [cited by applicant]
US 10510219B1 · Zalewski · 2019 [cited by examiner]
US 10535146B1 · Buibas et al. · 2020 [cited by applicant]
US 10635844B1 · Roose et al. · 2020 [cited by applicant]
US 10699421B1 · Cherevatsky et al. · 2020 [cited by applicant]
US 10839203B1 · Guigues et al. · 2020 [cited by applicant]
US 11195146B2 · Fisher et al. · 2021 [cited by applicant]
US 11232294B1 · Banerjee et al. · 2022 [cited by applicant]
US 11270260B2 · Fisher et al. · 2022 [cited by applicant]
US 11284041B1 · Bergamo et al. · 2022 [cited by applicant]
US 11367083B1 · Saurabh et al. · 2022 [cited by applicant]
US 11468698B1 · Kim et al. · 2022 [cited by applicant]
US 11482045B1 · Kim et al. · 2022 [cited by applicant]
US 11538186B2 · Fisher et al. · 2022 [cited by applicant]
US 20030002712A1 · Steenburgh et al. · 2003 [cited by applicant]
US 20030002717A1 · Hamid · 2003 [cited by applicant]
US 20030107649A1 · Flickner et al. · 2003 [cited by applicant]
US 20030128337A1 · Jaynes et al. · 2003 [cited by applicant]
US 20040181467A1 · Raiyani et al. · 2004 [cited by applicant]
US 20050251347A1 · Perona et al. · 2005 [cited by applicant]
US 20060018516A1 · Masoud et al. · 2006 [cited by applicant]
US 20060061583A1 · Spooner et al. · 2006 [cited by applicant]
US 20060222206A1 · Garoutte · 2006 [cited by applicant]
US 20070092133A1 · Luo · 2007 [cited by applicant]
US 20070156625A1 · Visel · 2007 [cited by applicant]
US 20070182818A1 · Buehler · 2007 [cited by applicant]
US 20070242066A1 · Rosenthal · 2007 [cited by applicant]
US 20070276776A1 · Sagher et al. · 2007 [cited by applicant]
US 20080055087A1 · Horii et al. · 2008 [cited by applicant]
US 20080077511A1 · Zimmerman · 2008 [cited by applicant]
US 20080109114A1 · Orita et al. · 2008 [cited by applicant]
US 20080137989A1 · Ng et al. · 2008 [cited by applicant]
US 20080159634A1 · Sharma et al. · 2008 [cited by applicant]
US 20080166019A1 · Lee · 2008 [cited by applicant]
US 20080193010A1 · Eaton et al. · 2008 [cited by applicant]
US 20080195315A1 · Hu et al. · 2008 [cited by applicant]
US 20090060352A1 · Distante et al. · 2009 [cited by applicant]
US 20090083815A1 · McMaster et al. · 2009 [cited by applicant]
US 20090121017A1 · Cato et al. · 2009 [cited by applicant]
US 20090132371A1 · Strietzel et al. · 2009 [cited by applicant]
US 20090210367A1 · Armstrong et al. · 2009 [cited by applicant]
US 20090245573A1 · Saptharishi et al. · 2009 [cited by applicant]
US 20090276705A1 · Ozdemir et al. · 2009 [cited by applicant]
US 20100002082A1 · Buehler et al. · 2010 [cited by applicant]
US 20100033574A1 · Ran et al. · 2010 [cited by applicant]
US 20110011936A1 · Morandi et al. · 2011 [cited by applicant]
US 20110205022A1 · Cavallaro et al. · 2011 [cited by applicant]
US 20120148103A1 · Hampel et al. · 2012 [cited by applicant]
US 20120159290A1 · Pulsipher et al. · 2012 [cited by applicant]
US 20120257789A1 · Lee et al. · 2012 [cited by applicant]
US 20120284132A1 · Kim et al. · 2012 [cited by applicant]
US 20120327220A1 · Ma · 2012 [cited by applicant]
US 20130076898A1 · Philippe et al. · 2013 [cited by applicant]
US 20130095961A1 · Marty et al. · 2013 [cited by applicant]
US 20130156260A1 · Craig · 2013 [cited by applicant]
US 20130253700A1 · Carson et al. · 2013 [cited by applicant]
US 20130322767A1 · Chao et al. · 2013 [cited by applicant]
US 20140139633A1 · Wang et al. · 2014 [cited by applicant]
US 20140139655A1 · Mimar · 2014 [cited by applicant]
US 20140259056A1 · Grusd · 2014 [cited by applicant]
US 20140279294A1 · Field-Darragh et al. · 2014 [cited by applicant]
US 20140282162A1 · Fein et al. · 2014 [cited by applicant]
US 20140334675A1 · Chu et al. · 2014 [cited by applicant]
US 20140362195A1 · Ng-Thow-Hing et al. · 2014 [cited by applicant]
US 20140362223A1 · LaCroix et al. · 2014 [cited by applicant]
US 20140379296A1 · Nathan et al. · 2014 [cited by applicant]
US 20150019391A1 · Kumar et al. · 2015 [cited by applicant]
US 20150039458A1 · Reid · 2015 [cited by applicant]
US 20150073907A1 · Purves et al. · 2015 [cited by applicant]
US 20150131851A1 · Bernal et al. · 2015 [cited by applicant]
US 20150199824A1 · Kim et al. · 2015 [cited by applicant]
US 20150206188A1 · Tanigawa et al. · 2015 [cited by applicant]
US 20150262116A1 · Katircioglu et al. · 2015 [cited by applicant]
US 20150269143A1 · Park et al. · 2015 [cited by applicant]
US 20150294483A1 · Wells et al. · 2015 [cited by applicant]
US 20160003636A1 · Ng-Thow-Hing et al. · 2016 [cited by applicant]
US 20160125245A1 · Saitwal et al. · 2016 [cited by applicant]
US 20160127641A1 · Gove · 2016 [cited by applicant]
US 20160292881A1 · Bose et al. · 2016 [cited by applicant]
US 20160307335A1 · Perry et al. · 2016 [cited by applicant]
US 20170116473A1 · Sashida et al. · 2017 [cited by applicant]
US 20170206669A1 · Saleemi et al. · 2017 [cited by applicant]
US 20170262994A1 · Kudriashov et al. · 2017 [cited by applicant]
US 20170278255A1 · Shingu et al. · 2017 [cited by applicant]
US 20170309136A1 · Schoner · 2017 [cited by applicant]
US 20170323376A1 · Glaser et al. · 2017 [cited by applicant]
US 20170345165A1 · Stanhill et al. · 2017 [cited by applicant]
US 20170353661A1 · Kawamura · 2017 [cited by applicant]
US 20180025175A1 · Kato · 2018 [cited by applicant]
US 20180070056A1 · DeAngelis et al. · 2018 [cited by applicant]
US 20180084242A1 · Rublee et al. · 2018 [cited by applicant]
US 20180164103A1 · Hill · 2018 [cited by applicant]
US 20180165728A1 · McDonald et al. · 2018 [cited by applicant]
US 20180218515A1 · Terekhov et al. · 2018 [cited by applicant]
US 20180315329A1 · D'Amato et al. · 2018 [cited by applicant]
US 20180343442A1 · Yoshikawa et al. · 2018 [cited by applicant]
US 20190043003A1 · Fisher · 2019 [cited by examiner]
US 20190073627A1 · Nakdimon · 2019 [cited by examiner]
US 20190102044A1 · Wang et al. · 2019 [cited by applicant]
US 20190156274A1 · Fisher et al. · 2019 [cited by applicant]
US 20190156277A1 · Fisher et al. · 2019 [cited by applicant]
US 20190158801A1 · Matsubayashi · 2019 [cited by applicant]
US 20190236531A1 · Adato et al. · 2019 [cited by applicant]
US 20190315329A1 · Adamski et al. · 2019 [cited by applicant]
US 20200005490A1 · Paik et al. · 2020 [cited by applicant]
US 20200043086A1 · Sorensen · 2020 [cited by applicant]
US 20200090484A1 · Chen et al. · 2020 [cited by applicant]
US 20200279382A1 · Zhang et al. · 2020 [cited by applicant]
US 20200320287A1 · Porikli et al. · 2020 [cited by applicant]
US 20200380274A1 · Shin et al. · 2020 [cited by applicant]
US 20210027485A1 · Zhang · 2021 [cited by examiner]
US 20210124936A1 · Mirza et al. · 2021 [cited by applicant]
US 20210125341A1 · Mirza et al. · 2021 [cited by applicant]
US 20210182922A1 · Zheng et al. · 2021 [cited by applicant]
CN 104778690B · 2017 [cited by applicant]
EP 1574986B1 · 2008 [cited by applicant]
JP 2013196199A · 2013 [cited by applicant]
JP 201489626A · 2014 [cited by applicant]
JP 2018207336A · 2018 [cited by applicant]
JP 2019018743A · 2019 [cited by applicant]
JP 2019096996A · 2019 [cited by applicant]
KR 20170006097A · 2017 [cited by applicant]
WO 0021021A1 · 2000 [cited by applicant]
WO 02059836A2 · 2002 [cited by applicant]
WO 2017151241A2 · 2017 [cited by applicant]
Black, J. et al., “Multi View Image Surveillance and Tracking,” IEEE Proceedings of the Workshop on Motion and Video Computing, 2002, https://www.researchgate.net/publication/4004539_Multi_view_image_surveillance_and_tr… [cited by applicant]
Harville, M., “Stereo Person Tracking with Adaptive Plan-View Templates of Height and Occupancy Statistics,” Image and Vision Computing, vol. 22, Issue 2, Feb. 1, 2004, https://www.researchgate.net/publication/223214495… [cited by applicant]
Huang, K. S. et al. “Driver's View and Vehicle Surround Estimation Using Omnidirectional Video Stream,” IEEE V2003 Intelligent Vehicles Symposium. Proceedings (Cal. No. 03TH8683), Jun. 9-11, 2003, http://cvrr.ucsd.edu/V… [cited by applicant]
Longuet-Higgins, H.C., “A Computer Algorithm for Reconstructing a Scene from Two Projections,” Nature 293, Sep. 10, 1981, https://cseweb.ucsd.edu/classes/fa01/cse291/hclh/SceneReconstruction.pdf, pp. 133-135. [cited by applicant]
Rossi, M. and Bozzoli, E. A., “Tracking and Counting Moving People,” IEEE Int'l Conf. on Image Processing, ICIP-94, Nov. 13-16, 1994, http://citeseerx.ist.psu.edu/viewdoc/download;isessionid=463D09F419FA5595DBF9DEF30D7E… [cited by applicant]
Vincze, M., “Robust Tracking of Ellipses at Frame Rate,” Pattern Recognition, vol. 34, Issue 2, Feb. 2001, pp. 487-498. [cited by applicant]
Zhang, Z., “A Flexible New Technique for Camera Calibration,” IEEE Transactions on Pattern Analysis and Machine Intelligence, vol. 22, No. 11, Nov. 2000, 5 pages. [cited by applicant]
Zhang, Z., “A Flexible New Technique for Camera Calibration,” Technical Report MSR-TR-98-71, Microsoft Research, Microsoft Corporation, microsoft.com/en-us/research/wp-content/uploads/2016/02/tr98-71.pdf, 22 pages. [cited by applicant]
Lee, K. and Kacorri, H., (2019, May), “Hands Holding Clues for Object Recognition in Teachable Machines”, In Proceedings of the 2019 CHI Conference on Human Factors in Computing Systems (pp. 1-12). [cited by applicant]
Liu, C., et al. “Accelerating Vanishing Point-Based Line Sampling Scheme for Real-Time People Localization”, IEEE Transactions on Circuits and Systems for Video Technology. vol 27. No. Mar. 3, 2017 (Year: 2017). [cited by applicant]
Abhaya Asthana et al., “An Indoor Wireless System for Personalized Shopping Assistance”, Proceedings of IEEE Workshop on Mobile Computing Systems and Applications, 1994, pp. 69-74, Publisher: IEEE Computer Society Press. [cited by applicant]
Ciplak G, Telceken S., “Moving Object Tracking Within Surveillance Video Sequences Based on EDContours,” 2015 9th International Conference on Electrical and Electronics Engineering (ELECO), Nov. 2, 20156 (pp. 720-723). … [cited by applicant]
Cristian Pop, “Introduction to the BodyCom Technology”, Microchip AN1391, May 2, 2011, pp. 1-24, vol. AN1391, No. DS01391A, Publisher: 2011 Microchip Technology Inc. [cited by applicant]
Fuentes et al., “People tracking in surveillance applications,” Proceedings 2nd IEEE Int. Workshop on PETS, Kauai, Hawaii, USA, Dec. 9, 2001, 6 pages. [cited by applicant]
Manocha et al., “Object Tracking Techniques for Video Tracking: A Survey,” The International Journal of Engineering and Science (IJES), vol. 3, Issue 6, pp. 25-29, 2014. [cited by applicant]
Phalke K, Hegadi R., “Pixel Based Object Tracking,” 2015 2nd International Conference on Signal Processing and Integrated Networks (SPIN), Feb. 1, 20159 (pp. 575-578). IEEE. [cited by applicant]
Sikdar A, Zheng YF, Xuan D., “Robust Object Tracking in the X-Z Domain,” 2016 IEEE International Conference on Multisensor Fusion and Integration for Intelligent Systems (MFI), Sep. 19, 2016 (pp. 499-504). IEEE. [cited by applicant]
Grinciunaite, A., et al., “Human Pose Estimation in Space and Time Using 3D CNN,” ECCV Workshop on Brave New Ideas for Motion Representations in Videos, Oct. 19, 2016, URL: https://arxiv.org/pdf/1609.00036.pdf, 7 pages. [cited by applicant]
He, K., et al., “Identity Mappings in Deep Residual Networks,” ECCV 2016 Camera-Ready, URL: https://arxiv.org/pdf/1603.05027.pdf, Jul. 25, 2016, 15 pages. [cited by applicant]
Redmon, J., et al., “You Only Look Once: Unified, Real-Time Object Detection,” University of Washington, Allen Institute for AI, Facebook Al Research, URL: https://arxiv.org/pdf/1506.02640.pdf, May 9, 2016, 10 pages. [cited by applicant]
Redmon, Joseph and Ali Farhadi, “YOLO9000: Better, Faster, Stronger,” URL: https:/arxiv.org/pdf/1612.08242.pdf, Dec. 25, 2016, 9 pages. [cited by applicant]
Toshev, Alexander and Christian Szegedy, “DeepPose: Human Pose Estimation via Deep Neural Networks,” IEEE Conference on Computer Vision and Pattern Recognition, Aug. 20, 2014, URL: https://arxiv.org/pdf/1312.4659.pdf, 9… [cited by applicant]