IP Library › Granted Patent US 11,106,901
Granted Patent B2
US 11,106,901 · App. 16/713,479 · Granted Aug 31, 2021

Method and system for recognizing user actions with respect to objects

Inventors: Hualin He (Hangzhou, CN); Chunxiang Pan (Hangzhou, CN); Yabo Ni (Hangzhou, CN); Anxiang Zeng (Hangzhou, CN)
Assignee: ALIBABA GROUP HOLDING LIMITED
G06K9/00335G06F3/017G06F3/0304G06K9/00375G06T2207/10016
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,106,901
App. No.
16/713,479
Granted
Aug 31, 2021
Kind
B2
Abstract

The specification discloses a computer-implemented method for user action determination, comprising: recognizing an item displacement action performed by a user; determining a first time and a first location of the item displacement action; recognizing a target item in a non-stationary state; determining a second time when the target item is in the non-stationary state and a second location where the target item is in the non-stationary state; and in response to determining that the first time matches the second time and the first location matches the second location, determining that the item displacement action of the user is performed with respect to the target item.

Claims (67)

1. A computer-implemented method for user action determination, comprising:

recognizing an item displacement action performed by a user;

determining a first time and a first location of the item displacement action;

recognizing a target item in a non-stationary state;

determining a second time when the target item is in the non-stationary state and a second location where the target item is in the non-stationary state; and

in response to determining that the first time matches the second time and the first location matches the second location, determining that the item displacement action of the user is performed with respect to the target item.

2. The method according to claim 1 , wherein the item displacement action performed by the user comprises flipping one or more items, moving one or more items, or picking up one or more items.

3. The method according to claim 1 , wherein recognizing a target item in a non-stationary state comprises:

receiving radio frequency (RF) signal sensed by an RF sensor from a radio frequency identification (RFID) tag attached to the target item; and

in response to detecting fluctuation in the RF signal, determining the target item in the non-stationary state.

4. The method according to claim 3 , wherein the RF signal comprises a tag identifier, a signal strength, or a phase value.

5. The method according to claim 3 , wherein determining a second time when the target item is in a non-stationary state and a second location where the target item is in a non-stationary state comprises:

determining the second time and the second location according to the RF signal sensed by the RF sensor.

6. The method according to claim 1 , wherein recognizing the item displacement action performed by the user comprises:

obtaining a video recording the item displacement action of the user;

cropping the video to obtain a video of hands including images of the user's hands;

inputting the video of hands into a classification model; and

obtaining from the classification model a determination regarding whether the user has performed the item displacement action.

7. The method according to claim 6 , wherein determining a first time and a first location of the item displacement action comprises:

determining the first time and the first location of the item displacement action according to the video recording the item displacement action of the user.

8. The method according to claim 6 , wherein the classification model is training with samples comprising a video of hands of users performing item displacement actions and a video of hands of users not performing item displacement actions.

9. The method according to claim 8 , wherein the classification model includes a 3D convolution model.

10. The method according to claim 6 , wherein the video of hands comprises a video of a left hand and a video of a right hand, and cropping the video to obtain a video of hands including images of the user's hands comprises:

determining coordinates of left and right wrists of the user respectively in a plurality of image frames in the video;

cropping each of the plurality of image frames to obtain a left hand image and a right hand image according to the coordinates of the left and right wrists of the user, respectively;

combining a plurality of left hand images obtained from the plurality of image frames to obtain a left hand video; and

combining a plurality of right hand images obtained from the plurality of image frames to obtain a right hand video.

11. The method according to claim 10 , wherein determining coordinates of left and right wrists of the user respectively in a plurality of image frames in the video comprises:

using a human posture detection algorithm to determine the coordinates of the left and right wrists of the user in one of the plurality of image frames in the video;

determine a first enclosing rectangle of the user;

using a target tracking algorithm to determine a second enclosing rectangle of the user in a next one of the plurality of image frames in the video based on the first enclosing rectangle of the user; and

using the human posture detection algorithm to determine the coordinates of the left and right wrists of the user in the next one of the plurality of image frames in the video based on the left and right wrists in the second enclosing rectangle of the user.

12. The method according to claim 10 , further comprising:

in response to a determination from the classification model that the user has performed the item displacement action, determining an average of the coordinates of the left or right wrist of the user in the plurality of image frames in the video, and determining whether the average of the coordinates is within a target area; and

in response to determining that the average of the coordinates is not within the target area, determining that the user had not performed the item displacement action.

13. The method according to claim 12 , wherein the target area includes a shelf area in a shop, and the target item includes a shelved merchandise.

14. An apparatus for user action determination, comprising one or more processors and one or more non-transitory computer-readable memories coupled to the one or more processors and configured with instructions executable by the one or more processors to cause the apparatus to perform operations comprising:

recognizing an item displacement action performed by a user;

determining a first time and a first location of the item displacement action;

recognizing a target item in a non-stationary state;

determining a second time when the target item is in the non-stationary state and a second location where the target item is in the non-stationary state; and

in response to determining that the first time matches the second time and the first location matches the second location, determining that the item displacement action of the user is performed with respect to the target item.

15. The apparatus according to claim 14 , wherein the item displacement action performed by the user comprises flipping one or more items, moving one or more items, or picking up one or more items.

16. The apparatus according to claim 14 , wherein recognizing the item displacement action performed by the user comprises:

obtaining a video recording the item displacement action of the user;

cropping the video to obtain a video of hands including images of the user's hands;

inputting the video of hands into a classification model; and

obtaining from the classification model a determination regarding whether the user has performed the item displacement action.

17. The apparatus according to claim 16 , wherein the video of hands comprises a video of a left hand and a video of a right hand, and cropping the video to obtain a video of hands including images of the user's hands comprises:

determining coordinates of left and right wrists of the user respectively in a plurality of image frames in the video;

cropping each of the plurality of image frames to obtain a left hand image and a right hand image according to the coordinates of the left and right wrists of the user, respectively;

combining a plurality of left hand images obtained from the plurality of image frames to obtain a left hand video; and

combining a plurality of right hand images obtained from the plurality of image frames to obtain a right hand video.

18. The apparatus according to claim 17 , wherein determining coordinates of left and right wrists of the user respectively in a plurality of image frames in the video comprises:

using a human posture detection algorithm to determine the coordinates of the left and right wrists of the user in one of the plurality of image frames in the video;

determine a first enclosing rectangle of the user;

using a target tracking algorithm to determine a second enclosing rectangle of the user in a next one of the plurality of image frames in the video based on the first enclosing rectangle of the user; and

using the human posture detection algorithm to determine the coordinates of the left and right wrists of the user in the next one of the plurality of image frames in the video based on the left and right wrists in the second enclosing rectangle of the user.

19. The apparatus according to claim 17 , wherein the operations further comprise:

in response to a determination from the classification model that the user has performed the item displacement action, determining an average of the coordinates of the left or right wrist of the user in the plurality of image frames in the video, and determining whether the average of the coordinates is within a target area; and

in response to determining that the average of the coordinates is not within the target area, determining that the user had not performed the item displacement action.

20. A non-transitory computer-readable storage medium configured with instructions executable by one or more processors to cause the one or more processors to perform operations comprising:

recognizing an item displacement action performed by a user;

determining a first time and a first location of the item displacement action;

recognizing a target item in a non-stationary state;

determining a second time when the target item is in the non-stationary state and a second location where the target item is in the non-stationary state; and

in response to determining that the first time matches the second time and the first location matches the second location, determining that the item displacement action of the user is performed with respect to the target item.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Mar 19, 2020
From: HE, HUALI; PAN, CHUNXIANG; NI, YABO; ZENG, ANXIANG
To: ALIBABA GROUP HOLDING LIMITED
Reel/Frame 052170/0306 →
Priority Claims (1)
CN 201811533006.X · Dec 14, 2018 · national
Continuity (1)
Related Publication 20200193148A1 · Jun 18, 2020
Cited By (1)
US 12,646,291