IP Library Granted Patent US 11,636,744
Granted Patent B2
US 11,636,744 · App. 17/169,191 · Granted Apr 25, 2023

Retail inventory shrinkage reduction via action recognition

Inventors: Weilin Huang (Shenzhen, CN); Shiwen Zhang (Shenzhen, CN); Limin Wang (Shenzhen, CN); Sheng Guo (Shenzhen, CN); Matthew Robert Scott (Shenzhen, CN)
Assignee: Shenzhen Malong Technologies Co., Ltd.
G08B13/19613G06V20/41G06V20/46G06V20/52G06V40/20G06Q10/087G06Q50/26G06V20/44
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,636,744
App. No.
17/169,191
Granted
Apr 25, 2023
Kind
B2
Abstract

This disclosure includes technologies for action recognition in general. The disclosed system may automatically detect various types of actions in a video, including reportable actions that cause shrinkage in a practical application for loss prevention in the retail industry. Further, appropriate responses may be invoked if a reportable action is recognized. In some embodiments, a three-branch architecture may be used in a machine learning model for action and/or activity recognition. The three-branch architecture may include a main branch for action recognition, an auxiliary branch for learning/identifying an actor (e.g., human parsing) related to an action, and an auxiliary branch for learning/identifying a scene related to an action. In this three-branch architecture, the knowledge of the actor and the scene may be integrated in two different levels for action and/or activity recognition.

Claims (63)

1. A computer-implemented method for action recognition, comprising:

processing video data by one or more auxiliary branches of a network;

identifying, by the one or more auxiliary branches, intermediate auxiliary features corresponding to the video data;

integrating intermediate action features from an action branch and the intermediate auxiliary features from the one or more auxiliary branches of the network;

generating, based on the integrated intermediate action features and intermediate auxiliary features, a set of high-level action features by the action branch;

integrating high-level auxiliary features from the one or more auxiliary branches and the high-level action features from the action branch; and

classifying an action based on the integrated high-level auxiliary features and the high-level action features.

2. The computer-implemented method of claim 1 , wherein the one or more auxiliary branches are selected from the group consisting of a human parsing branch and a scene recognition branch.

3. The computer-implemented method of claim 1 , further comprising:

parsing a subject identified in the video data;

providing, by a teacher network, a pseudo ground truth of the subject; and

comparing an output from the one or more auxiliary branches to the pseudo ground truth.

4. The computer implemented method of claim 1 , further comprising:

determining the action is a reportable action; and

based on the determined reportable action, providing an indicator of the action to a user.

5. The computer implemented method of claim 4 , wherein the indicator includes a portion of video data corresponding to the determined reportable action.

6. The computer implemented method of claim 1 , further comprising tracking an object of interest in a scene of a retail environment;

determining a subject identified in the video data has performed a reportable action associated with a shrinkage event with the object of interest; and

providing an indicator to a user of the reportable action.

7. The computer-implemented method of claim 1 , wherein the intermediate action features correspond to determined movements in the video data and are associated with one or more shrinkage events.

8. A system comprising:

one or more processors;

one or more memory devices storing instructions thereon, that when executed by the one or more processors, cause the one or more processors to execute operations comprising:

processing video data, by one or more auxiliary branches of a network, wherein the video data includes images of a scene in a retail environment;

identifying, by the one or more auxiliary branches, intermediate auxiliary features corresponding to the video data, wherein the intermediate auxiliary features include subject features corresponding to a subject in the scene of the retail environment and scene features corresponding to the scene of the retail environment;

integrating intermediate action features corresponding to subject movements from an action branch and the intermediate auxiliary features from the one or more auxiliary branches of the network;

generating, based on the integrated intermediate action features and intermediate auxiliary features, a set of high-level action features corresponding to the subject movements by the action branch;

integrating high-level auxiliary features from the one or more auxiliary branches and the high-level action features from the action branch;

classifying a reportable action based on the integrated high-level auxiliary features and the high-level action features, wherein the reportable action is associated with an inventory shrinkage event;

based on the classified reportable action, determining the subject identified in the video data has performed the reportable action; and

providing an indicator to a user, wherein the indicator includes information relating to the inventory shrinkage event.

9. The system of claim 8 , wherein the one or more auxiliary branches are selected from the group consisting of a human parsing branch and a scene recognition branch.

10. The system of claim 8 , further comprising:

parsing the subject identified in the video data from one or more segments of a data stream;

providing, by a teacher network, a pseudo ground truth of the subject; and

comparing an output from the one or more auxiliary branches to the pseudo ground truth.

11. The system of claim 8 , further comprising:

based on the determined reportable action, generating a report of the reportable action, wherein the report includes data corresponding to the reportable action.

12. The system of claim 8 , wherein the indicator is provided to the subject corresponding to the determined reportable action.

13. The system of claim 8 , further comprising:

tracking an object of interest in the scene of the retail environment;

determining the subject identified in the video data has performed the reportable action with the object of interest; and

providing the indicator to the user of the reportable action, wherein the indicator includes image data corresponding to the object of interest.

14. The system of claim 8 , wherein the shrinkage event is a theft.

15. One or more non-transitory computer storage media having computer-executable instructions embodied thereon that, when executed, by one or more processors, cause the one or more processors to perform a method for implementing action recognition systems, the method comprising:

processing video data, by one or more auxiliary branches of a network;

identifying, by the one or more auxiliary branches, intermediate auxiliary features corresponding to the video data;

integrating intermediate action features from an action branch and the intermediate auxiliary features from the one or more auxiliary branches of the network;

generating, based on the integrated intermediate action features and intermediate auxiliary features, a set of high-level action features by the action branch;

integrating high-level auxiliary features from the one or more auxiliary branches and the high-level action features from the action branch; and

classifying an action based on the integrated high-level auxiliary features and the high-level action features.

16. The non-transitory media of claim 15 , wherein the one or more auxiliary branches are selected from the group consisting of a human parsing branch and a scene recognition branch.

17. The non-transitory media of claim 15 , further comprising:

parsing a subject identified in the video data;

providing, by a teacher network, a pseudo ground truth of the subject; and

comparing an output from the one or more auxiliary branches to the pseudo ground truth.

18. The non-transitory media of claim 15 , further comprising:

determining the action is a reportable action; and

based on the determined reportable action, providing an indicator of the action to a user.

19. The non-transitory media of claim 18 , wherein the indicator includes a portion of video data corresponding to the determined reportable action.

20. The non-transitory media of claim 1 , further comprising tracking an object of interest in a scene;

determining a subject identified in the video data has performed a reportable action with the object of interest; and

providing an indicator to a user of the reportable action.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Mar 4, 2021
From: HUANG, WEILIN; ZHANG, SHIWEN; WANG, LIMIN; GUO, SHENG; SCOTT, MATTHEW ROBERT
To: SHENZHEN MALONG TECHNOLOGIES CO., LTD.
Reel/Frame 055494/0681 →
Continuity (2)
Provisional Application 62971189 · Feb 6, 2020
Related Publication 20210248885A1 · Aug 12, 2021