IP Library › Granted Patent US 11,710,351
Granted Patent B2
US 11,710,351 · App. 17/321,237 · Granted Jul 25, 2023

Action recognition method and apparatus, and human-machine interaction method and apparatus

Inventors: Jingmin Luo (Shenzhen, CN); Liang Qiao (Shenzhen, CN); Xiaolong Zhu (Shenzhen, CN)
Assignee: TENCENT TECHNOLOGY (SHENZHEN) COMPANY LIMITED
G06V40/23G06T7/74G06V10/754G06V10/82G06V20/41G06V20/46G06V20/48G06V40/103G06V10/422
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,710,351
App. No.
17/321,237
Granted
Jul 25, 2023
Kind
B2
Abstract

A computer device extracts a plurality of target windows from a target video. Each of the target windows comprises a respective plurality of consecutive video frames. For each of the target windows, the device performs action recognition on the respective plurality of consecutive video frames corresponding to the target window to obtain respective first action feature information of the target window. The device obtains a similarity between the first action feature information of the target window and preset feature information. The device determines, from the respective obtained similarities corresponding to the plurality of target windows, a highest first similarity and a first target window corresponding to the highest first similarity. The device also determines a dynamic action corresponding to the highest first similarity as the preset dynamic action in accordance with threshold settings.

Claims (65)

1. An action recognition method, performed by an electronic device, the method comprising:

extracting a plurality of target windows from a target video, each of the plurality of target windows comprising a respective plurality of consecutive video frames;

for each of the plurality of target windows:

performing action recognition on the respective plurality of consecutive video frames corresponding to the target window to obtain respective first action feature information of the target window, the first action feature information comprising movement of one or more body parts of a subject and is used for describing a dynamic action comprised in the target window; and

obtaining a similarity between the first action feature information of the target window and preset feature information, the preset feature information being used for describing a preset dynamic action;

determining, from the respective obtained similarities corresponding to the plurality of target windows, a highest first similarity and a first target window corresponding to the highest first similarity; and

determining a dynamic action corresponding to the highest first similarity as the preset dynamic action when (i) the highest first similarity is greater than a first preset threshold and (ii) a difference between the highest first similarity and a second similarity is greater than a second preset threshold, the second similarity being a similarity between respective first action feature information of a target window adjacent to the first target window and the preset feature information.

2. The method according to claim 1 , wherein extracting the plurality of target windows from a target video further comprises:

classifying the plurality of video frames.

3. The method according to claim 1 , wherein performing the action recognition on the respective plurality of consecutive video frames further comprises:

extracting, for each of the video frames, a plurality of body key points in the video frame;

performing action recognition on the video frame according to a distribution of the plurality of body key points, to obtain second action feature information of the video frame, the second action feature information is used for describing a static action comprised in the video frame; and

combining a plurality of second action feature information of the video frames in the target window, to obtain the first action feature information of the target window.

4. The method according to claim 3 , wherein the preset dynamic action is performed using at least two preset body parts; and performing action recognition on the video frame according to the distribution comprises one or more of:

determining an angle between any two preset body parts in the video frame according to coordinates of the plurality of body key points in the video frame and body parts to which the plurality of body key points belong, and using the angle as the second action feature information;

obtaining a displacement amount between at least one of the plurality of body key points and a body key point corresponding to a reference video frame, and using the displacement amount as the second action feature information, the reference video frame being previous to the each video frame by a third preset quantity of video frames; and

obtaining a size of a reference body part in any two preset body parts and a distance between the any two preset body parts, and using a ratio of the distance to the size of the reference body part as the second action feature information.

5. The method according to claim 1 , wherein the first action feature information is a first action matrix comprising M first action vectors, the preset feature information is a preset action matrix comprising N preset action vectors, M and N being positive integers, and the obtaining a similarity between the first action feature information of the each target window and preset feature information comprises:

creating a similarity matrix, the similarity matrix having M rows and N columns, or N rows and M columns;

obtaining, for a specified position corresponding to an i th first action vector and a j th preset action vector in the similarity matrix, a sum of a maximum similarity among similarities of a first position, a second position, and a third position and a similarity between the i th first action vector and the j th preset action vector as a similarity of the specified position, the first position being a position corresponding to an (i−1) th first action vector and the j th preset action vector, the second position being a position corresponding to the (i−1) th first action vector and a (j−1) th preset action vector, and the third position being a position corresponding to the i th first action vector and the (j−1) th preset action vector, i being a positive integer not less than 1 and not greater than M, and j being a positive integer not less than 1 and not greater than N; and

determining a similarity of a position corresponding to an M th first action vector and an N th preset action vector in the similarity matrix as the similarity between the first action feature information and the preset feature information.

6. An electronic device, comprising:

one or more processors; and

memory storing one or more programs that, when executed by the one or more processors, cause the one or more processors to perform operations comprising:

extracting a plurality of target windows from a target video, each of the plurality of target windows comprising a respective plurality of consecutive video frames;

for each of the plurality of target windows:

performing action recognition on the respective plurality of consecutive video frames corresponding to the target window to obtain respective first action feature information of the target window, the first action feature information comprising movement of one or more body parts of a subject and is used for describing a dynamic action comprised in the target window; and

obtaining a similarity between the first action feature information of the target window and preset feature information, the preset feature information being used for describing a preset dynamic action;

determining, from the respective obtained similarities corresponding to the plurality of target windows, a highest first similarity and a first target window corresponding to the highest first similarity; and

determining a dynamic action corresponding to the highest first similarity as the preset dynamic action when (i) the highest first similarity is greater than a first preset threshold and (ii) a difference between the highest first similarity and a second similarity is greater than a second preset threshold, the second similarity being a similarity between respective first action feature information of a target window adjacent to the first target window and the preset feature information.

7. The electronic device according to claim 6 , wherein extracting the plurality of target windows from a target video further comprises:

classifying the plurality of video frames.

8. The electronic device according to claim 6 , wherein performing the action recognition on the respective plurality of consecutive video frames further comprises:

extracting, for each of the video frames, a plurality of body key points in the video frame;

performing action recognition on the video frame according to a distribution of the plurality of body key points, to obtain second action feature information of the video frame, the second action feature information is used for describing a static action comprised in the video frame; and

combining a plurality of second action feature information of the video frames in the target window, to obtain the first action feature information of the target window.

9. The electronic device according to claim 8 , wherein the preset dynamic action is performed using at least two preset body parts; and performing action recognition on the video frame according to the distribution comprises one or more of:

determining an angle between any two preset body parts in the video frame according to coordinates of the plurality of body key points in the video frame and body parts to which the plurality of body key points belong, and using the angle as the second action feature information;

obtaining a displacement amount between at least one of the plurality of body key points and a body key point corresponding to a reference video frame, and using the displacement amount as the second action feature information, the reference video frame being previous to the each video frame by a third preset quantity of video frames; and

obtaining a size of a reference body part in any two preset body parts and a distance between the any two preset body parts, and using a ratio of the distance to the size of the reference body part as the second action feature information.

10. The electronic device according to claim 6 , wherein the first action feature information is a first action matrix comprising M first action vectors, the preset feature information is a preset action matrix comprising N preset action vectors, M and N being positive integers, and the obtaining a similarity between the first action feature information of the each target window and preset feature information comprises:

creating a similarity matrix, the similarity matrix having M rows and N columns, or N rows and M columns;

obtaining, for a specified position corresponding to an i th first action vector and a j th preset action vector in the similarity matrix, a sum of a maximum similarity among similarities of a first position, a second position, and a third position and a similarity between the i th first action vector and the j th preset action vector as a similarity of the specified position, the first position being a position corresponding to an (i−1) th first action vector and the j th preset action vector, the second position being a position corresponding to the (i−1) th first action vector and a (j−1) th preset action vector, and the third position being a position corresponding to the i th first action vector and the (j−1) th preset action vector, i being a positive integer not less than 1 and not greater than M, and j being a positive integer not less than 1 and not greater than N; and

determining a similarity of a position corresponding to an M th first action vector and an N th preset action vector in the similarity matrix as the similarity between the first action feature information and the preset feature information.

11. A non-transitory computer-readable storage medium storing instructions that, when executed by one or more processors of an electronic device, cause the one or more processors to perform operations comprising:

extracting a plurality of target windows from a target video, each of the plurality of target windows comprising a respective plurality of consecutive video frames;

for each of the plurality of target windows:

performing action recognition on the respective plurality of consecutive video frames corresponding to the target window to obtain respective first action feature information of the target window, the first action feature information comprising movement of one or more body parts of a subject and is used for describing a dynamic action comprised in the target window; and

obtaining a similarity between the first action feature information of the target window and preset feature information, the preset feature information being used for describing a preset dynamic action;

determining, from the respective obtained similarities corresponding to the plurality of target windows, a highest first similarity and a first target window corresponding to the highest first similarity; and

determining a dynamic action corresponding to the highest first similarity as the preset dynamic action when (i) the highest first similarity is greater than a first preset threshold and (ii) a difference between the highest first similarity and a second similarity is greater than a second preset threshold, the second similarity being a similarity between respective first action feature information of a target window adjacent to the first target window and the preset feature information.

12. The non-transitory computer-readable storage medium according to claim 11 , wherein extracting the plurality of target windows from a target video further comprises:

classifying the plurality of video frames.

13. The non-transitory computer-readable storage medium according to claim 11 , wherein performing the action recognition on the respective plurality of consecutive video frames further comprises:

extracting, for each of the video frames, a plurality of body key points in the video frame;

performing action recognition on the video frame according to a distribution of the plurality of body key points, to obtain second action feature information of the video frame, the second action feature information is used for describing a static action comprised in the video frame; and

combining a plurality of second action feature information of the video frames in the target window, to obtain the first action feature information of the target window.

14. The non-transitory computer-readable storage medium according to claim 13 , wherein the preset dynamic action is performed using at least two preset body parts; and performing action recognition on the video frame according to the distribution comprises one or more of:

determining an angle between any two preset body parts in the video frame according to coordinates of the plurality of body key points in the video frame and body parts to which the plurality of body key points belong, and using the angle as the second action feature information;

obtaining a displacement amount between at least one of the plurality of body key points and a body key point corresponding to a reference video frame, and using the displacement amount as the second action feature information, the reference video frame being previous to the each video frame by a third preset quantity of video frames; and

obtaining a size of a reference body part in any two preset body parts and a distance between the any two preset body parts, and using a ratio of the distance to the size of the reference body part as the second action feature information.

15. The non-transitory computer-readable storage medium according to claim 11 , wherein the first action feature information is a first action matrix comprising M first action vectors, the preset feature information is a preset action matrix comprising N preset action vectors, M and N being positive integers, and the obtaining a similarity between the first action feature information of the each target window and preset feature information comprises:

creating a similarity matrix, the similarity matrix having M rows and N columns, or N rows and M columns;

obtaining, for a specified position corresponding to an i th first action vector and a j th preset action vector in the similarity matrix, a sum of a maximum similarity among similarities of a first position, a second position, and a third position and a similarity between the i th first action vector and the j th preset action vector as a similarity of the specified position, the first position being a position corresponding to an (i−1) th first action vector and the j th preset action vector, the second position being a position corresponding to the (i−1) th first action vector and a (j−1) th preset action vector, and the third position being a position corresponding to the i th first action vector and the (j−1) th preset action vector, i being a positive integer not less than 1 and not greater than M, and j being a positive integer not less than 1 and not greater than N; and

determining a similarity of a position corresponding to an M th first action vector and an N th preset action vector in the similarity matrix as the similarity between the first action feature information and the preset feature information.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Apr 21, 2022
From: LUO, JINGMIN; QIAO, LIANG; ZHU, XIAOLONG
To: TENCENT TECHNOLOGY (SHENZHEN) COMPANY LIMITED
Reel/Frame 059667/0419 →
Priority Claims (1)
CN 201910345010.1 · Apr 26, 2019 · national
Continuity (2)
Continuation PCTCN2020084996 · Apr 16, 2020
Related Publication 20210271892A1 · Sep 2, 2021