IP Library Granted Patent US 12,443,285
Granted Patent B2
US 12,443,285 · App. 18/345,631 · Granted Oct 14, 2025

Human-machine interaction method and human-machine interaction apparatus

Inventors: Shuaihua Peng (Shanghai, CN); Hao Wu (Shanghai, CN)
G06F3/017B60K35/00B60W60/00253G06F3/0304B60K35/28B60K35/29B60K2360/176B60K2360/191
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,443,285
App. No.
18/345,631
Granted
Oct 14, 2025
Kind
B2
Abstract

This application provides a human-machine interaction method and the like. In one aspect, gesture action information of the user is detected by using an optical sensor of the object device; motion track information of the mobile terminal is detected by using a motion sensor of the mobile terminal. When the gesture action information matches the terminal motion track information, corresponding first control is executed.

Claims (66)

1. A human-machine interaction method, comprising:

obtaining motion track information of a mobile terminal, wherein the motion track information is obtained by using a motion sensor of the mobile terminal;

in response to determining that a predefined operation is performed on the mobile terminal, obtaining first gesture action information of a user, wherein the first gesture action information is obtained by using an optical sensor of an object device that interacts with the user, wherein the first gesture action information comprises gesture action form information and gesture action time information, and the motion track information comprises motion track form information and motion track time information;

determining whether a similarity of a first form of a motion of the mobile and a second form of a gesture exists by processing the gesture action form information and the motion track form information using a machine learning model;

determining whether a consistency between a first time of the motion of the mobile and a second time of the gesture by comparing a preset threshold with a difference between the gesture action form information and the motion track time information;

determining that the first gesture action information matches the motion track information in response to determining that the similarity exists and the consistency exists; and

executing first control when the first gesture action information matches the motion track information, wherein the first control comprises control executed according to a control instruction corresponding to the first gesture action information.

2. The human-machine interaction method according to claim 1 , further comprising:

recognizing, by using the optical sensor, the user corresponding to the first gesture action information; and

when the first gesture action information matches the motion track information, authenticating the user corresponding to the first gesture action information as a valid user.

3. The human-machine interaction method according to claim 2 , further comprising:

obtaining second gesture action information of the valid user by using the optical sensor, wherein the second gesture action information is later than the first gesture action information in terms of time, and the first control comprises control executed according to a control instruction corresponding to the second gesture action information.

4. The human-machine interaction method according to claim 2 , wherein

the object device is a vehicle, the vehicle comprises a display, and the first control comprises displaying, on the display, an environment image comprising the valid user, wherein the valid user is highlighted in the environment image.

5. The human-machine interaction method according to claim 2 , wherein

the object device is a vehicle, and the first control comprises enabling the vehicle to move autonomously toward the valid user.

6. The human-machine interaction method according to claim 1 , wherein

the obtaining first gesture action information comprises:

obtaining location information of the mobile terminal from the mobile terminal; and

adjusting the optical sensor based on the location information to include the mobile terminal within a detection range of the optical sensor.

7. The human-machine interaction method according to claim 1 , wherein

the obtaining first gesture action information comprises:

when the motion track information is obtained but the first gesture action information is not obtained within predefined time, sending, to the mobile terminal, information for requesting to perform a first gesture action.

8. The human-machine interaction method according to claim 1 ,

further comprising: authenticating validity of an identity (ID) of the mobile terminal, wherein

the obtaining motion track information of a mobile terminal comprises obtaining motion track information of a mobile terminal that has a valid ID.

9. A human-machine interaction apparatus, comprising:

at least one processor; and

a memory coupled to the at least one processor and storing programming instructions for execution by the at least one processor, wherein the programming instructions instruct the at least one processor to perform the following operations:

obtaining motion track information of a mobile terminal, wherein the motion track information is obtained by using a motion sensor of the mobile terminal;

in response to determining that a predefined operation is performed on the mobile terminal, obtaining first gesture action information of a user, wherein the first gesture action information is obtained by using an optical sensor of an object device that interacts with the user, wherein the first gesture action information comprises gesture action form information and gesture action time information, and the motion track information comprises motion track form information and motion track time information;

determining whether a similarity of a first form of a motion of the mobile and a second form of a gesture exists by processing the gesture action form information and the motion track form information using a machine learning model;

determining whether a consistency between a first time of the motion of the mobile and a second time of the gesture by comparing a preset threshold with a difference between the gesture action form information and the motion track time information;

determining that the first gesture action information matches the motion track information in response to determining that the similarity exists and the consistency exists; and

executing first control when the first gesture action information matches the motion track information, wherein the first control comprises control executed according to a control instruction corresponding to a first gesture action.

10. The human-machine interaction apparatus according to claim 9 , wherein the operations further comprise:

recognizing, by using the optical sensor, the user corresponding to the first gesture action information; and

when the first gesture action information matches the motion track information, authenticating the user corresponding to the first gesture action information as a valid user.

11. The human-machine interaction apparatus according to claim 10 , wherein the operations further comprise:

obtaining second gesture action information of the valid user by using the optical sensor, wherein the second gesture action information is later than the first gesture action information in terms of time, and

the first control comprises control executed according to a control instruction corresponding to the second gesture action information.

12. The human-machine interaction apparatus according to claim 10 , wherein

the object device is a vehicle, the vehicle comprises a display, and the first control comprises displaying, on the display, an environment image comprising the valid user, wherein the valid user is highlighted in the environment image.

13. The human-machine interaction apparatus according to claim 10 , wherein

the object device is a vehicle, and the first control comprises enabling the vehicle to move autonomously toward the valid user.

14. The human-machine interaction apparatus according to claim 9 , wherein the programming instructions instruct the at least one processor to perform the following operation:

obtaining location information of the mobile terminal from the mobile terminal; and

adjusting the optical sensor based on the location information to include the mobile terminal within a detection range of the optical sensor.

15. The human-machine interaction apparatus according to claim 9 , wherein

when the motion track information is obtained but the first gesture action information is not obtained within predefined time, information for requesting to perform the first gesture action is sent to the mobile terminal.

16. The human-machine interaction apparatus according to claim 9 , wherein the programming instructions instruct the at least one processor to perform the following operation:

authenticating validity of an identity (ID) of the mobile terminal, wherein

obtaining motion track information of a mobile terminal that has a valid ID.

17. One or more non-transitory computer-readable media containing instructions which, when executed, cause an electronic device to perform operations comprising:

obtaining motion track information of a mobile terminal, wherein the motion track information is obtained by using a motion sensor of the mobile terminal;

in response to determining that a predefined operation is performed on the mobile terminal, obtaining first gesture action information of a user, wherein the first gesture action information is obtained by using an optical sensor of an object device that interacts with the user, wherein the first gesture action information comprises gesture action form information and gesture action time information, and the motion track information comprises motion track form information and motion track time information;

determining whether a similarity of a first form of a motion of the mobile and a second form of a gesture exists by processing the gesture action form information and the motion track form information using a machine learning model;

determining whether a consistency between a first time of the motion of the mobile and a second time of the gesture by comparing a preset threshold with a difference between the gesture action form information and the motion track time information;

determining that the first gesture action information matches the motion track information in response to determining that the similarity exists and the consistency exists; and

executing first control when the first gesture action information matches the motion track information, wherein the first control comprises control executed according to a control instruction corresponding to the first gesture action information.

18. The one or more non-transitory computer-readable media according to claim 17 , wherein the operations further comprise:

recognizing the user corresponding to the first gesture action information; and

when the first gesture action information matches the motion track information, authenticating the user corresponding to the first gesture action information as a valid user.

19. The one or more non-transitory computer-readable media according to claim 18 , wherein the operations further comprise:

obtaining second gesture action information of the valid user by using the optical sensor, wherein the second gesture action information is later than the first gesture action information in terms of time, and the first control comprises control executed according to a control instruction corresponding to the second gesture action information.

20. The one or more non-transitory computer-readable media according to claim 18 , wherein the object device is a vehicle, the vehicle comprises a display, and the first control comprises displaying, on the display, an environment image comprising the valid user, wherein the valid user is highlighted in the environment image.

Assignments (3)
CHANGE OF NAME Recorded Apr 28, 2026
From: SHENZHEN YINWANG INTELLIGENT TECHNOLOGIES CO., LTD.
To: YINWANG INTELLIGENT TECHNOLOGIES CO., LTD.
Reel/Frame 075492/0796 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Nov 12, 2024
From: HUAWEI TECHNOLOGIES CO., LTD.
To: SHENZHEN YINWANG INTELLIGENT TECHNOLOGIES CO., LTD.
Reel/Frame 069336/0125 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Sep 5, 2023
From: PENG, SHUAIHUA; WU, HAO
To: HUAWEI TECHNOLOGIES CO., LTD.
Reel/Frame 064795/0972 →
Continuity (2)
Continuation PCTCN2021070188 · Jan 4, 2021
Related Publication 20230350498A1 · Nov 2, 2023
References Cited (19)
US 20130069931A1 · Wilson · 2013 [cited by examiner]
US 20170136842A1 · Anderson · 2017 [cited by examiner]
US 20220164035A1 · Katz · 2022 [cited by examiner]
US 20220404949A1 · Berquam · 2022 [cited by examiner]
CN 105313835A · 2016 [cited by applicant]
CN 111332314A · 2020 [cited by applicant]
CN 111399636A · 2020 [cited by applicant]
CN 111741167A · 2020 [cited by applicant]
CN 111913559A · 2020 [cited by applicant]
DE 102018100335A1 · 2019 [cited by applicant]
Chang et al. “Eye on You: Fusing gesture data from depth camera and inertial for person identification” (Year: 2018). [cited by examiner]
Chang et al., “Eye on You: Fusing Gesture Data from Depth Camera and Inertial Sensors for Person Identification,” 2018 IEEE International Conference On Robotics and Automation (ICRA), May 21, 2018 , pp. 2021-2026. [cited by applicant]
Partial Supplementary European Search Report in European Appln No. 21912431.0, dated Jan. 3, 2024, 12 pages. [cited by applicant]
Aslan et al., “Design and exploration of mid-air authentication gestures,” ACM Transactions on Interactive Intelligent Systems (TiiS), Sep. 14, 2016, 6(3):1-22. [cited by applicant]
Liu et al., “Dynamic-hand-gesture authentication dataset and benchmark,” IEEE Transactions on Information Forensics and Security, Nov. 5, 2020, 16:1550-62. [cited by applicant]
Extended European Search Report in European Appln No. 21912431.0, dated Mar. 15, 2024, 16 pages. [cited by applicant]
Zhang, “Vision-based markerless gesture recognition,” Changchun Jilin University Press, Jun. 2016, 4 pages (with English abstract). [cited by applicant]
Xu et al., “Network Security and Virus Prevention (Sixth Edition),” Shanghai Jiao Tong University Press, 21st edition, Jul. 1, 2016, 5 pages (with English abstract). [cited by applicant]
International Search Report and Written Opinion in International Appln. No. PCT/CN2021/070188, mailed on Oct. 13, 2021, 15 pages (with English translation). [cited by applicant]