IP Library Granted Patent US 11,557,150
Granted Patent B2
US 11,557,150 · App. 16/641,828 · Granted Jan 17, 2023

Gesture control for communication with an autonomous vehicle on the basis of a simple 2D camera

Inventors: Erwin Kraft (Frankfurt, DE); Nicolai Harich (Mainz, DE); Sascha Semmler (Ulm, DE); Pia Dreiseitel (Eschborn, DE)
G06V40/20G06K9/6218G06K9/6269G06V20/58G06V40/10
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,557,150
App. No.
16/641,828
Granted
Jan 17, 2023
Kind
B2
Abstract

A method of recognizing gestures of a person from at least one image from a monocular camera, e.g. a vehicle camera, includes comp the steps: a) detecting key points of the person in the at least one image, b) connecting the key points to form a skeleton-like representation of body parts of the person, wherein the skeleton-like representation represents a relative position and a relative orientation of the respective body parts of the person, c) recognizing a gesture of the person from the skeleton-like representation of the person, and d) outputting a signal indicating the gesture.

Claims (32)

1. A method comprising the steps:

a) detecting key points of body parts of a person in at least one 2D image from a monocular vehicle camera mounted on a vehicle,

b) connecting the key points to form a skeleton-like representation of the body parts of the person, wherein the skeleton-like representation represents a relative position and a relative orientation of respective individual ones of the body parts of the person,

c) forming a first group of a first subset of the body parts, and forming a second group of a second subset of the body parts, wherein the first subset and the second subset include different body parts,

d) determining a first partial gesture of the person based on the first subset and generating a first feature vector based on the first partial gesture, determining a second partial gesture of the person based on the second subset and generating a second feature vector based on the second partial gesture, and

e) recognizing a final gesture of the person based on a final feature vector generated by merging the first feature vector and the second feature vector, wherein the detecting of the key points, the connecting of the key points, and the recognizing of the final gesture is performed based on 2D information from the at least one 2D image without any depth information,

f) producing a signal indicating the final gesture, and

g) actuating a control system of the vehicle or outputting a humanly perceivable information signal from the vehicle, automatically in response to and dependent on the signal indicating the final gesture.

2. The method according to claim 1 , wherein at least one of the body parts belongs to more than one of the groups.

3. The method according to claim 1 , wherein the partial gestures are static gestures, and further comprising adjusting a number of the groups.

4. The method according to claim 1 , further comprising assigning a respective feature vector respectively to each one of the groups, wherein the forming of the groups comprises combining the key points associated with the related ones of the body parts respectively in each respective one of the groups, and wherein said feature vector of a respective one of the groups is based on coordinates of the key points which are combined in the respective group.

5. The method according to claim 1 , wherein the recognizing of the final gesture is based on classifying the final feature vector.

6. The method according to claim 5 , wherein at least one of the body parts belongs to more than one of the groups.

7. The method according to claim 1 , further comprising estimating a viewing direction of the person based on the skeleton-like representation.

8. The method according to claim 7 , further comprising checking whether the viewing direction of the person is directed toward the monocular vehicle camera.

9. The method according to claim 7 , further comprising classifying the person as a distracted road user when the final gesture and the viewing direction indicate that the person has lowered his or her head and is looking at his or her hand.

10. The method according to claim 1 , wherein the recognizing of the final gesture is based on a gesture classification which has previously been trained.

11. The method according to claim 1 , wherein a number of the key points of the body parts of the person is a maximum of 20.

12. The method according to claim 1 , wherein the step g) comprises the actuating of the control system of the vehicle automatically in response to and dependent on the signal indicating the final gesture.

13. The method according to claim 1 , wherein the step g) comprises the outputting of the humanly perceivable information signal, which communicates, from the vehicle to the person, a warning or an acknowledgment indicating that the person has been detected, automatically in response to and dependent on the signal indicating the final gesture.

14. The method according to claim 1 , wherein the partial gestures are static gestures.

15. The method according to claim 1 , wherein the at least one image is a single still monocular image.

16. The method according to claim 1 , wherein the body parts of the person include at least one body part selected from the group consisting of an upper body, shoulders, upper arms, elbows, legs, thighs, hips, knees, and ankles.

17. A device configured:

a) to detect key points of body parts of a person in at least one 2D image from a monocular vehicle camera mounted on a vehicle,

b) to connect the key points to form a skeleton-like representation of the body parts of the person, wherein the skeleton-like representation represents a relative position and a relative orientation of respective individual ones of the body parts of the person,

c) to form a first group of a first subset of the body parts, and to form a second group of a second subset of the body parts, wherein the first subset and the second subset include different body parts,

d) to determine a first partial gesture of the person based on the first subset and generating a first feature vector based on the first partial gesture, to determine a second partial gesture of the person based on the second subset and generate a second feature vector based on the second partial gesture, and

e) to recognize a final gesture of the person based on a final feature vector generated by merging the first feature vector and the second feature vector, wherein the detecting of the key points, the connecting of the key points, and the recognizing of the final gesture is performed based on 2D information from the at least one 2D image without any depth information,

f) to produce a signal indicating the final gesture, and

g) to actuate a control system of the vehicle or to output a humanly perceivable information signal from the vehicle, automatically in response to and dependent on the signal indicating the final gesture.

18. A vehicle having a monocular vehicle camera and a device according to claim 17 .

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Feb 25, 2020
From: KRAFT, ERWIN; HARICH, NICOLAI; SEMMLER, SASCHA; DREISEITEL, PIA
To: CONTI TEMIC MICROELECTRONIC GMBH
Reel/Frame 051920/0518 →
Priority Claims (1)
DE 10 2017 216 000.4 · Sep 11, 2017 · national
Continuity (1)
Related Publication 20200394393A1 · Dec 17, 2020
Cited By (2)
US 12,456,333 US 12,505,686