IP Library Granted Patent US 12670735
Granted Patent B2
US 12670735 · App. 18/714,663 · Granted Jun 30, 2026

Method and apparatus for optical detection and analysis in a movement environment

Inventors: Pascal Siekmann (Bueren, DE); Gerhard Dick (Lippstadt, DE)
Assignee: AIRIS GmbH
G06V20/64G06T7/292G06T7/70G06V10/147G06V10/26G06V10/32G06V10/751G06V10/82G06V10/955G06V40/10G06T2207/10016G06T2207/20084G06T2207/20132G06T2207/30196
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12670735
App. No.
18/714,663
Granted
Jun 30, 2026
Kind
B2
Abstract

Provided is a device and a corresponding method for optical recognition and analysis, in particular of bodies, body parts and/or joints of a person, surfaces and objects in a 3-dimensional movement area in real time, with one or more cameras with fisheye lens, which are arranged at a distance and position from the movement area specifically selected for the application, for determining 2D positions of the bodies, body parts and/or joints, surfaces and objects by means of camera images, at least one data processing unit being provided, with devices for calculating neural networks and artificial intelligence in real time, for applying several AIs on captured camera images, and with at least one interface for one or more audio/video feedback units for outputting data or audio/video feedback on the analysis results by means of the audio/video feedback unit.

Claims (50)

1 . A method for optical recognition and analysis of bodies, body parts and/or joints of a person, surfaces and objects in 3-dimensional space in a motion range in real time, comprising the following steps:

a. Providing one or more cameras with fisheye lens,

b. Arranging the one or more cameras at a distance and position from the movement area specifically selected for the application,

c. Providing at least one data processing unit equipped with devices for calculating neural networks and artificial intelligences in real time,

d. Providing at least one audio/video feedback unit,

e. Determining a 2D positions of bodies, body parts and/or joints, surfaces and objects using camera images on the respective camera or cameras,

f. Applying several AIs to captured camera images for recognition and/or analysis with the steps

i. Recognizing and cropping bodies/objects on camera images,

ii. Recognizing body part/joint and/or object positions on image crops,

iii. Determination of the 3-dimensional positions of the recognized body part, joint and/or object positions in the movement area,

iv. Analyzing the movement of the 3-dimensional body part/joint and/or object positions in the movement area,

g. Outputting data or audio/video feedback on the analysis results using the audio/video feedback unit,

wherein

the AIs comprise object recognition AIs and wherein determining the 2D positions of the bodies, body parts, joints and/or surfaces and objects in the range of motion by camera images on the respective camera or cameras comprises the following steps:

h. Preparing the camera images for input into the object recognition AI by reducing the size of the images using GPU (Graphic Processor Unit) acceleration,

i. Cropping of bodies, joints, body parts and/or objects, correction of the orientation and scaling of the image section size using GPU acceleration,

j. Recognizing body part and/or joint positions on the image sections of the bodies by means of further object recognition AIs that are trained to recognize body parts and joints, and

k. Reverse calculating the detected 2D body parts and/or joint points to the original image size and orientation of the camera image.

2 . The method according to claim 1 , wherein the determination of the 3-dimensional positions of the detected body part/joint and/or object positions in the range of motion is carried out by 3D real-time estimation.

3 . The method according to claim 1 , wherein at least two cameras with fisheye lens are provided and the determination of the 3-dimensional positions of the recognized body part, joint and/or object positions in the range of motion is carried out by triangulation.

4 . The method according to claim 1 , wherein the at least one camera with fisheye lens has a diagonal field of view (FOV) of at least 220°.

5 . The method according to claim 4 , wherein the at least one camera with a fisheye lens has a field of view of up to 180° horizontally and 180° vertically.

6 . The method according to claim 1 , wherein the distance and position of the one or more cameras specifically selected for the application is chosen to cover the area in front of the camera or cameras in a radius of up to 10 meters.

7 . The method according to claim 1 , wherein, when using several cameras, the body parts and joints are additionally recognized by means of feature matching, wherein the joint points recognized by the AIs are searched for by a first camera by means of feature matching on all further cameras.

8 . The method according to claim 7 , wherein the projection of the epipolar line of the first camera is calculated on all further cameras to minimize the computational effort and areas to be analyzed are defined, wherein the feature matching of the respectively detected joint point takes place along these epipolar line areas.

9 . The method according to claim 3 , wherein the recognition of 3-D surfaces is carried out using a further AI, which receives two camera images as input and creates a spatial model from them.

10 . The method according to any claim 1 , wherein the object recognition AI is provided with data on the relations of human joints to each other and performs a pose correction.

11 . The method according to any claim 3 , wherein the cameras are arranged such that their fields of view intersect.

12 . The method according to claim 1 , wherein a second data processing unit with graphics processors is provided for controlling games and sports programs based on the 3D data with the body, joints, objects, surfaces and for displaying the feedback by means of the audio/video feedback unit.

13 . A device for carrying out a method for optical recognition and analysis, in particular of bodies, body parts and/or joints of a person, surfaces and objects in the 3-dimensional motion range in real time, comprising:

a. a housing;

b. one or more cameras ( 2 ) with fisheye lens ( 3 ), which are arranged in the housing ( 34 ) at a distance and position from the movement area specifically selected for the application, for determining 2D positions of the bodies, body parts and/or joints, surfaces and objects in the motion area by means of camera images;

c. at least one data processing unit provided with means for calculating neural networks and artificial intelligences in real time, for applying several AIs to captured camera images for recognition and/or analysis, comprising the steps of

i. Recognizing and cropping bodies/objects on camera images;

ii. Recognizing body parts, joints and/or object positions on image sections;

iii. Determining the 3-dimensional positions of the recognized body part, joint and/or object positions;

iv. Analyzing the movement of the 3-dimensional body part, joint and/or object positions; and

d. at least one interface for one or more audio/video feedback units for outputting data or audio/video feedback on the analysis results by means of the audio/video feedback unit,

wherein

the AIs comprise object recognition AIs for determining 2D positions of the bodies, body parts, joints and/or surfaces and objects by camera images on the respective camera or cameras ( 2 ) by means of the following steps:

e. Preparing the camera images for input into the object recognition AI, by reducing the size of the images using GPU (Graphic Processor Unit) acceleration,

f. Cropping of bodies, joints, body parts and/or objects, correction of the orientation and scaling of the image section size using GPU acceleration,

g. Recognizing body part and/or joint positions on the image sections of the bodies by means of further object recognition AIs that are trained to recognize body parts and joints, and

h. Reverse calculating the detected 2D body part and/or joint points to the original image size and orientation of the camera image.

14 . The device according to claim 13 , wherein at least two cameras ( 2 ) with fisheye lens are provided for determining the 3-dimensional positions of the recognized body part, joint and/or object positions in the motion range by triangulation.

15 . The device according to claim 13 , wherein the at least one camera with fisheye lens has a diagonal field of view (FOV) of up to 220°.

16 . The device according to claim 13 , wherein the at least one camera with fisheye lens has a field of view of up to 180° horizontally and 180° vertically.

17 . The device according to claim 13 , wherein the distance and position of the one or more cameras in the housing specifically selected for the application is chosen to cover the movement space in front of the one or more cameras in a radius of up to 10 meters.

18 . The device according to claim 13 , wherein the object recognition AI is provided with data on the relations of human joints to each other and performs a pose correction.

19 . The device according to claim 13 , wherein the cameras are arranged in the housing in such a way that their fields of view intersect.