IP Library › Granted Patent US 12,572,220
Granted Patent B2
US 12,572,220 · App. 18/899,741 · Granted Mar 10, 2026

Systems and methods for multi-modal interaction analysis

Inventors: Dongfang Zhao (San Diego, CA); Sobhan Soleymani (San Diego, CA); Razieh Kaviani Baghbaderani (San Jose, CA); Yangwen Liang (San Diego, CA); Shuangquan Wang (San Diego, CA); Kee-Bong Song (San Diego, CA); Yanlin Zhou (San Diego, CA); Mostafa El-Khamy (San Diego, CA)
Assignee: Samsung Electronics Co., Ltd.
G06F3/017G06V10/32G06V10/82G06V20/64G06V40/28G06V10/12
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,572,220
App. No.
18/899,741
Granted
Mar 10, 2026
Kind
B2
Abstract

A method and device are disclosed for interaction analysis. The method includes receiving, by an interaction-analysis circuit of a device, input image data from an interaction sensor, generating, by the interaction-analysis circuit, a cropped-and-resized image based on detecting a three-dimensional (3D) object in the input image data, generating, by the interaction-analysis circuit, a 3D pose-estimation structure based on the cropped-and-resized image, determining, by the interaction-analysis circuit, a gesture classification based on the 3D pose-estimation structure, and transmitting the gesture classification.

Claims (68)

1 . A method for interaction analysis, the method comprising:

receiving, by an interaction-analysis circuit of a device, input image data from an interaction sensor;

generating, by the interaction-analysis circuit, a cropped-and-resized image based on detecting a three-dimensional (3D) object in the input image data;

generating, by the interaction-analysis circuit, a 3D pose-estimation structure based on receiving, as inputs to the interaction-analysis circuit:

the cropped-and-resized image; and

one or more interaction-sensor parameters associated with determining depth information, the one or more interaction-sensor parameters comprising at least one of a translation parameter, a focal length, a pixel size, a rotation parameter, a camera position, or a camera orientation;

determining, by the interaction-analysis circuit, a gesture classification based on the 3D pose-estimation structure; and

transmitting the gesture classification.

2 . The method of claim 1 , wherein the one or more interaction-sensor parameters are associated with determining depth information associated with the cropped-and-resized image.

3 . The method of claim 1 , wherein:

the interaction sensor comprises a stereo camera; and

the input image data comprises a first image associated with a left sensor and a second image associated with a right sensor.

4 . The method of claim 1 , wherein:

the interaction sensor comprises a depth sensing camera; and

the input image data comprises depth image data.

5 . The method of claim 1 , wherein:

the interaction sensor comprises a light detection and ranging (LiDAR) sensor; and

the input image data comprises cloud point data.

6 . The method of claim 1 , further comprising generating, by a first machine-learning model, a bounding-box image and a confidence level, based on the input image data,

wherein the detecting the 3D object is performed by the first machine-learning model.

7 . The method of claim 6 , wherein the first machine-learning model comprises a convolutional neural network (CNN).

8 . The method of claim 6 , further comprising:

cropping and resizing the bounding-box image to generate the cropped-and-resized image; and

inputting the cropped-and-resized image to a second machine-learning model,

wherein the generating the 3D pose-estimation structure is performed, by the second machine-learning model, based on the cropped-and-resized image.

9 . The method of claim 8 , further comprising outputting, by a third machine-learning model, a classification of a gesture,

wherein the determining the gesture is performed by the third machine-learning model and based on a plurality of outputs from the second machine-learning model.

10 . A device for interaction analysis, the device comprising:

a memory; and

a processor communicatively coupled to the memory, wherein the processor is configured to:

receive input image data from an interaction sensor;

generate a cropped-and-resized image based on detecting a three-dimensional (3D) object in the input image data;

generate a 3D pose-estimation structure based on receiving, as inputs:

the cropped-and-resized image; and

one or more interaction-sensor parameters associated with determining depth information, the one or more interaction-sensor parameters comprising at least one of a translation parameter, a focal length, a pixel size, a rotation parameter, a camera position, or a camera orientation;

determine a gesture classification based on the 3D pose-estimation structure; and

transmit the gesture classification.

11 . The device of claim 10 , wherein the one or more interaction-sensor parameters are associated with the cropped-and-resized image.

12 . The device of claim 10 , wherein the processor is configured to:

detect the 3D object based on a first machine-learning model; and

generate a bounding-box image and a confidence level, based on providing the input image data to the first machine-learning model.

13 . The device of claim 10 , wherein the processor is configured to:

crop and resize a bounding-box image to generate the cropped-and-resized image; and

generate the 3D pose-estimation structure based on providing the cropped-and-resized image to a second machine-learning model.

14 . The device of claim 13 , wherein the processor is configured to:

determine the gesture classification based on providing the 3D pose-estimation structure to a third machine-learning model; and

output the gesture classification from the third machine-learning model.

15 . The device of claim 14 , wherein the processor is configured to determine the gesture classification based on a plurality of outputs from the second machine-learning model.

16 . A system for interaction analysis, the system comprising:

a memory; and

a means for processing configured to:

receive input image data from an interaction sensor;

generate a cropped-and-resized image based on detecting a three-dimensional (3D) object in the input image data;

generate a 3D pose-estimation structure based on receiving, as inputs:

the cropped-and-resized image; and

one or more interaction-sensor parameters associated with determining depth information, the one or more interaction-sensor parameters comprising at least one of a translation parameter, a focal length, a pixel size, a rotation parameter, a camera position, or a camera orientation;

determine a gesture classification based on the 3D pose-estimation structure; and

transmit the gesture classification.

17 . The system of claim 16 , wherein the one or more interaction-sensor parameters are associated with determining depth information associated with the cropped-and-resized image.

18 . The system of claim 16 , wherein the means for processing is configured to:

detect the 3D object based on a first machine-learning model; and

generate a bounding-box image and a confidence level, based on providing the input image data to the first machine-learning model.

19 . The system of claim 17 , wherein the means for processing is configured to:

crop and resize a bounding-box image to generate the cropped-and-resized image; and

generate the 3D pose-estimation structure based on providing the cropped-and-resized image to a second machine-learning model.

20 . The system of claim 18 , wherein the means for processing is configured to:

determine the gesture classification based on providing the 3D pose-estimation structure to a third machine-learning model; and

output the gesture classification from the third machine-learning model.

Assignments (2)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jan 15, 2026
From: SAMSUNG SEMICONDUCTOR, INC.
To: SAMSUNG ELECTRONICS CO., LTD.
Reel/Frame 073489/0870 →
SAMSUNG SEMICONDUCTOR INC. EMPLOYEE AGREEMENT CONFIDENTIAL INFORMATION AND INVENTIONS Recorded Jan 15, 2026
From: WANG, SHUANGQUAN
To: SAMSUNG SEMICONDUCTOR, INC.
Reel/Frame 074388/0063 →
Continuity (2)
Provisional Application 63636601 · Apr 19, 2024
Related Publication 20250328198A1 · Oct 23, 2025
References Cited (18)
US 10796482B2 · Ge et al. · 2020 [cited by applicant]
US 10891473B2 · Zhang et al. · 2021 [cited by applicant]
US 11551374B2 · Li et al. · 2023 [cited by applicant]
US 11841920B1 · Marsden et al. · 2023 [cited by applicant]
US 11854308B1 · Marsden et al. · 2023 [cited by applicant]
US 20110107216A1 · Bi · 2011 [cited by examiner]
US 20110291926A1 · Gokturk · 2011 [cited by examiner]
US 20180052520A1 · Amores Llopis · 2018 [cited by examiner]
US 20210180942A1 · Lin · 2021 [cited by examiner]
US 20220291752A1 · Agu · 2022 [cited by examiner]
US 20220343687A1 · Zhou et al. · 2022 [cited by applicant]
US 20230044664A1 · Sinha et al. · 2023 [cited by applicant]
US 20230206613A1 · Mahbub · 2023 [cited by examiner]
US 20230252737A1 · Dreyer · 2023 [cited by examiner]
US 20240126381A1 · Korrapati · 2024 [cited by examiner]
Zhang, Fan, et al. Google Research. Mediapipe hands: On-device real-time hand tracking. arXiv preprint arXiv:2006.10214 (2020). [cited by applicant]
Gao, Qing, et al. “Dynamic hand gesture recognition based on 3D hand pose estimation for human-robot interaction.” IEEE Sensors Journal 22.18 (2021): 17421-17430. [cited by applicant]
Sharma, Ashish, et al. “Hand gesture recognition using image processing and feature extraction techniques.” Procedia Computer Science 173 (2020): 181-190. [cited by applicant]