IP Library › Granted Patent US 12,361,359
Granted Patent B2
US 12,361,359 · App. 18/752,176 · Granted Jul 15, 2025

Vision-based hand grip recognition method and system for industrial ergonomics risk identification

Inventors: Julia Penfield (Seattle, WA); Francis Seunghyun Baek (Ann Arbor, MI); Richard Thomas Barker (West Chester, OH); Daeho Kim (Toronto, CA); SangHyun Lee (Ann Arbor, MI)
Assignee: VelocityEHS Holdings, Inc.
G06Q10/0635G06Q10/06398G06T7/70G06V10/764G06V40/11
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,361,359
App. No.
18/752,176
Granted
Jul 15, 2025
Kind
B2
Abstract

A system, comprising: a computing device configured to obtain video signals of a worker performing a hand-related job at a workplace; and a computing server system configured to receive and process the video signals to identify hand grips and wrist bending involved in the job, determine a hand grip type for each identified hand grip, obtain force information relating to each identified hand grip, determine neutral or hazardous wrist bending based at least upon the wrist bending and the hand grip force information, calculate a percent maximum strength for each identified hand grip, calculate frequencies and durations of each identified hand grip and wrist bending, and determine ergonomic risks of the hand-related job accordingly.

Claims (53)

1. A system comprising:

a computing device comprising:

a non-transitory computer-readable storage medium configured to store an application program; and

a processor coupled to the non-transitory computer-readable storage medium and configured to control a plurality of modules to execute instructions of the application program to obtain video signals of a worker performing a work task at a workplace; and

a computing server system configured to:

receive the video signals,

obtain one or more image frames from the video signals,

incorporate a deep learning model to identify joint positions and postures of a body part involved in the work task based at least upon the one or more image frames, wherein the deep learning model is trained and validated by:

recording a number of body part-related activities to obtain images capturing different joint positions and postures of the body part,

determining 3D coordinates of joints in each image based at least upon a tracking confidence set for at least one detected body part in sequential images, and

normalizing the 3D coordinates of joints to calculate confidence values corresponding to a number of postures of the body part,

calculate a frequency and a duration of each identified posture of the body part,

determine ergonomic risks of the work task based at least upon the frequency and the duration of each identified posture of the body part, and

present, via the application program of the computing device, assessment results of the ergonomic risks to a user.

2. The system of claim 1 , wherein the application program of the computing device is configured to directly record the video signals of the worker performing the work task.

3. The system of claim 1 , wherein the application program of the computing device is configured to obtain the video signals of the worker performing the work task captured by another computing or imaging device.

4. The system of claim 1 , wherein the computing server system is further configured to determine a level of the ergonomic risks associated with the body part, wherein the level of the ergonomic risks associated with the body part is one of a low, medium or high level.

5. The system of claim 1 , wherein the computing server system is configured to use the deep learning model to identify the joint positions of the body part by determining X, Y, and Z coordinates of each joint position of the body part in the one or more image frames.

6. The system of claim 5 , wherein the X, Y, and Z coordinates of each joint position are overlaid on the body part to visualize the ergonomic risks.

7. The system of claim 1 , wherein the computing server system is further configured to provide ergonomic risk control recommendations to mitigate the ergonomic risks.

8. A computer-implemented method comprising:

obtaining, by a processor of a computing device deployed within a communication network, video signals of a worker performing a work task at a workplace;

receiving, by a computing server system deployed within the communication network, the video signals;

obtaining, by the computing server system, one or more image frames from the video signals;

incorporating, by the computing server system, a deep learning model to identify joint positions and postures of a body part involved in the work task based at least upon the one or more image frames, wherein the deep learning model is trained and validated by:

recording a number of body part-related activities to obtain images capturing different joint positions and postures of the body part,

determining 3D coordinates of joints in each image based at least upon a tracking confidence set for at least one detected body part in sequential images, and

normalizing the 3D coordinates of joints to calculate confidence values corresponding to a number of postures of the body part;

calculating a frequency and a duration of each identified posture of the body part;

determining ergonomic risks of the work task based at least upon the frequency and the duration of each identified posture of the body part; and

presenting, via the computing device, assessment results of the ergonomic risks to a user.

9. The computer-implemented method of claim 8 , wherein obtaining, by the computing device, the video signals comprises directly recording, by an application program of the computing device, the video signals of the worker performing the work task.

10. The computer-implemented method of claim 8 , wherein the video signals are obtained from another computing or imaging device.

11. The computer-implemented method of claim 8 , further comprising determining a level of the ergonomic risks associated with the body part, wherein the level of the ergonomic risks associated with the body part is one of a low, medium or high level.

12. The computer-implemented method of claim 8 , wherein incorporating, by the computing server system, the deep learning model to identify the joint positions of the body part comprises determining X, Y, and Z coordinates of each joint position of the body part in the one or more image frames.

13. The computer-implemented method of claim 12 , wherein the X, Y, and Z coordinates of each joint position are overlaid on the body part to visualize the ergonomic risks.

14. The computer-implemented method of claim 8 , further comprising providing, by the computing server system, ergonomic risk control recommendations to mitigate the ergonomic risks.

15. A non-transitory computer readable medium storing computer executable instructions for a system deployed within a communication network, the instructions being configured for:

obtaining, by a processor of a computing device deployed within a communication network, video signals of a worker performing a work task at a workplace;

receiving, by a computing server system deployed within the communication network, the video signals;

obtaining, by the computing server system, one or more image frames from the video signals;

incorporating, by the computing server system, a deep learning model to identify joint positions and postures of a body part involved in the work task based at least upon the one or more image frames, wherein the deep learning model is trained and validated by:

recording a number of body part-related activities to obtain images capturing different joint positions and postures of the body part,

determining 3D coordinates of joints in each image based at least upon a tracking confidence set for at least one detected body part in sequential images, and

normalizing the 3D coordinates of joints to calculate confidence values corresponding to a number of postures of the body part;

calculating a frequency and a duration of each identified posture of the body part;

determining ergonomic risks of the work task based at least upon the frequency and the duration of each identified posture of the body part; and

presenting, via the computing device, assessment results of the ergonomic risks to a user.

16. The non-transitory computer readable medium of claim 15 , wherein the instructions for obtaining, by the computing device, the video signals comprise instructions for directly recording, by an application program of the computing device, the video signals of the worker performing the work task.

17. The non-transitory computer readable medium of claim 15 , wherein the video signals are obtained from another computing or imaging device.

18. The non-transitory computer readable medium of claim 15 , further comprising determining a level of the ergonomic risks associated with the body part, wherein the level of the ergonomic risks associated with the body part is one of a low, medium or high level.

19. The non-transitory computer readable medium of claim 15 , wherein the instructions for incorporating, by the computing server system, the deep learning model to identify the joint positions of the body part comprise instructions for determining X, Y, and Z coordinates of each joint position of the body part in the one or more image frames, wherein the X, Y, and Z coordinates of each joint position are overlaid on the body part to visualize the ergonomic risks.

20. The non-transitory computer readable medium of claim 15 , further comprising instructions for providing, by the computing server system, ergonomic risk control recommendations to mitigate the ergonomic risks.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jun 13, 2025
From: PENFIELD, JULIA; BAEK, FRANCIS SEUNGHYUN; BARKER, RICHARD THOMAS; KIM, DAEHO; LEE, SANGHYUN
To: VELOCITYEHS HOLDINGS, INC.
Reel/Frame 071410/0260 →
Continuity (2)
Continuation 18457855 · Aug 29, 2023
Related Publication 20250078003A1 · Mar 6, 2025
References Cited (35)
US 7804998B2 · Mundermann et al. · 2010 [cited by applicant]
US 8139067B2 · Anguelov et al. · 2012 [cited by applicant]
US 8180714B2 · Corazza et al. · 2012 [cited by applicant]
US 8384714B2 · De Aguiar et al. · 2013 [cited by applicant]
US 11324439B2 · Diaz-Arias et al. · 2022 [cited by applicant]
US 11482048B1 · Diaz-Arias · 2022 [cited by examiner]
US 11763235B1 · Penfield · 2023 [cited by examiner]
US 20080031512A1 · Mundermann · 2008 [cited by examiner]
US 20080180448A1 · Anguelov · 2008 [cited by examiner]
US 20100020073A1 · Corazza et al. · 2010 [cited by applicant]
US 20110208444A1 · Solinsky · 2011 [cited by examiner]
US 20200327465A1 · Baek · 2020 [cited by examiner]
US 20220079510A1 · Robillard et al. · 2022 [cited by applicant]
US 20220237537A1 · Baek · 2022 [cited by examiner]
US 20220386942A1 · Diaz-Arias · 2022 [cited by examiner]
US 20230097454A1 · Kim · 2023 [cited by examiner]
WO 2009140261A1 · 2009 [cited by applicant]
WO WO2022249555A1 · 2022 [cited by examiner]
JoonOh Seo, SangHyun Lee, “Automated postural ergonomic risk assessment using vision-based posture classification”, Automation in Construction (Year: 2021). [cited by examiner]
Eldar et al.; “Ergonomic design visualization mapping-developing an assistive model for design activities”; International journal of industrial ergonomics (Year: 2019). [cited by examiner]
Muendermann et al.; A New Accurate Method of 3D Full Body Motion Capture for Animation; Standford Office of Technology Licensing, https://techfinder.stanford.edu/technology/new-accurate-method-3d-full-body-motion-captur… [cited by applicant]
Markerless Motion Capture, BioMotion Laboratory Mechanical Engineering, Stanford University, http://web.stanford.edu/group/biomotion/markerless.html. [cited by applicant]
Automatic Generation of Human Models for Motion Capture, Biomechanics and Animation, ioMotion Laboratory Mechanical Engineering, Stanford University, https://techfinder.stanford.edu/technology/automatic-generation-human… [cited by applicant]
Mesh-based Performance Capture from Multi-view Video, BioMotion Laboratory Mechanical Engineering, Stanford University, https://techfinder.stanford.edu/technology/mesh-based-performance-capture-multi-view-video. [cited by applicant]
Anguelov et al., SCAPE: Shape Completion and Animation of People, In ACM Transactions on Graphics (TOG) (vol. 24, No. 3, pp. 408-416). ACM. [cited by applicant]
Corazza et al. A markerless motion capture system to study musculoskeletal biomechanics: visual hull and simulated annealing approach, Annals of Biomedical Engineering, 2006,34(6):1019-29. [cited by applicant]
Corazza et al., A Framework For The Functional Identification Of Joint Centers Using Markerless Motion Capture, Validation For The Hip Joint, Journal of Biomechanics, 2007. [cited by applicant]
Mundermann et al., The Evolution of methods for the capture of human movement leading to markerless motion capture for biomechanical applications. Journal of NeuroEngineering and Rehabilitation, 3(1), 2006. [cited by applicant]
Mundermann, et al., Accurately measuring human movement using articulated ICP with soft-joint constraints and a repository of articulated models, CVPR 2007. [cited by applicant]
Robson, Motion-capture system adds costume to the drama, New Scientist, Technology, May 29, 2008. [cited by applicant]
De Aguia et al., Performance capture from sparse multi-view video. ACM Trans. Graph. 27, 3 (Aug. 2008), 1-10. https://doi.org/10.1145/1360612.1360697. [cited by applicant]
Rahman, et al. WERA: an observational tool develop to investigate the physical risk factor associated with WMSDs, Journal of human ergology 40 (1_2) (2011) 19-36. [cited by applicant]
Li et al. Applying the BRIEF survey in Taiwan's high-tech industries, International Journal of the Computer, The Internet and Management 11 (2) (2003) 78. [cited by applicant]
Kim, et al. “Analysis of risk factors for work-related musculoskeletal disorders in radiological technologists.” Journal of physical therapy science 26.9 (2014). [cited by applicant]
Hwang et al. “A deep learning-based method for grip strength prediction: Comparison of multilayer perceptron and polynomial regression approaches”; NIH (2021). [cited by applicant]