IP Library › Granted Patent US 12,567,173
Granted Patent B2
US 12,567,173 · App. 18/287,518 · Granted Mar 3, 2026

Infant 2D pose estimation and posture detection system

Inventors: Sarah Ostadabbas (Watertown, MA); Xiaofei Huang (Lynnfield, MA); Nihang Fu (Malden, MA); Shuangjun Liu (Boston, MA)
Assignee: Northeastern University
G06T7/74G06T7/75G06T15/04G06T2207/10016G06T2207/10024G06T2207/10028G06T2207/10048G06T2207/20081G06T2207/20084G06T2207/30196
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,567,173
App. No.
18/287,518
Filed
Oct 19, 2023
Granted
Mar 3, 2026
Kind
B2
Art Unit
2615
USPC
345/419
Abstract

Methods are provided for estimating a pose of an infant using image analysis and artificial intelligence. A classifier is trained using a dataset containing hybrid synthetic and real infant pose data. Multi-stage invariant representation machine learning strategies are employed that transfer knowledge from adjacent domains of adult poses and synthetic infant images into a fine-tuned domain-adapted infant pose estimation model.

Claims (53)

1 . A system for estimating a pose of an infant, the system comprising:

one or more processors; and

at least one memory including instructions for a pose estimator, the pose estimator comprising a trained model for estimating infant poses trained on an adult pose dataset and an augmented dataset containing real infant pose data and synthetic infant pose data using domain adaptation to align features of the synthetic infant pose data with the real infant pose data; and

wherein the at least one memory further includes instructions that, when executed by the one or more processors, cause the system to:

receive one or more images of an infant transmitted from an imaging device; and

determine a pose of the infant using the trained model.

2 . The system of claim 1 further comprising one or more imaging devices for acquiring images of the infant.

3 . The system of claim 1 , wherein the augmented dataset is provided by a process comprising:

fitting infant model pose and shape parameters into a real infant image to form a reposed model, the infant model pose and shape parameters including body joints and shape coefficients; and

generating synthetic output images as the synthetic infant pose data.

4 . The system of claim 3 , wherein the synthetic output images are generated by an imaging process as a function of the infant model pose and shape parameters, imaging device parameters, body texture maps, and background images;

wherein the imaging device parameters include principal point and focal length of a camera; and

wherein the body texture maps include infant textures from infant clothing images augmented with adult textures from adult clothing images.

5 . The system of claim 3 , wherein the synthetic output images are augmented by minimizing a cost function, wherein the cost function is a sum of loss terms weighted by imaging device parameters, the loss terms including:

a joint-based data term comprising a distance between ground truth two dimensional joints and a two dimensional projection of corresponding posed three dimensional joints of the infant model pose for each joint;

a mixture of Gaussian pose priors learned from adult poses;

a shape penalty comprising a distance between a shape prior of the infant model and the shape parameters being optimized; and

a pose prior penalizing elbows and knees.

6 . The system of claim 1 , wherein the pose estimator further comprises:

a pose estimation component comprising a feature extractor and a pose predictor; and

a domain confusion component in communication with the pose estimation component to share the feature extractor and including a domain classifier operative to distinguish whether a feature on an input image belongs to a real image or synthetic image.

7 . The system of claim 6 , wherein the pose estimation component further comprises a residual neural network as an encoder and a pose head estimator as a decoder.

8 . The system of claim 6 , wherein in a first stage, the domain confusion component is fine-tuned using real infant pose data and synthetic infant pose data, to obtain a domain classifier for predicting whether features of an input image are from a synthetic infant image or a real infant image, using an optimization of a loss function to determine a binary cross entropy loss.

9 . The system of claim 8 , wherein in the first stage, the pose estimation component is locked.

10 . The system of claim 9 , wherein in a second stage, the pose estimation component is fine-tuned using the domain classifier to extract body keypoint information independently of differences between a real domain and a synthetic domain.

11 . The system of claim 10 , wherein in the second stage, weight of real data and synthetic data is balanced by maximizing a domain loss function, wherein features representing the real domain and features representing the synthetic domain become more similar.

12 . The system of claim 6 , wherein the feature extractor and the domain classifier are operative to enforce mapping of input images in a real domain or a synthetic domain into a same feature space after feature extraction.

13 . The system of claim 6 , wherein the domain classifier comprises a binary classifier with three fully connected layers operative to distinguish whether an input feature belongs to a real image or a synthetic image.

14 . The system of claim 6 , wherein the pose estimation component is pre-trained with real adult data, and the domain confusion component is pre-trained with real adult data and synthetic adult data.

15 . The system of claim 2 , wherein the one or more imaging devices are selected from a video camera, a motion capture device, a red-green-blue (RGB) camera, a long-wavelength infrared (LWIR) imaging device, and a depth sensor.

16 . A method of estimating a pose of an infant comprising:

receiving one or more images of an infant transmitted from an imaging device; and

determining a pose or posture of the infant by processing the one or more images with a pose estimator, wherein the pose estimator includes a trained model for estimating infant poses trained on an adult pose dataset and an augmented dataset containing real infant pose data and synthetic infant pose data using domain adaptation to align features of the synthetic infant pose data with the real infant pose data.

17 . The method of claim 16 , wherein the

one or more images of the infant are captured using the imaging device.

18 . The method of claim 16 , wherein the augmented dataset is generated by:

fitting infant model pose and shape parameters into a real infant image to form a reposed model, the infant model pose and shape parameters including body joints and shape coefficients, the shape coefficients including representations of height, length, fatness, thinness, and head-to-body ratio; and

generating synthetic output images, wherein the synthetic output images are generated by an imaging process as a function of the pose and shape parameters, camera parameters, body texture maps, and background images, the camera parameters including camera principal point and focal length, and the body texture maps including infant textures from infant clothing images augmented with adult textures from adult clothing images.

19 . The method of claim 18 , wherein the synthetic output images are augmented by minimizing a cost function, wherein the cost function is a sum of loss terms weighted by camera parameters, the loss terms including:

a joint-based data term comprising a distance between ground truth two dimensional joints and a two dimensional projection of corresponding posed three dimensional joints of the infant model pose for each joint;

a mixture of Gaussian pose priors learned from adult poses;

a shape penalty comprising a distance between a shape prior of the infant model and the shape parameters being optimized; and

a pose prior penalizing elbows and knees.

20 . The method of claim 16 , wherein:

the pose estimator comprises:

a pose estimation component including a feature extractor and a pose predictor; and

a domain confusion component in communication with the pose estimation component to share the feature extractor and including a domain classifier operative to distinguish whether a feature on an input image belongs to a real image or synthetic image; and

wherein the method further comprises:

training the pose estimator in a first stage, wherein the domain confusion component is fine-tuned using real infant pose data and synthetic infant pose data, to obtain a domain classifier for predicting whether features of an input image are from a synthetic infant image or a real infant image, using an optimization of a loss function to determine a binary cross entropy loss;

training the pose estimator in a second stage, wherein the pose estimation component is fine-tuned using the domain classifier to extract body keypoint information independently of differences between a real domain and a synthetic domain; and

balancing weights of real data and synthetic data in the second stage by maximizing a domain loss function, wherein features representing the real domain and features representing the synthetic domain become more similar.

21 . The method of claim 20 , further comprising pre-training the pose estimation component with real adult data, and pre-training the domain confusion component with real adult data and synthetic adult data.

22 . The system of claim 3 , wherein the shape coefficients include representations of height, length, fatness, thinness, and head-to-body ratio.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Oct 30, 2023
From: OSTADABBAS, SARAH; HUANG, XIAOFEI; FU, NIHANG; LIU, SHUANGJUN
To: NORTHEASTERN UNIVERSITY
Reel/Frame 065380/0153 →
Continuity (2)
Provisional Application 63185435 · May 7, 2021
Related Publication 20240193809A1 · Jun 13, 2024
References Cited (68)
US 8548260B2 · Okada · 2013 [cited by examiner]
US 9311713B2 · Han · 2016 [cited by examiner]
US 10429944B2 · Hebbalaguppe · 2019 [cited by examiner]
US 10447972B2 · Patil · 2019 [cited by examiner]
US 10830584B2 · Boulkenafed · 2020 [cited by examiner]
US 10942256B2 · Achour · 2021 [cited by examiner]
US 11003949B2 · Lan · 2021 [cited by examiner]
US 11055580B2 · Amon · 2021 [cited by examiner]
US 11069150B2 · Sminchisescu · 2021 [cited by examiner]
US 11138416B2 · Lin · 2021 [cited by examiner]
US 11138422B2 · Hu · 2021 [cited by examiner]
US 11170581B1 · Marek · 2021 [cited by examiner]
US 11182620B2 · Weinzaepfel · 2021 [cited by examiner]
US 11256962B2 · Biswas · 2022 [cited by examiner]
US 11270455B2 · Zhang · 2022 [cited by examiner]
US 11295168B2 · Ye · 2022 [cited by examiner]
US 11410378B1 · Ghosh · 2022 [cited by examiner]
US 11501162B2 · Karg · 2022 [cited by examiner]
US 11501435B2 · Xu · 2022 [cited by examiner]
US 11568109B2 · Rejeb Sfar · 2023 [cited by examiner]
US 11568595B2 · Logothetis · 2023 [cited by examiner]
US 11574198B2 · Son · 2023 [cited by examiner]
US 11704804B2 · Abrol · 2023 [cited by examiner]
US 11727086B2 · Tan · 2023 [cited by examiner]
US 11763541B2 · Jie · 2023 [cited by examiner]
US 11790228B2 · Swami · 2023 [cited by examiner]
US 11842511B2 · Sato · 2023 [cited by examiner]
US 11937918B2 · Otsuki · 2024 [cited by examiner]
US 11977976B2 · Rejeb Sfar · 2024 [cited by examiner]
US 12087447B2 · Otsuki · 2024 [cited by examiner]
US 12141700B2 · Murray · 2024 [cited by examiner]
US 12142033B2 · Picon Ruiz · 2024 [cited by examiner]
US 12148201B2 · Kroenke · 2024 [cited by examiner]
US 20080180448A1 · Anguelov et al. · 2008 [cited by applicant]
US 20090232353A1 · Sundaresan et al. · 2009 [cited by applicant]
US 20190340803A1 · Comer · 2019 [cited by applicant]
US 20200058137A1 · Pujades et al. · 2020 [cited by applicant]
US 20200134382A1 · Zhuravlev · 2020 [cited by examiner]
US 20200241646A1 · Hebbalaguppe · 2020 [cited by examiner]
US 20200320345A1 · Nikolenko · 2020 [cited by examiner]
US 20210259582A1 · Matic · 2021 [cited by examiner]
US 20220157048A1 · Ting · 2022 [cited by examiner]
US 20220269948A1 · Grigorescu · 2022 [cited by examiner]
US 20230169754A1 · Irie · 2023 [cited by examiner]
WO WO2019177732A1 · 2019 [cited by examiner]
WO WO2020014286A1 · 2020 [cited by examiner]
WO WO2020103068A1 · 2020 [cited by examiner]
WO WO2020108362A1 · 2020 [cited by examiner]
WO WO2020249961A1 · 2020 [cited by examiner]
Ganin et al., “Unsupervised domain adaptation by backpropagation”, International conference on machine learning, 11 pages, PMLR, (2015). [cited by applicant]
Huang et al., “Invatiant Representation Learning for Infant Pose Estimation with Small Data”, Dec. 2, 2020 (Dec. 2, 2020), [retrieved on Jul. 21, 2022], Retrieved from the Internet: //arxiv.org/pdf/2010.06100v3.pdf pp. … [cited by applicant]
Fang et al., “RMPE: Regional Multi-Person Pose Estimation”, ICCV, (2017), pp. 1-10. [cited by applicant]
Hesse et al., “Body Pose Estimation in Depth Images for Infant Motion Analysis”, 2017 39th Annual International Conference of the IEEE Engineering in Medicine and Biology Society (EMBC), 4 pages. IEEE, (2017). [cited by applicant]
Hesse et al., “Computer Vision for Medical Infant Motion Analysis: State of the Art and RGB-D Data Set”, Proceedings of the European Conference on Computer Vision (ECCV), 18 pages, (2018). [cited by applicant]
Hesse et al., “Learning an infant body model from RGB-D data for accurate full body motion analysis”, International Conference on Medical Image Computing and Computer-Assisted Intervention, pp. 792-800. Springer, (2018). [cited by applicant]
Huang et al., “The Devil is in the Details: Delving into Unbiased Data Processing for Human Pose Estimation”, The IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), Jun. 2020, pp. 5700-5709. [cited by applicant]
Liu et al., “A Vision-Based System for In-Bed Posture Tracking”, Proceedings of the IEEE International Conference on Computer Vision Workshops, 13 pages, (2017). [cited by applicant]
Liu et al., “A Semi-Supervised Data Augmentation Approach using 3D Graphical Engines”, European Conference on Computer Vision, pp. 395-408, (2018). [cited by applicant]
Ma et al., “Pose Guided Person Image Generation”, Advances in neural information processing systems, 11 pages, (2017). [cited by applicant]
Rhodin et al., “Unsupervised Geometry-Aware Representation for 3D Human Pose Estimation”, Proceedings of the European Conference on Computer Vision (ECCV), 18 pages, (2018). [cited by applicant]
Su et al., “Render for CNN: Viewpoint Estimation in Images Using CNNs Trained with Rendered 3D Model Views”, Proceedings of the IEEE International Conference on Computer Vision, pp. 2686-2694, (2015). [cited by applicant]
Varol et al., “Learning from Synthetic Humans”, Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, pp. 109-117, (2017). [cited by applicant]
Vyas et al., “Recognition of Atypical Behavior in Autism Diagnosis From Video Using Pose Estimation Over Time”, 2019 IEEE 29th International Workshop on Machine Learning for Signal Processing (MLSP), 6 pages, IEEE, (201… [cited by applicant]
Zhang et al., “Distribution-Aware Coordinate Representation for Human Pose Estimation”, Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pp. 7093-7102, (2020). [cited by applicant]
Pavlakos et al., “Expressive Body Capture: 3D Hands, Face, and Body from a Single Image”, Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, pp. 10975-10985, (2019). [cited by applicant]
Pavllo et al., “3D human pose estimation in video with temporal convolutions and semi-supervised training”, Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, pp. 7753-7762, (2019). [cited by applicant]
Yu et al., “LSUN: Construction of a Large-Scale Image Dataset using Deep Learning with Humans in the Loop”, arXiv preprint arXiv:1506.03365, 9 pages, (2015). [cited by applicant]
Huang et al., “AH-CoLT: an AI-Human Co-Labeling Toolbox to Augment Efficient Groundtruth Generation”, IEEE 29th International Workshop on Machine Learning for Signal Processing (MLSP), 6 pages, IEEE, 2019. [cited by applicant]