IP Library › Granted Patent US 12,586,361
Granted Patent B2
US 12,586,361 · App. 18/364,728 · Granted Mar 24, 2026

Online adaptation of segmentation machine learning systems

Inventors: Kambiz Azarian Yazdi (San Diego, CA); Debasmit Das (San Diego, CA); Hyojin Park (San Diego, CA); Fatih Murat Porikli (San Diego, CA)
Assignee: QUALCOMM Incorporated
G06V10/778G06N3/0895G06V10/267G06V10/82
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,586,361
App. No.
18/364,728
Granted
Mar 24, 2026
Kind
B2
Abstract

Techniques and systems are provided for performing online adaptation of machine learning model(s). For example, a process may include obtaining features extracted from a image by a machine learning model during inference and determining, by the machine learning model based on the features during inference, a plurality of keypoint estimates in the image and/or a bounding region estimate associated with an object in the image. The process may further include generating pseudo-label(s) based on the plurality of keypoint estimates and/or the bounding region estimate. The process may include determining at least one self-supervised loss based on the plurality of keypoint estimates and/or the bounding region estimate. The process may further include adapting, based on the at least one self-supervised loss, parameter(s) of the machine learning model. The process may include generating, using the machine learning model with the adapted parameter(s), a segmentation mask for the image (or another image).

Claims (70)

1 . A processor-implemented method, comprising:

obtaining features extracted from a first image by a machine learning model during inference;

determining, by the machine learning model based on the features during inference, at least one of a plurality of keypoint estimates in the first image or a bounding region estimate associated with an object in the first image;

generating one or more pseudo-labels based on at least one of the plurality of keypoint estimates or the bounding region estimate;

determining at least one self-supervised loss based on at least one of the plurality of keypoint estimates or the bounding region estimate;

adapting, based on the at least one self-supervised loss, one or more parameters of the machine learning model; and

generating, using the machine learning model with the adapted one or more parameters, a segmentation mask for the first image or a second image;

wherein the features are extracted from the first image by a feature extraction engine of the machine learning model; and

wherein adapting, based on the at least one self-supervised loss, the one or more parameters of the machine learning model comprises adapting one or more parameters of the feature extraction engine.

2 . The method of claim 1 , further comprising:

processing, using the machine learning model with the adapted one or more parameters, the first image or a second image to generate updated features representing the first image or to generate features representing the second image.

3 . The method of claim 2 , wherein the segmentation mask is generated based on processing the updated features representing the first image or the features representing the second image.

4 . The method of claim 1 , wherein the plurality of keypoint estimates is determined in the first image by a keypoint detection engine of the machine learning model.

5 . The method of claim 1 , wherein the bounding region estimate is determined in the first image by an object detection engine of the machine learning model.

6 . The method of claim 1 , wherein the segmentation mask is generated by a segmentation mask engine of the machine learning model.

7 . The method of claim 1 , wherein the plurality of keypoint estimates includes at least one heat map indicating probabilities that portions of the first image correspond to a particular keypoint.

8 . The method of claim 7 , wherein the at least one heat map includes a plurality of heat maps, each heat map of the plurality of heat maps being associated with a respective keypoint of a plurality of keypoints.

9 . The method of claim 1 , wherein the plurality of keypoint estimates include coordinates associated with a plurality of keypoints.

10 . The method of claim 1 , wherein generating the one or more pseudo-labels comprises:

determining one or more keypoint estimates from the plurality of keypoint estimates that are greater than an individual keypoint score threshold; and

selecting the one or more keypoint estimates as the one or more pseudo-labels based on the one or more keypoint estimates being greater than the individual keypoint score threshold.

11 . The method of claim 1 , wherein generating the one or more pseudo-labels comprises:

determining the plurality of keypoint estimates are greater than a total score threshold; and

selecting the plurality of keypoint estimates as the one or more pseudo-labels based on the plurality of keypoint estimates being greater than the total score threshold.

12 . The method of claim 1 , wherein the one or more parameters include weights of the machine learning model.

13 . An apparatus for performing online adaptation of one or more machine learning models, the apparatus comprising:

at least one memory; and

at least one processor coupled to the at least one memory and configured to:

obtain features extracted from a first image by a machine learning model during inference;

determine, using the machine learning model based on the features during inference, at least one of a plurality of keypoint estimates in the first image or a bounding region estimate associated with an object in the first image;

generate one or more pseudo-labels based on at least one of the plurality of keypoint estimates or the bounding region estimate;

determine at least one self-supervised loss based on at least one of the plurality of keypoint estimates or the bounding region estimate;

adapt, based on the at least one self-supervised loss, one or more parameters of the machine learning model; and

generate, using the machine learning model with the adapted one or more parameters, a segmentation mask for the first image or a second image;

wherein the features are extracted from the first image by a feature extraction engine of the machine learning model; and

wherein adapting, based on the at least one self-supervised loss, the one or more parameters of the machine learning model comprises adapting one or more parameters of the feature extraction engine.

14 . The apparatus of claim 13 , wherein the at least one processor is configured to:

process, using the machine learning model with the adapted one or more parameters, the first image or a second image to generate updated features representing the first image or to generate features representing the second image.

15 . The apparatus of claim 14 , wherein the at least one processor is configured to generate the segmentation mask based on processing the updated features representing the first image or the features representing the second image.

16 . The apparatus of claim 13 , wherein the at least one processor is configured to determine the plurality of keypoint estimates in the first image using a keypoint detection engine of the machine learning model.

17 . The apparatus of claim 13 , wherein the at least one processor is configured to determine the bounding region estimate in the first image using an object detection engine of the machine learning model.

18 . The apparatus of claim 13 , wherein the at least one processor is configured to generate the segmentation mask using a segmentation mask engine of the machine learning model.

19 . The apparatus of claim 13 , wherein the plurality of keypoint estimates includes at least one heat map indicating probabilities that portions of the first image correspond to a particular keypoint.

20 . The apparatus of claim 19 , wherein the at least one heat map includes a plurality of heat maps, each heat map of the plurality of heat maps being associated with a respective keypoint of a plurality of keypoints.

21 . The apparatus of claim 13 , wherein the plurality of keypoint estimates includes coordinates associated with a plurality of keypoints.

22 . The apparatus of claim 13 , wherein, to generate the one or more pseudo-labels, the at least one processor is configured to:

determine one or more keypoint estimates from the plurality of keypoint estimates that are greater than an individual keypoint score threshold; and

select the one or more keypoint estimates as the one or more pseudo-labels based on the one or more keypoint estimates being greater than the individual keypoint score threshold.

23 . The apparatus of claim 13 , wherein, to generate the one or more pseudo-labels, the at least one processor is configured to:

determine the plurality of keypoint estimates are greater than a total score threshold; and

select the plurality of keypoint estimates as the one or more pseudo-labels based on the plurality of keypoint estimates being greater than the total score threshold.

24 . The apparatus of claim 13 , wherein the one or more parameters include weights of the machine learning model.

25 . A non-transitory computer-readable medium having stored thereon instructions that, when executed by at least one processor, cause the at least one processor to:

obtain features extracted from a first image by a machine learning model during inference;

determine, using the machine learning model based on the features during inference, at least one of a plurality of keypoint estimates in the first image or a bounding region estimate associated with an object in the first image;

generate one or more pseudo-labels based on at least one of the plurality of keypoint estimates or the bounding region estimate;

determine at least one self-supervised loss based on at least one of the plurality of keypoint estimates or the bounding region estimate;

adapt, based on the at least one self-supervised loss, one or more parameters of the machine learning model; and

generate, using the machine learning model with the adapted one or more parameters, a segmentation mask for the first image or a second image;

wherein the features are extracted from the first image by a feature extraction engine of the machine learning model; and

wherein adapting, based on the at least one self-supervised loss, the one or more parameters of the machine learning model comprises adapting one or more parameters of the feature extraction engine.

26 . An apparatus for performing online adaptation of one or more machine learning models, the apparatus comprising:

means for obtaining features extracted from a first image by a machine learning model during inference;

means for determining, by the machine learning model based on the features during inference, at least one of a plurality of keypoint estimates in the first image or a bounding region estimate associated with an object in the first image;

means for generating one or more pseudo-labels based on at least one of the plurality of keypoint estimates or the bounding region estimate;

means for determining at least one self-supervised loss based on at least one of the plurality of keypoint estimates or the bounding region estimate;

means for adapting, based on the at least one self-supervised loss, one or more parameters of the machine learning model; and

means for generating, using the machine learning model with the adapted one or more parameters, a segmentation mask for the first image or a second image;

wherein the features are extracted from the first image by a feature extraction engine of the machine learning model; and

wherein adapting, based on the at least one self-supervised loss, the one or more parameters of the machine learning model comprises adapting one or more parameters of the feature extraction engine.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Sep 13, 2023
From: AZARIAN YAZDI, KAMBIZ; DAS, DEBASMIT; PARK, HYOJIN; PORIKLI, FATIH MURAT
To: QUALCOMM INCORPORATED
Reel/Frame 064895/0616 →
Continuity (2)
Provisional Application 63401611 · Aug 27, 2022
Related Publication 20240078797A1 · Mar 7, 2024
References Cited (22)
US 11410315B2 · Homayounfar et al. · 2022 [cited by applicant]
US 11449788B2 · Perona · 2022 [cited by examiner]
US 12159410B2 · Ye · 2024 [cited by examiner]
US 20210279950A1 · Phalak · 2021 [cited by examiner]
US 20220066544A1 · Kwon · 2022 [cited by examiner]
US 20220067940A1 · Heng · 2022 [cited by examiner]
US 20220261593A1 · Yu · 2022 [cited by examiner]
US 20220383505A1 · Homayounfar · 2022 [cited by examiner]
US 20230071437A1 · Kim · 2023 [cited by examiner]
US 20230196755A1 · He · 2023 [cited by examiner]
US 20240010225A1 · Huang · 2024 [cited by examiner]
CN 110503097A · 2019 [cited by applicant]
EP 4091093B1 · 2025 [cited by examiner]
JP 2019125340A · 2019 [cited by examiner]
JP 2023132818A · 2023 [cited by examiner]
Liangzhi Li et al. , “Article Self-Supervised Keypoint Detection and Cross-Fusion Matching Networks for Multimodal Remote Sensing Image Registration,” Jul. 27, 2022,Remote Sens. 2022, 14, 3599,pp. 1-14. [cited by examiner]
Tahir Rizwan et al. , “Neural Network Approach for 2-Dimension Person Pose Estimation With Encoded Mask and Keypoint Detection,” Jun. 10, 2020, IEEEAcess,vol. 8, pp. 1-10. [cited by examiner]
Kai Fang et al.,“A Multitarget Interested Region Extraction Method for Wrist X-Ray Images Based on Optimized AlexNet and Two-Class Combined Model,” Dec. 15, 2021, IEEE Transactions on Computational Social Systems, vol. … [cited by examiner]
Yunjie Xu et al.,“Dual attention-based method for occluded person re-identification,” Nov. 11, 2020, Knowledge-Based Systems 212 (2021),pp. 1-11. [cited by examiner]
Bernhard Mueller,“Machine Learning—based Image Segmentation,” Apr. 2, 2018, Technical University of Munich, Germany Chair of Computational Modeling and Simulation,pp. 20-32. [cited by examiner]
Johann Sawatzky et al.,“Harvesting Information from Captions forWeakly Supervised Semantic Segmentation,” Oct. 2019, Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV), 2019,pp. 1-7. [cited by examiner]
International Search Report and Written Opinion—PCT/US2023/071691—ISA/EPO—Nov. 15, 2023. [cited by applicant]