IP Library › Granted Patent US 12,493,462
Granted Patent B2
US 12,493,462 · App. 18/357,375 · Granted Dec 9, 2025

Systems and methods for detecting usage of undermaintained open-source packages and versions

Inventors: Anne Jackson (Waban, MA); Nicolaus Christiaan Thirion (Scottsdale, AZ); Ziqian Huang (Buffalo Grove, IL); Jack R. Freier (Saint Paul, MN)
Assignee: Optum, Inc.
G06F8/71G06F8/77
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,493,462
App. No.
18/357,375
Granted
Dec 9, 2025
Kind
B2
Abstract

Various embodiments of the present disclosure provide methods, apparatus, systems, computing devices, computing entities, and/or the like for identifying stale or vibrant open-source packages by training a predictive machine learning model with a labeled dataset, wherein the labeled dataset is created by generating package-basis features and version-basis features based on repository data, generating package-basis clusters based on the package-basis features, generating version-basis clusters based on the version-basis features, and generating labels for the labeled dataset based on the package-basis clusters and the version-basis clusters.

Claims (53)

1 . A computer-implemented method comprising:

receiving, by one or more processors, a prediction input comprising an identification of one or more of an open-source package or an open source version;

generating, by the one or more processors and using a predictive machine learning model, a prediction output based on the prediction input, the prediction output comprising a stale rating or a vibrancy rating for the one or more of the open-source package or the open source version associated with the prediction input, wherein the predictive machine learning model is trained based on a labeled dataset generated by:

(a) determining a package-basis feature associated with a plurality of training open-source packages based on repository data,

(b) generating, using a package-basis clustering machine learning model, a plurality of package-basis clusters based on the package-basis feature, wherein the plurality of package-basis clusters comprises a package-basis cluster data object associated with the plurality of training open-source packages,

(c) generating a first assignment of a package-basis cluster label of a plurality of package-basis cluster labels to a package-basis cluster of the plurality of package-basis clusters,

(d) determining a version-basis feature associated with a plurality of training open-source versions based on the repository data,

(e) generating, using a version-basis clustering machine learning model, a plurality of version-basis clusters based on the version-basis feature, wherein the plurality of version-basis clusters comprises a version-basis cluster data object associated with the plurality of training open-source versions,

(f) generating a second assignment of a version-basis cluster label of a plurality of version-basis cluster labels to a version-basis cluster of the plurality of version-basis clusters,

(g) labeling a plurality of training data objects associated with at least a portion of the plurality of training open-source packages and the plurality of training open-source versions based on the first assignment and the second assignment;

in response to a determination of the stale rating for the one or more of the open-source package or the open source version based at least in part on the prediction output, determining, by the one or more processors, a computing system comprising the one or more of the open-source package or the open source version; and

performing, by the one or more processors and based on the prediction output, an update to the one or more of the open-source package or the open source version on the computing system to comply with an operational or coding practice.

2 . The computer-implemented method of claim 1 , wherein at least one of the package-basis clustering machine learning model or the version-basis clustering machine learning model comprises a Gaussian mixture machine learning model.

3 . The computer-implemented method of claim 1 , wherein the predictive machine learning model comprises a semi-supervised machine learning model.

4 . The computer-implemented method of claim 3 , wherein the predictive machine learning model is trained based on a combination of a machine-labeled dataset and a developer-labeled dataset.

5 . The computer-implemented method of claim 1 , wherein labeling the plurality of training data objects further comprises labeling the plurality of training data objects with one or more of vibrant, stale, or unlabeled labels.

6 . The computer-implemented method of claim 1 , wherein the repository data comprises package release metadata and external repository metrics.

7 . The computer-implemented method of claim 1 , wherein the prediction output comprises at least one of a probability of staleness or vibrancy associated with the one or more of the open-source package or the open source version.

8 . A system comprising one or more processors and at least one memory storing processor-executable instructions that, when executed by any of the one or more processors, cause the one or more processors to perform operations comprising:

receiving a prediction input comprising an identification of one or more of an open-source package or an open source version;

generating, using a predictive machine learning model, a prediction output based on the prediction input, the prediction output comprising a stale rating or a vibrancy rating for the one or more of the open-source package or the open source version associated with the prediction input, wherein the predictive machine learning model is trained based on a labeled dataset generated by:

(a) determining a package-basis feature associated with a plurality of training open-source packages based on repository data,

(b) generating, using a package-basis clustering machine learning model, a plurality of package-basis clusters based on the package-basis feature, wherein the plurality of package-basis clusters comprises a package-basis cluster data object associated with the plurality of training open-source packages,

(c) generating a first assignment of a package-basis cluster label of a plurality of package-basis cluster labels to a package-basis cluster of the plurality of package-basis clusters,

(d) determining a version-basis feature associated with a plurality of training open-source versions based on the repository data,

(e) generating, using a version-basis clustering machine learning model, a plurality of version-basis clusters based on the version-basis feature, wherein the plurality of version-basis clusters comprises a version-basis cluster data object associated with the plurality of training open-source versions,

(f) generating a second assignment of a version-basis cluster label of a plurality of version-basis cluster labels to a version-basis cluster of the plurality of version-basis clusters,

(g) labeling a plurality of training data objects associated with at least a portion of the plurality of training open-source packages and the plurality of training open-source versions based on the first assignment and the second assignment;

in response to a determination of the stale rating for the one or more of the open-source package or the open source version based at least in part on the prediction output, determining a computing system comprising the one or more of the open-source package or the open source version; and

performing, based on the prediction output, an update to the one or more of the open-source package or the open source version on the computing system to comply with an operational or coding practice.

9 . The system of claim 8 , wherein at least one of the package-basis clustering machine learning model or the version-basis clustering machine learning model comprises a Gaussian mixture machine learning model.

10 . The system of claim 8 , wherein the predictive machine learning model comprises a semi-supervised machine learning model.

11 . The system of claim 10 , wherein the predictive machine learning model is trained based on a combination of a machine-labeled dataset and a developer-labeled dataset.

12 . The system of claim 8 , wherein labeling the plurality of training data objects further comprises labeling the plurality of training data objects with one or more of vibrant, stale, or unlabeled labels.

13 . The system of claim 8 , wherein the repository data comprises package release metadata and external repository metrics.

14 . The system of claim 8 , wherein the prediction output comprises at least one of a probability of staleness or vibrancy associated with the one or more of the open-source package or the open-source version.

15 . One or more non-transitory computer-readable storage media storing instructions that, when executed by one or more processors, cause the one or more processors to perform operations comprising:

receiving a prediction input comprising an identification of one or more of an open-source package or an open source version;

generating, using a predictive machine learning model, a prediction output based on the prediction input, the prediction output comprising a stale rating or a vibrancy rating for the one or more of the open-source package or the open-source version associated with the prediction input, wherein the predictive machine learning model is trained based on a labeled dataset generated by:

(a) determining a package-basis feature associated with a plurality of training open-source packages based on repository data,

(b) generating, using a package-basis clustering machine learning model, a plurality of package-basis clusters based on the package-basis feature, wherein the plurality of package-basis clusters comprises a package-basis cluster data object associated with the plurality of training open-source packages,

(c) generating a first assignment of a package-basis cluster label of a plurality of package-basis cluster labels to a package-basis cluster of the plurality of package-basis clusters,

(d) determining a version-basis feature associated with a plurality of training open-source versions based on the repository data,

(e) generating, using a version-basis clustering machine learning model, a plurality of version-basis clusters based on the version-basis feature, wherein the plurality of version-basis clusters comprises a version-basis cluster data object associated with the plurality of training open-source versions,

(f) generating a second assignment of a version-basis cluster label of a plurality of version-basis cluster labels to a version-basis cluster of the plurality of version-basis clusters,

(g) labeling a plurality of training data objects associated with at least a portion of the plurality of training open-source packages and the plurality of training open-source versions based on the first assignment and the second assignment;

in response to a determination of the stale rating for the one or more of the open-source package or the open source version based at least in part on the prediction output, determining a computing system comprising the one or more of the open-source package or the open source version; and

performing, based on the prediction output, an update to the one or more of the open-source package or the open source version on the computing system to comply with an operational or coding practice.

16 . The one or more non-transitory computer-readable storage media of claim 15 , wherein at least one of the package-basis clustering machine learning model or the version-basis clustering machine learning model comprises a Gaussian mixture machine learning model.

17 . The one or more non-transitory computer-readable storage media of claim 15 , wherein the predictive machine learning model comprises a semi-supervised machine learning model.

18 . The one or more non-transitory computer-readable storage media of claim 17 , wherein the predictive machine learning model is trained based on a combination of a machine-labeled dataset and a developer-labeled dataset.

19 . The one or more non-transitory computer-readable storage media of claim 17 , wherein labeling the plurality of training data objects further comprises labeling the plurality of training data objects with one or more of vibrant, stale, or unlabeled labels.

20 . The one or more non-transitory computer-readable storage media of claim 17 , wherein the repository data comprises package release metadata and external repository metrics.

Assignments (2)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Oct 16, 2024
From: HANNAFORD, MARTIN
To: MOJO TECHNOLOGIES SAS
Reel/Frame 068918/0123 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jul 24, 2023
From: JACKSON, ANNE; THIRION, NICOLAUS CHRISTIAAN; HUANG, ZIQIAN; FREIER, JACK R.
To: OPTUM, INC.
Reel/Frame 064357/0045 →
Continuity (1)
Related Publication 20250036400A1 · Jan 30, 2025
References Cited (19)
US 10268684B1 · Denkowski · 2019 [cited by examiner]
US 10528741B1 · Collins et al. · 2020 [cited by applicant]
US 10579509B2 · Banuelos · 2020 [cited by examiner]
US 11023356B2 · Anders et al. · 2021 [cited by applicant]
US 11281998B2 · Ben-Arie · 2022 [cited by examiner]
US 11586436B1 · Jennings · 2023 [cited by examiner]
US 12197912B2 · Balasubramanian · 2025 [cited by examiner]
US 12323449B1 · Graves · 2025 [cited by examiner]
US 20140075336A1 · Curtis · 2014 [cited by examiner]
US 20200097608A1 · Xiu · 2020 [cited by examiner]
US 20200226490A1 · Abdulaal · 2020 [cited by examiner]
US 20210374563A1 · Jezewski · 2021 [cited by examiner]
US 20220114490A1 · Das · 2022 [cited by examiner]
US 20220261597A1 · Panwar · 2022 [cited by examiner]
US 20230125150A1 · Cutler · 2023 [cited by examiner]
US 20230196184A1 · Dong · 2023 [cited by examiner]
US 20230222386A1 · Li · 2023 [cited by examiner]
US 20230297886A1 · Glaser · 2023 [cited by examiner]
US 20240160997A1 · Hertzberg · 2024 [cited by examiner]