IP Library › Granted Patent US 12,586,020
Granted Patent B2
US 12,586,020 · App. 18/381,988 · Granted Mar 24, 2026

Determining impacts of work items on repositories

Inventors: Thomas Kenneth Monson (Chicago, IL); Hitheshwar Peddamekala (Glen Allen, VA); Evelio Sosa (Raleigh, NC); Emily Otero (Durham, NC); Keith Gregory Frost (Delaware, OH)
Assignee: International Business Machines Corporation
G06Q10/0633G06F40/40G06F8/71
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,586,020
App. No.
18/381,988
Granted
Mar 24, 2026
Kind
B2
Abstract

A computer-implemented method, according to one approach, includes: receiving a new work item, and extracting topics from the new work item. A trained machine learning model is used to determine a first set of values representing correlation strengths between the new work item and the topics. Moreover, the first set of values are compared to a second set of values, where the second set of values represents correlation strengths between the topics and multiple files. A third set of values representing correlation strengths between the new work item and the multiple files is also generated, and output.

Claims (86)

1 . A computer-implemented method, comprising:

receiving a new work item;

extracting topics from the new work item;

using a trained machine learning model to determine a first set of values representing correlation strengths between the new work item and the topics;

causing a set of pull requests (PRs) which are configured to solve the new work item to be produced by a confidence weighting module;

comparing the first set of values to a second set of values, the second set of values representing correlation strengths between the topics and multiple files;

generating a third set of values representing correlation strengths between the new work item and the multiple files;

using the third set of values to identify one or more PRs in the set of PRs to use to satisfy the new work item, wherein the one or more PRs are identified using the confidence weighting module;

outputting the third set of values, by:

converting the third set of values into a heatmap, and

causing the heatmap to be displayed on a monitor;

in response to determining a correlation strength between the new work item and a given one of the files is in a predetermined range, using the third set of values to preload the given one of the files into cache; and

causing the new work item to be satisfied using the identified one or more PRs and the preloaded one of the files.

2 . The computer-implemented method of claim 1 , wherein the machine learning model is trained by:

receiving a set of work items;

extracting topics from the received set of work items;

generating a first training set of values that represent a correlation between (i) each of the work items in the received set, and (ii) each of the topics extracted from the received set of work items; and

comparing the first training set of values to PRs configured to solve each of the respective work items in the received set.

3 . The computer-implemented method of claim 2 , wherein training the machine learning model includes:

generating a second training set of values that represent a correlation between (i) each of the PRs configured to solve the respective work items in the received set, and (ii) each of the topics extracted from the received set of work items; and

using the second training set of values to generate the second set of values.

4 . The computer-implemented method of claim 3 , wherein using the second training set of values to generate the second set of values includes performing a file-to-topic association.

5 . The computer-implemented method of claim 2 , wherein extracting topics from the received set of work items includes using one or more trained natural language processing models.

6 . The computer-implemented method of claim 5 , wherein extracting topics from the new work item received includes using the one or more trained natural language processing models.

7 . The computer-implemented method of claim 1 , wherein generating the third set of values includes performing a topic-to-file association by:

determining a first group of correlation strengths between the multiple files and the new work item; and

determining a second group of correlation strengths between the new work item and the multiple files.

8 . The computer-implemented method of claim 1 , comprising:

using the first set of values representing correlation strengths between the new work item and the topics to dynamically retrain machine learning models that are configured to generate the third set of values representing correlation strengths between the new work item and the multiple files.

9 . A computer program product, comprising a computer readable storage medium having program instructions embodied therewith, the program instructions readable by a processor, executable by the processor, or readable and executable by the processor, to cause the processor to:

receive a new work item;

extract topics from the new work item;

use a trained machine learning model to determine a first set of values representing correlation strengths between the new work item and the topics;

cause a set of pull requests (PRs) which are configured to solve the new work item to be produced by a confidence weighting module;

compare the first set of values to a second set of values, the second set of values representing correlation strengths between the topics and multiple files;

generate a third set of values representing correlation strengths between the new work item and the multiple files;

use the third set of values to identify one or more PRs in the set of PRs to use to satisfy the new work item, wherein the one or more PRs are identified using the confidence weighting module;

output the third set of values, by:

converting the third set of values into a heatmap, and

causing the heatmap to be displayed on a monitor;

in response to determining a correlation strength between the new work item and a given one of the files is in a predetermined range, use the third set of values to preload the given one of the files into cache; and

cause the new work item to be satisfied using the identified one or more PRs and the preloaded one of the files.

10 . The computer program product of claim 9 , wherein the machine learning model is trained by:

receiving a set of work items;

extracting topics from the received set of work items;

generating a first training set of values that represent a correlation between (i) each of the work items in the received set, and (ii) each of the topics extracted from the received set of work items; and

comparing the first training set of values to PRs configured to solve each of the respective work items in the received set.

11 . The computer program product of claim 10 , wherein training the machine learning model includes:

generating a second training set of values that represent a correlation between (i) each of the PRs configured to solve the respective work items in the received set, and (ii) each of the topics extracted from the received set of work items; and

using the second training set of values to generate the second set of values.

12 . The computer program product of claim 11 , wherein using the second training set of values to generate the second set of values includes performing a file-to-topic association.

13 . The computer program product of claim 10 , wherein extracting topics from the received set of work items includes using one or more trained natural language processing models.

14 . The computer program product of claim 13 , wherein extracting topics from the new work item received includes using the one or more trained natural language processing models.

15 . The computer program product of claim 9 , wherein generating the third set of values includes:

using a file-to-topic association module to perform a file-to-topic association and determine files that are modified and/or created by the one or more PRs; and

using a topic-to-file association module to perform a topic-to-file association and generate a group of confidence scores for the multiple file.

16 . The computer program product of claim 9 , wherein the program instructions are readable and/or executable by the processor to cause the processor to:

use the first set of values representing correlation strengths between the new work item and the topics to dynamically retrain machine learning models that are configured to generate the third set of values representing correlation strengths between the new work item and the multiple files.

17 . A system, comprising:

a processor; and

logic integrated with the processor, executable by the processor, or integrated with and executable by the processor, the logic being configured to:

receive a new work item;

extract topics from the new work item;

use a trained machine learning model to determine a first set of values representing correlation strengths between the new work item and the topics;

cause a set of pull requests (PRs) which are configured to solve the new work item to be produced by a confidence weighting module;

compare the first set of values to a second set of values, the second set of values representing correlation strengths between the topics and multiple files;

generate a third set of values representing correlation strengths between the new work item and the multiple files;

use the third set of values to identify one or more PRs in the set of PRs to use to satisfy the new work item, wherein the one or more PRs are identified using the confidence weighting module;

output the third set of values, by:

converting the third set of values into a heatmap, and

causing the heatmap to be displayed on a monitor;

in response to determining a correlation strength between the new work item and a given one of the files is in a predetermined range, use the third set of values to preload the given one of the files into cache; and

cause the new work item to be satisfied using the identified one or more PRs and the preloaded one of the files.

18 . The system of claim 17 , wherein the machine learning model is trained by:

receiving a set of work items;

extracting topics from the received set of work items;

generating a first training set of values that represent a correlation between (i) each of the work items in the received set, and (ii) each of the topics extracted from the received set of work items; and

comparing the first training set of values to PRs configured to solve each of the respective work items in the received set,

wherein the logic is further configured to:

use the first set of values representing correlation strengths between the new work item and the topics to dynamically retrain machine learning models that are configured to generate the third set of values representing correlation strengths between the new work item and the multiple files;

generate a second training set of values that represent a correlation between (i) each of the PRs configured to solve the respective work items in the received set, and (ii) each of the topics extracted from the received set of work items; and

use the second training set of values to generate the second set of values,

wherein using the second training set of values to generate the second set of values includes performing a file-to-topic association,

wherein generating the third set of values includes:

using a file-to-topic association module to perform a file-to-topic association and determine files that are modified and/or created by the one or more PRs; and

using a topic-to-file association module to perform a topic-to-file association and generate a group of confidence scores for the multiple file.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Oct 24, 2023
From: MONSON, THOMAS KENNETH; PEDDAMEKALA, HITHESHWAR; SOSA, EVELIO; OTERO, EMILY; FROST, KEITH GREGORY
To: INTERNATIONAL BUSINESS MACHINES CORPORATION
Reel/Frame 065321/0560 →
Continuity (1)
Related Publication 20250131356A1 · Apr 24, 2025
References Cited (23)
US 8984485B2 · Elshishiny et al. · 2015 [cited by applicant]
US 11064074B2 · Erhart et al. · 2021 [cited by applicant]
US 11100438B2 · Somech et al. · 2021 [cited by applicant]
US 11610145B2 · Rogynskyy et al. · 2023 [cited by applicant]
US 20160357519A1 · Vargas · 2016 [cited by applicant]
US 20180114177A1 · Somech et al. · 2018 [cited by applicant]
US 20190026697A1 · Burton · 2019 [cited by examiner]
US 20190347282A1 · Cai · 2019 [cited by examiner]
US 20200089761A1 · Guerra · 2020 [cited by examiner]
US 20200387819A1 · Rogynskyy et al. · 2020 [cited by applicant]
US 20210029249A1 · Erhart et al. · 2021 [cited by applicant]
US 20220004479A1 · McCawley · 2022 [cited by examiner]
Catolino, et at., “Not all bugs are the same: Understanding, characterizing, and classifying bug types,” 2019, The Journal of Systems and Software, vol. 152, pp. 165-181 (Year: 2019). [cited by examiner]
Jeswani, et al, “Minimizing latency in serving requests through differential template caching in a cloud,” 2012, In 2012 IEEE Fifth International Conference on Cloud Computing, pp. 269-276 (Year: 2012). [cited by examiner]
Panichella et al., “How to Effectively Use Topic Models for Software Engineering Tasks? An Approach Based on Genetic Algorithms,” 35th International Conference on Software Engineering, May 2013, 10 pages, retrieved from… [cited by applicant]
Dit et al., “Configuring Topic Models for Software Engineering Tasks in TraceLab,” IEEE International Workshop on Traceability in Emerging Forms of Software Engineering (TEFSE), Jun. 2013, pp. 105-109. [cited by applicant]
Licorish et al., “Exploring software developers' work practices: Task differences, participation, engagement, and speed of task resolution,” Information & Management, vol. 54, No. 3, 2017, 21 pages. [cited by applicant]
Liang et al., “TaskMatrix.AI: Completing Tasks by Connecting Foundation Models with Millions of APIs,” arXiv, Mar. 2023, 27 pages, retrieved from https://arxiv.org/abs/2303.16434. [cited by applicant]
Stanley et al., “Distributed Ensemble Learning for Provisioning Indices, Actions an Inputs for Assurance and Performance,” IP.com Prior Art Database, Technical Disclosure No. IPCOM000252743D, Feb. 6, 2018, 12 pages. [cited by applicant]
Anonymous, “System and Method to Use a Directed Graph and Artificial Intelligence to Identify a Public Cloud for Set of Objectives and Constraints,” IP.com Prior Art Database, Technical Disclosure No. IPCOM000272162D, A… [cited by applicant]
Github, “Your AI pair programmer,” GitHub, 2022, 12 pages, retrieved from https://github.com/features/copilot. [cited by applicant]
Codescene, “4 key factors behind high-performing software development,” CodeScene, 2023, 15 pages, retrieved from https://codescene.com/. [cited by applicant]
Sonar, “clean code for teams and enterprises with {SonarQube},” Sonar, 2023, 11 pages, retrieved from https://www.sonarsource.com/products/sonarqube/. [cited by applicant]