IP Library › Granted Patent US 11,726,775
Granted Patent B2
US 11,726,775 · App. 17/349,154 · Granted Aug 15, 2023

Source code issue assignment using machine learning

Inventors: Prabal Mahanta (Bangalore, IN); Vipul Khullar (New Delhi, IN)
Assignee: SAP SE
G06F8/71G06F8/24G06F11/362G06F11/3604G06F18/2113G06N20/00G06V10/751
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,726,775
App. No.
17/349,154
Granted
Aug 15, 2023
Kind
B2
Abstract

Technologies are provided for assigning developers to source code issues using machine learning. A machine learning model can be generated based on multiple versions of source code objects (such as source code files, classes, modules, packages, etc.), such as those that are managed by a version control system. The versions of the source code objects can reflect changes that are made to the source code objects over time. Associations between developers and source code object versions can be analyzed and used to train the machine learning model. Patterns of similar changes to various source code objects can be detected and can also be used to train the machine learning model. When an issue is detected in a version of a source code object, the model can be used to identify a developer to assign to the issue. Feedback data regarding the developer assignment can be used to re-train the model.

Claims (50)

1. A method, comprising:

generating a machine learning model using multiple versions of a source code object and identifiers of a plurality of developers associated with the multiple versions of the source code object, wherein the machine learning model is generated by generating feature sets for the multiple versions of the source code object, comparing the feature sets to one another, generating similarity scores for the multiple versions of the source code object based on the comparing, and incorporating the similarity scores into training of the machine learning model;

receiving an additional version of the source code object;

detecting a source code issue in the additional version of the source code object; and

using the machine learning model to identify a developer, of the plurality of developers, as a candidate to correct the source code issue in the additional version of the source code object.

2. The method of claim 1 , wherein the generating the machine learning model comprises using multiple versions of a plurality of source code objects, including the source code object.

3. The method of claim 2 , wherein the generating the machine learning model comprises:

comparing multiple source code objects, of the plurality of source code objects, that are associated with a same developer identifier; and

determining a coding signature for a developer associated with the developer identifier.

4. The method of claim 2 , wherein the generating the machine learning model comprises:

generating feature sets for the multiple versions of the plurality of source code objects;

comparing the feature sets to one another; and

generating similarity scores for the multiple versions of the plurality of source code objects based on the comparing.

5. The method of claim 1 , wherein the generating the machine learning model comprises:

identifying one or more source code objects related to the source code object; and

analyzing change histories for the multiple versions of the source code object and the one or more source code objects.

6. The method of claim 5 , wherein identifying the one or more source code objects that are related to the source code object comprises determining that the one or more source code objects are in a hierarchical relationship with the source code object.

7. The method of claim 1 , wherein the source code issue comprises a static code issue detected during a static source code analysis of the additional version of the source code object.

8. A system, comprising:

a computing device comprising a processor and a memory storing instructions that, when executed by the processor, cause the computing device to perform operations, the operations comprising:

storing a machine learning model generated using multiple versions of a source code object and identifiers of a plurality of developers associated with the multiple versions of the source code object, wherein the machine learning model is generated by generating feature sets for the multiple versions of the source code object, comparing the feature sets to one another, generating similarity scores for the multiple versions of the source code object based on the comparing, and incorporating the similarity scores into training of the machine learning model;

receiving an additional version of the source code object;

detecting a source code issue in the additional version of the source code object; and

using the machine learning model to identify a developer, of the plurality of developers, as a candidate to correct the source code issue in the additional version of the source code object.

9. The system of claim 8 , wherein the generating the machine learning model comprises:

identifying one or more source code objects related to the source code object; and

analyzing the multiple versions of the source code object and multiple versions of the one or more source code objects related to the source code object.

10. The system of claim 9 , wherein the identifying the one or more source code objects related to the source code object comprises determining that the one or more source code objects are in a hierarchical relationship with the source code object.

11. The system of claim 8 , wherein the generating the features sets for the multiple versions of the source code object comprises generating abstract syntax trees for the multiple versions of the source code object and creating the feature sets using the abstract syntax trees.

12. The system of claim 8 , wherein the generating the machine learning model comprises:

identifying versions of the source code object, of the multiple versions of the source code object, that are associated with a same developer identifier; and

determining a coding signature for a developer associated with the developer identifier.

13. One or more computer-readable storage media storing instructions that, when executed by one or more processors, cause the one or more processors to perform operations, the operations comprising:

generating a machine learning model using multiple versions of a source code object and identifiers of a plurality of developers associated with the multiple versions of the source code object, wherein the machine learning model is generated by generating feature sets for the multiple versions of the source code object, comparing the feature sets to one another, generating similarity scores for the multiple versions of the source code object based on the comparing, and incorporating the similarity scores into training of the machine learning model;

receiving an additional version of the source code object;

detecting a source code issue in the additional version of the source code object, wherein the source code issue comprises an error detected by static source code analysis or an error or exception detected by dynamic analysis of an application, library, or module containing a compiled representation of the additional version of the source code object; and

using the machine learning model to identify a developer, of the plurality of developers, as a candidate to correct the source code issue in the additional version of the source code object.

14. The one or more computer-readable storage media of claim 13 , wherein the generating the machine learning model comprises using multiple versions of a plurality of source code objects, including the source code object.

15. The one or more computer-readable storage media of claim 14 , wherein the generating the machine learning model further comprises:

comparing multiple source code objects, of the plurality of source code objects, that are associated with a same developer identifier; and

determining a coding signature for a developer associated with the developer identifier.

16. The one or more computer-readable storage media of claim 14 , wherein the generating the machine learning model comprises:

generating feature sets for the multiple versions of the plurality of source code objects;

comparing the feature sets to one another; and

generating similarity scores for the multiple versions of the plurality of source code objects based on the comparing.

17. The one or more computer-readable storage media of claim 16 , wherein the generating the features sets for the multiple versions of the plurality of source code objects comprises generating abstract syntax trees for the plurality of source code objects and creating the feature sets using the abstract syntax trees.

18. The one or more computer-readable storage media of claim 13 , wherein the generating the machine learning model comprises:

identifying one or more source code objects related to the source code object; and

analyzing change histories for the multiple versions of the source code object and the one or more source code objects related to the source code object.

19. The one or more computer-readable storage media of claim 18 , wherein identifying the one or more source code objects related to the source code object comprises determining that the one or more source code objects are in a hierarchical relationship with the source code object.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jul 6, 2021
From: MAHANTA, PRABAL; KHULLAR, VIPUL
To: SAP SE
Reel/Frame 056767/0196 →
Continuity (1)
Related Publication 20220405091A1 · Dec 22, 2022
Cited By (2)
US 12,229,552 US 12,645,561