IP Library › Granted Patent US 12,159,127
Granted Patent B2
US 12,159,127 · App. 18/064,620 · Granted Dec 3, 2024

Systems and methods for detecting code duplication in codebases

Inventors: Rohan Saphal (Glasgow, GB); Fanny Silavong (London, GB); Sean Moran (London, GB); Antonios Georgiadis (London, GB); Sanat Saha (Mumbai, IN); Gaurav Singh (Mumbai, IN); Pierre Osselin (Le Chetelet-en-Brie, FR); Rob Otter (Witham, GB)
Assignee: JPMORGAN CHASE BANK, N.A.
G06F8/4435G06F8/427
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,159,127
App. No.
18/064,620
Filed
Dec 12, 2022
Granted
Dec 3, 2024
Kind
B2
Examiner
LEE, MARINA
Art Unit
2192
USPC
717/144
Abstract

Systems and methods for detecting code duplication are disclosed. In one embodiment, a method for detecting exact code snippet duplicates may include: (1) representing, by a code duplication detection computer program, each of a plurality of code snippets in a codebase as an abstract syntax trees; (2) featurizing, by the code duplication detection computer program, the abstract syntax trees into corpus feature vectors by converting the abstract syntax tree into vector representations; (3) generating, by the code duplication detection computer program, dense feature vectors from the corpus feature vectors using a dimension reduction technique; (4) identifying, by the code duplication detection computer program, exact duplicate code snippet matches by apply density-based clustering to the dense feature vectors; and (5) tagging, by the code duplication detection computer program, the exact duplicate code snippets.

Claims (50)

1. A method for detecting exact code snippet duplicates, comprising:

representing, by a code duplication detection computer program, each of a plurality of code snippets in a codebase as abstract syntax trees;

applying, by the code duplication detection computer program, a de-noising filter to the plurality of code snippets or the abstract syntax trees;

featurizing, by the code duplication detection computer program, the abstract syntax trees into corpus feature vectors by converting the abstract syntax tree into vector representations;

generating, by the code duplication detection computer program, dense feature vectors from the corpus feature vectors using a dimension reduction technique;

clustering, by the code duplication detection computer program, the dense feature vectors into dendrograms, each dendrogram having a different value for a cluster distance metric;

applying, by the code duplication detection computer program, cross-correlation thresholding to identify an optimal value for the cluster distance metric;

applying, by the code duplication detection computer program, iterative density-based clustering to the dendrogram for the optimal value for the cluster distance metric;

identifying, by the code duplication detection computer program, exact duplicate code snippet matches by apply density-based clustering to the dense feature vectors; and

tagging, by the code duplication detection computer program, the exact duplicate code snippets.

2. The method of claim 1 , further comprising:

applying, by the code duplication detection computer program, Natural Language Processing (NLP) to generate features for the abstract syntax trees.

3. The method of claim 1 , wherein the de-noising filter filters code snippets or abstract syntax trees that are not actively used.

4. The method of claim 1 , wherein the de-noising filter filters code snippets or abstract syntax trees that are irrelevant.

5. The method of claim 1 , wherein the de-noising filter is based on a trained neural network.

6. The method of claim 1 , wherein the corpus feature vectors comprise a list of featurized abstract syntax trees from a code corpus.

7. The method of claim 1 , wherein the dimension reduction technique comprises truncated Singular Value Decomposition.

8. A method for detecting near code snippet duplicates, comprising:

representing, by a code duplication detection computer program, each of a plurality of code snippets in a codebase as abstract syntax trees;

applying, by the code duplication detection computer program, a de-noising filter to the plurality of code snippets or the abstract syntax trees;

featurizing, by the code duplication detection computer program, the abstract syntax trees into corpus feature vectors by converting the abstract syntax trees into vector representations;

generating, by the code duplication detection computer program, dense feature vectors from the corpus feature vectors using a dimension reduction technique;

clustering, by the code duplication detection computer program, the dense feature vectors into dendrograms, each dendrogram having a different value for a cluster distance metric;

applying, by the code duplication detection computer program, cross-correlation thresholding to identify an optimal value for the cluster distance metric;

applying, by the code duplication detection computer program, iterative density-based clustering to the dendrogram for the optimal value for the cluster distance metric;

tracking, by the code duplication detection computer program, data points in the dendrogram that have merged into a large cluster but were also present in small unique clusters, wherein the data points belonging to the unique small cluster to identify code snippets that are near duplicates of each other; and

tagging, by the code duplication detection computer program, the near duplicate code snippets.

9. The method of claim 8 , further comprising:

applying, by the code duplication detection computer program, Natural Language Processing (NLP) to generate features for the abstract syntax trees.

10. The method of claim 8 , wherein the de-noising filter filters code snippets or abstract syntax trees that are not actively used.

11. The method of claim 8 , wherein the de-noising filter is based on a trained neural network.

12. The method of claim 8 , wherein the corpus feature vectors comprise a list of featurized abstract syntax trees from a code corpus.

13. The method of claim 8 , wherein the dimension reduction technique comprises truncated Singular Value Decomposition.

14. A method for detecting exact code snippet duplicates, comprising:

loading, by a code duplication detection computer program, a near duplicate centroid, an exact duplicate centroid, a vectorizer, and a dimension reduction model;

producing, by the code duplication detection computer program, dense vectors from incremental functions;

representing, by a code duplication detection computer program, each of a plurality of incremental functions as abstract syntax trees;

featurizing, by the code duplication detection computer program, the abstract syntax trees into incremental function feature vectors by converting the abstract syntax trees into vector representations;

generating, by the code duplication detection computer program, dense feature vectors from the incremental function feature vectors using the dimension reduction model;

computing, by the code duplication detection computer program, a cosine similarity between the generated dense vectors and the near duplicate centroid and the exact duplicate centroid;

ranking, by the code duplication detection computer program, the near duplicate centroid and the exact duplicate centroid in descending order based on the similarity;

thresholding, by the code duplication detection computer program, the ranked near duplicate centroid and the exact duplicate centroid;

selecting, by the code duplication detection computer program, a top most ranked centroid; and

identifying, by the code duplication detection computer program, the incremental function as a duplicate of data points in a cluster of the top-most centroid.

15. The method of claim 14 , wherein the dimension reduction model comprises truncated Singular Value Decomposition.

16. The method of claim 14 , further comprising:

applying, by the code duplication detection computer program, Natural Language Processing (NLP) to generate features for the abstract syntax trees.

17. The method of claim 14 , further comprising:

applying, by the code duplication detection computer program, a de-noising filter to the plurality of the abstract syntax trees.

18. The method of claim 17 , wherein the de-noising filter filters abstract syntax trees that are irrelevant.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Aug 14, 2024
From: SAPHAL, ROHAN; SILAVONG, FANNY; MORAN, SEAN; GEORGIADIS, ANTONIOS; SAHA, SANAT; SINGH, GAURAV; OSSELIN, PIERRE; OTTER, ROB
To: JPMORGAN CHASE BANK, N.A.
Reel/Frame 068277/0323 →
Priority Claims (1)
GR 20210100873 · Dec 13, 2021 · national
Continuity (1)
Related Publication 20230185550A1 · Jun 15, 2023