IP Library › Granted Patent US 12,455,808
Granted Patent B2
US 12,455,808 · App. 17/970,262 · Granted Oct 28, 2025

Duplicate incident detection using dynamic similarity threshold

Inventors: David C. Sydow (Merrimack, NH); Anil Kumar Koluguri (Durham, NC); Jeremy Denis White (Londonderry, NH); Shobhit Nitinkumar Dutia (Westborough, MA); Duhita Mulky Avinash (Melrose, MA)
Assignee: Dell Products L.P.
G06F11/3616G06F8/71G06F16/215G06F16/2379
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,455,808
App. No.
17/970,262
Granted
Oct 28, 2025
Kind
B2
Abstract

Methods, apparatus, and processor-readable storage media for duplicate incident detection using a dynamic similarity threshold are provided herein. An example computer-implemented method includes obtaining a request including information associated with tracking at least a first incident in a database; generating a first representation of the first incident that encodes at least a portion of the information; computing a set of similarity scores for the first incident, where a given similarity score is based on a comparison between the first representation and a second representation generated for one of a plurality of additional incidents in the database; detecting that the first incident is a duplicate of at least one of the plurality of additional incidents based on a comparison of the set of similarity scores to a similarity threshold, where the similarity threshold is updated over time; and initiating an update in the database in response to the detecting.

Claims (56)

1. A computer-implemented method comprising:

obtaining a request comprising information associated with tracking at least a first incident in an incident database;

generating a first numeric representation of the first incident that encodes at least a portion of the information;

computing a set of one or more similarity scores for the first incident, wherein a given similarity score in the set is based at least in part on a comparison between the first numeric representation and a second numeric representation generated for one of a plurality of additional incidents in the incident database;

detecting that the first incident is a duplicate of at least one of the plurality of additional incidents based on a comparison of the set of similarity scores to a similarity threshold, wherein the similarity threshold is dynamically updated over time based on at least one of: one or more updates to one or more of the additional incidents in the incident database and one or more new incidents added to the incident database; and

in response to the detecting, initiating an update to the incident database, wherein the update comprises at least one of: (i) assigning a flag to at least one of: (a) the first incident and (b) the at least one of the plurality of additional incidents as being a duplicate incident, and (ii) merging the first incident and the at least one of the plurality of additional incidents in the incident database;

wherein the method is performed by at least one processing device comprising a processor coupled to a memory.

2. The computer-implemented method of claim 1 , wherein the first incident relates to a software defect in one or more software code repositories, and wherein the information associated with tracking the first incident comprises at least one of:

a title of the first incident;

one or more user comments related to the software defect;

a location of the software defect within software code associated with the one or more software code repositories; and

a release version corresponding to the first incident.

3. The computer-implemented method of claim 2 , wherein the similarity threshold is separately determined for at least one of: the release version corresponding to the first incident and the location of the software defect.

4. The computer-implemented method of claim 1 , wherein the first numeric representation comprises a vector representation that is generated using a term frequency-inverse document frequency process.

5. The computer-implemented method of claim 1 , wherein the similarity threshold is updated over time based at least in part on information provided by one or more users, wherein the information verifies that at least one of: (i) the first incident in the incident database is the duplicate of the at least one of the plurality of additional incidents and (two or more other incidents in the incident database are duplicates.

6. The computer-implemented method of claim 1 , wherein the similarity threshold is computed at least in part by:

generating at least one numeric representation for each of the plurality of additional incidents, wherein at least a portion of the plurality of additional incidents are labeled as duplicate incidents;

computing similarity scores between each pair of numeric representations of the plurality of numeric representations, wherein the computing comprises computing a cosine similarity between each pair of numeric representations of the plurality of numeric representations; and

calculating the similarity threshold based on the similarity scores corresponding to the pairs of incidents labeled as duplicates.

7. The computer-implemented method of claim 1 , wherein the similarity threshold is updated in response to detecting that the first incident is the duplicate of the at least one of the plurality of additional incidents.

8. The computer-implemented method of claim 1 , wherein the initiating the update to the incident database, further comprises:

adding one or more comments related to at least one of the first incident and the at least one of the plurality of additional incidents in the incident database.

9. The computer-implemented method claim 1 , wherein the similarity threshold is dynamically updated by recomputing a value based on an average similarity of incident pairs currently labeled as duplicates in the incident database.

10. The computer-implemented method of claim 1 , wherein the similarity threshold is applied across each incident in the incident database.

11. A non-transitory processor-readable storage medium having stored therein program code of one or more software programs, wherein the program code when executed by at least one processing device causes the at least one processing device:

to obtain a request comprising information associated with tracking at least a first incident in an incident database;

to generate a first numeric representation of the first incident that encodes at least a portion of the information;

to compute a set of one or more similarity scores for the first incident, wherein a given similarity score in the set is based at least in part on a comparison between the first numeric representation and a second numeric representation generated for one of a plurality of additional incidents in the incident database;

to detect that the first incident is a duplicate of at least one of the plurality of additional incidents based on a comparison of the set of similarity scores to a similarity threshold, wherein the similarity threshold is dynamically updated over time based on at least one of: one or more updates to one or more of the additional incidents in the incident database and one or more new incidents added to the incident database; and

in response to the detecting, to initiate an update to the incident database, wherein the update comprises at least one of: (i) assigning a flag to at least one of: (a) the first incident and (b) the at least one of the plurality of additional incidents as being a duplicate incident, and (ii) merging the first incident and the at least one of the plurality of additional incidents in the incident database.

12. The non-transitory processor-readable storage medium of claim 11 , wherein the first incident relates to a software defect in one or more software code repositories.

13. The non-transitory processor-readable storage medium of claim 12 , wherein the information associated with tracking the first incident comprises at least one of:

a title of the first incident;

one or more user comments related to the software defect;

a location of the software defect within software code associated with the one or more software code repositories; and

a release version corresponding to the first incident.

14. The non-transitory processor-readable storage medium of claim 13 , wherein the similarity threshold is separately determined for at least one of: the release version corresponding to the first incident and the location of the software defect.

15. The non-transitory processor-readable storage medium of claim 11 , wherein the similarity threshold is updated over time based at least in part on information provided by one or more users, wherein the information verifies that at least one of: (i) the first incident in the incident database is the duplicate of the at least one of the plurality of additional incidents and (ii) two or more other incidents in the incident database are duplicates.

16. The non-transitory processor-readable storage medium of claim 11 , wherein the similarity threshold is computed at least in part by:

generating at least one numeric representation for each of the plurality of additional incidents, wherein at least a portion of the plurality of additional incidents are labeled as duplicate incidents;

computing similarity scores between each pair of numeric representations of the plurality of numeric representations; and

calculating the similarity threshold based on the similarity scores corresponding to the pairs of incidents labeled as duplicates.

17. An apparatus comprising:

at least one processing device comprising a processor coupled to a memory;

the at least one processing device being configured:

to obtain a request comprising information associated with tracking at least a first incident in an incident database;

to generate a first numeric representation of the first incident that encodes at least a portion of the information;

to compute a set of one or more similarity scores for the first incident, wherein a given similarity score in the set is based at least in part on a comparison between the first numeric representation and a second numeric representation generated for one of a plurality of additional incidents in the incident database;

to detect that the first incident is a duplicate of at least one of the plurality of additional incidents based on a comparison of the set of similarity scores to a similarity threshold, wherein the similarity threshold is dynamically updated over time based on at least one of: one or more updates to one or more of the additional incidents in the incident database and one or more new incidents added to the incident database; and

in response to the detecting, to initiate an update to the incident database, wherein the update comprises at least one of: (i) assigning a flag to at least one of: (a) the first incident and (b) the at least one of the plurality of additional incidents as being a duplicate incident, and (ii) merging the first incident and the at least one of the plurality of additional incidents in the incident database.

18. The apparatus of claim 17 , wherein the first incident relates to a software defect in one or more software code repositories.

19. The apparatus of claim 17 , wherein the similarity threshold is updated over time based at least in part on information provided by one or more users, wherein the information verifies that at least one of: the first incident in the incident database is a duplicate and two or more other incidents in the incident database are duplicates.

20. The apparatus of claim 17 , wherein the similarity threshold is computed at least in part by:

generating at least one numeric representation for each of a plurality of incidents, wherein at least a portion of the plurality of incidents are labeled as duplicates;

computing similarity scores between each pair of numeric representations of the plurality of numeric representations; and

calculating the similarity threshold based on the similarity scores corresponding to the pairs of incidents labeled as duplicates.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Oct 20, 2022
From: SYDOW, DAVID C.; KOLUGURI, ANIL KUMAR; WHITE, JEREMY DENIS; DUTIA, SHOBHIT NITINKUMAR; AVINASH, DUHITA MULKY
To: DELL PRODUCTS L.P.
Reel/Frame 061486/0887 →
Continuity (2)
Related Publication 20240134774A1 · Apr 25, 2024
Related Publication 20240232048A9 · Jul 11, 2024
References Cited (14)
US 10089213B1 · Noble · 2018 [cited by examiner]
US 11379526B2 · Srinivas · 2022 [cited by applicant]
US 20110066908A1 · Bartz · 2011 [cited by examiner]
US 20210124876A1 · Kryscinski · 2021 [cited by applicant]
US 20210232387A1 · Hicks · 2021 [cited by examiner]
US 20220164200A1 · Ganhotra · 2022 [cited by examiner]
US 20220180247A1 · Chow · 2022 [cited by examiner]
US 20230419017A1 · Fabbri · 2023 [cited by applicant]
EP 4250133A1 · 2023 [cited by examiner]
Y.-B. Kang et al., “A computer-facilitated method for matching incident cases using semantic similarity measurement,” 2009 IFIP/IEEE International Symposium on Integrated Network Management—Workshops, New York, NY, USA,… [cited by examiner]
“How to Find Duplicates in Jira?”, https://reliex.com/blog/how-to-find-duplicates-in-jira/, Reliex OÜ, available at: https://reliex.com/blog/how-to-find-duplicates-in-jira/ (last accessed Oct. 20, 2022). [cited by applicant]
Devlin, Jacob, et al., “BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding,” arXiv preprint arXiv:1810.04805, Oct. 11, 2018. [cited by applicant]
Lewis, Mike, et al., “BART: Denoising Sequence-To-Sequence Pre-training for Natural Language Generation, Translation, and Comprehension,” arXiv preprint arXiv:1910.13461, Oct. 29, 2019. [cited by applicant]
Liu, Yinhan, et al., “RoBERTa: A Robustly Optimized BERT Pretraining Approach,” arXiv preprint arXiv:1907.11692, Jul. 26, 2019. [cited by applicant]