IP Library Granted Patent US 11,068,377
Granted Patent B2
US 11,068,377 · App. 16/586,625 · Granted Jul 20, 2021

Classifying warning messages generated by software developer tools

Inventors: Andrew Walenstein (Issaquah, WA); Andrew James Malton (Waterloo, CA); Jong Chun Park (Grapevine, TX); Hanyang Hu (Ottawa, CA)
Assignee: BlackBerry Limited
G06F11/362G06F8/43G06F11/36G06N5/003G06N20/00G06N20/20
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,068,377
App. No.
16/586,625
Granted
Jul 20, 2021
Kind
B2
Abstract

A method for classifying warning messages generated by software developer tools includes receiving a first data set. The first data set includes a first plurality of data entries, where each data entry is associated with a warning message generated based on a first set of software codes, includes indications for a plurality of features, and is associated with one of a plurality of class labels. A second data set is generated by sampling the first data set. Based on the second data set, at least one feature is selected from the plurality of features. A third data set is generated by filtering the second data set with the selected at least one feature. A machine learning classifier is determined based on the third data set. The machine learning classifier is used to classify a second warning message generated based on a second set of software codes to one of the plurality of class labels.

Claims (43)

1. A method, comprising:

receiving, by a hardware processor, a first data set, the first data set including a first plurality of data entries, wherein each data entry is associated with a warning message generated based on a first set of software codes, each data entry includes indications for a plurality of features, and each data entry is associated with one of a plurality of class labels;

generating, by the hardware processor, a second data set by sampling the first data set;

based on the second data set, selecting, by the hardware processor, at least one feature from the plurality of features, wherein selecting the at least one feature comprises selecting the at least one feature that is clustered above a cut-off value, wherein the cut-off value is a Spearman value;

generating, by the hardware processor, a third data set by filtering the second data set with the selected at least one feature;

determining, by the hardware processor, a machine learning classifier based on the third data set, wherein determining the machine learning classifier comprises dividing the third data set into a training data set and a testing data set; and

classifying, by the hardware processor, a second warning message generated based on a second set of software codes to one of the plurality of class labels using the machine learning classifier, wherein the second set of software codes is different than the first set of software codes.

2. The method of claim 1 , wherein the selecting the at least one feature comprises:

determining that more than one feature is clustered above the cut-off value; and

randomly selecting one feature from each cluster above the cut-off value.

3. The method of claim 1 , wherein the plurality of class labels includes a first class label for fixing the warning message and a second class label for ignoring the warning message.

4. The method of claim 1 , wherein the first data set is an imbalanced data set.

5. The method of claim 1 , wherein the plurality of features includes features associated with at least one of a software development process, a programming code, a software code change, or a fault finding tool analysis.

6. The method of claim 5 , wherein at least one of a stratified sampling or a stratified K-fold sampling is applied to the training data set.

7. A device, comprising:

a memory; and

at least one hardware processor communicatively coupled with the memory and configured to:

receive a first data set, the first data set including a first plurality of data entries, wherein each data entry is associated with a warning message generated based on a first set of software codes, each data entry includes indications for a plurality of features, and each data entry is associated with one of a plurality of class labels;

generate a second data set by sampling the first data set;

based on the second data set, select at least one feature from the plurality of features, wherein selecting the at least one feature comprises selecting the at least one feature that is clustered above a cut-off value, wherein the cut-off value is a Spearman value;

generate a third data set by filtering the second data set with the selected at least one feature;

determine a machine learning classifier based on the third data set, wherein determining the machine learning classifier comprises dividing the third data set into a training data set and a testing data set; and

classify a second warning message generated based on a second set of software codes to one of the plurality of class labels using the machine learning classifier, wherein the second set of software codes is different than the first set of software codes.

8. The device of claim 7 , wherein the selecting the at least one feature comprises:

determining that more than one feature is clustered above the cut-off value; and

randomly selecting one feature from each cluster above the cut-off value.

9. The device of claim 7 , wherein the plurality of class labels includes a first class label for fixing the warning message and a second class label for ignoring the warning message.

10. The device of claim 7 , wherein the first data set is an imbalanced data set.

11. The device of claim 7 , wherein the plurality of features includes features associated with at least one of a software development process, a programming code, a software code change, or a fault finding tool analysis.

12. The device of claim 11 , wherein at least one of a stratified sampling or a stratified K-fold sampling is applied to the training data set.

13. A non-transitory computer-readable medium containing instructions which, when executed, cause a computing device to perform operations comprising:

receiving a first data set, the first data set including a first plurality of data entries, wherein each data entry is associated with a warning message generated based on a first set of software codes, each data entry includes indications for a plurality of features, and each data entry is associated with one of a plurality of class labels;

generating a second data set by sampling the first data set;

based on the second data set, selecting at least one feature from the plurality of features, wherein selecting the at least one feature comprises selecting the at least one feature that is clustered above a cut-off value, wherein the cut-off value is a Spearman value;

generating a third data set by filtering the second data set with the selected at least one feature;

determining a machine learning classifier based on the third data set, wherein determining the machine learning classifier comprises dividing the third data set into a training data set and a testing data set; and

classifying a second warning message generated based on a second set of software codes to one of the plurality of class labels using the machine learning classifier, wherein the second set of software codes is different than the first set of software codes.

14. The non-transitory computer-readable medium of claim 13 , wherein the selecting the at least one feature comprises:

determining that more than one feature is clustered above the cut-off value; and

randomly selecting one feature from each cluster above the cut-off value.

15. The non-transitory computer-readable medium of claim 13 , wherein the plurality of class labels includes a first class label for fixing the warning message and a second class label for ignoring the warning message.

16. The non-transitory computer-readable medium of claim 13 , wherein the first data set is an imbalanced data set.

17. The non-transitory computer-readable medium of claim 13 , wherein the plurality of features includes features associated with at least one of a software development process, a programming code, a software code change, or a fault finding tool analysis.

Assignments (5)
NUNC PRO TUNC ASSIGNMENT Recorded Jun 19, 2023
From: BLACKBERRY LIMITED
To: MALIKIE INNOVATIONS LIMITED
Reel/Frame 064271/0199 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jun 16, 2023
From: BLACKBERRY LIMITED
To: MALIKIE INNOVATIONS LIMITED
Reel/Frame 064104/0103 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Apr 24, 2020
From: MALTON, ANDREW JAMES; HU, HANYANG
To: BLACKBERRY LIMITED
Reel/Frame 052492/0572 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Apr 24, 2020
From: WALENSTEIN, ANDREW; PARK, JONG CHUN
To: BLACKBERRY CORPORATION
Reel/Frame 052492/0658 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Apr 24, 2020
From: BLACKBERRY CORPORATION
To: BLACKBERRY LIMITED
Reel/Frame 052492/0767 →
Continuity (2)
Continuation 15725250 · Oct 4, 2017
Related Publication 20200026636A1 · Jan 23, 2020