IP Library › Granted Patent US 11,481,580
Granted Patent B2
US 11,481,580 · App. 15/994,788 · Granted Oct 25, 2022

Accessible machine learning

Inventors: Li Deng (Menlo Park, CA); Suhas Chelian (San Jose, CA); Ajay Chander (San Francisco, CA)
Assignee: FUJITSU LIMITED
G06K9/626G06K9/623G06K9/6282G06N20/00
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,481,580
App. No.
15/994,788
Granted
Oct 25, 2022
Kind
B2
Abstract

According to an aspect of an embodiment, a method may include obtaining a data set that includes categories (or features), and a target criteria. The method may further include obtaining a first decision tree model using the data set. The method may further include ranking the categories based on the first decision tree model and removing low-ranking categories from the data set. The method may further include generating a second decision tree model using the data set. The second decision tree model may include branch nodes. Each of branch nodes may represent a branch criteria. The method may further include pruning a branch node. The method may further include designating a remaining branch nodes as a rule node. The method may further include generating a rule based on the branch criteria of the rule node and presenting the rule in a graphical user interface.

Claims (86)

1. A method comprising:

obtaining a data set that includes a plurality of records, each of the records including values in a plurality of categories, one of the plurality of categories being a target category;

obtaining an indication of a target criteria where a first set of records of the plurality of records each include a first target value of the target category that meets the target criteria;

obtaining a first decision tree model using the data set, the first decision tree model representing relationships between the values of the categories of the records and the target criteria;

ranking the plurality of categories based on the first decision tree model and based on relationships between values of the ranked categories and the target criteria;

removing one or more low-ranking categories from the records of the data set based on the ranking of the low-ranking categories;

generating a second decision tree model using the data set with the low-ranking categories removed from the data set, the second decision tree model including a root node, a plurality of leaf nodes, and a plurality of branch nodes, each of the plurality of branch nodes representing a branch criteria of one of the plurality of categories, the branch criteria for each of the plurality of branch nodes selected based on relationships between the first target values that meet the target criteria and values of the one of the plurality of categories that meet the branch criteria;

pruning a branch node of the plurality of branch nodes, the pruned branch node selected for pruning based on a second set of records of the data set associated with the pruned branch node including more records that include second target values that do not meet the target criteria than records of the first set of record that include the first target values that meet the target criteria;

designating at least one of the remaining branch nodes as a rule node;

generating a rule based on the branch criteria of the rule node; and

presenting the rule on a display in a graphical user interface.

2. The method of claim 1 , further comprising:

presenting the plurality of categories of the data set on the display in the graphical user interface;

receiving from the graphical user interface an indication of a category; and

removing the category from records of the data set.

3. The method of claim 1 , wherein the method further comprises presenting on the display in the graphical user interface a visual representation of values of a selected category of the data set based on relationships between the values of the selected category and target values.

4. The method of claim 1 , further comprising:

generating a list of nodes that includes the plurality of branch nodes and the plurality of leaf nodes that include a subset of records associated therewith that includes more records that include third target values that do not meet the target criteria than records of the first set of records that include the first target values that meet the target criteria;

removing a first child node from the list of nodes based on the first child node having a first parent node in the list of nodes;

removing a second child node from the list of nodes; and

adding a second parent node to the list of nodes, the second parent node being a parent node of the second child node.

5. The method of claim 1 , wherein the rule includes a precondition based on a parent nodes of the rule node and a post-condition based on the branch criteria of the rule node.

6. The method of claim 1 , the method further comprising:

receiving from the graphical user interface an indication of a first percentage of records that have values of the target category that meet the target criteria;

selecting a second set of records from the data set, the second set of records having a second percentage of records that meet the target criteria, wherein the second percentage is within a threshold distance of the first percentage; and

presenting on the display in the graphical user interface one or more values of the second set of records.

7. The method of claim 1 , the method further comprising:

receiving a plurality of values of an additional record not included in the data set, each of the plurality of values corresponding to a different one of a subset of the categories of the plurality of records, the subset of categories not including the target category;

predicting whether the additional record is likely to include a third target value that meets the target criteria based on the plurality of values of the additional record; and

displaying at the graphical user interface results of the predicting whether the additional record is likely to meet the target criteria.

8. The method of claim 7 , the method further comprising:

receiving from the graphical user interface a new value for one of the subset of the categories, of the additional record;

predicting whether the additional record including the new value is likely to include a fourth target value that meets the target criteria based on the plurality of values of the additional record including the new value; and

displaying at the graphical user interface results of the predicting whether the additional record including the new value is likely to meet the target criteria.

9. At least one non-transitory computer-readable media configured to store one or more instructions that when executed by at least one processor cause or direct a system to perform the method of claim 1 .

10. At least one non-transitory computer-readable media configured to store one or more instructions that when executed by at least one processor cause or direct a system to perform operations, the operations comprising:

obtaining a data set that includes a plurality of records, each of the records including values in a plurality of categories, one of the plurality of categories being a target category;

obtaining an indication of a target criteria where a first set of records of the plurality of records each include a first target value of the target category that meets the target criteria;

generating a decision tree model using the data set, the decision tree model including a root node, a plurality of leaf nodes, and a plurality of branch nodes, each of the plurality of branch nodes representing a branch criteria of one of the plurality of categories, the branch criteria for each of the plurality of branch nodes selected based on relationships between the first target values that meet the target criteria and values of the one of the plurality of categories that meet the branch criteria;

pruning a branch node of the plurality of branch nodes, the pruned branch node selected for pruning based on a second set of records of the data set associated with the pruned branch node including more records that include second target values that do not meet the target criteria than records of the first set of records that include the first target values that meet the target criteria;

designating at least one of the remaining branch nodes as a rule node;

generating a rule based on the branch criteria of the rule node; and

presenting the rule on a display in a graphical user interface.

11. The non-transitory computer-readable media of claim 10 , wherein the operations further comprise prior to the generating the decision tree model:

obtaining a second decision tree model using the data set;

ranking the plurality of categories based on the second decision tree model and based on relationships between values of the ranked categories and the target criteria; and

removing one or more low-ranking categories from the records of the data set based on the ranking of the low-ranking categories,

wherein generating the decision tree model using the data set is based on the data set with the low-ranking categories removed from the data set.

12. The non-transitory computer-readable media of claim 10 , wherein the operations further comprise:

presenting the plurality of categories of the data set on the display in the graphical user interface;

receiving from the graphical user interface an indication of a category; and

removing the category from records of the data set.

13. The non-transitory computer-readable media of claim 10 , wherein the operations further comprise presenting on the display in the graphical user interface a visual representation of values of a selected category of the data set based on relationships between the values of the selected category and target values.

14. The non-transitory computer-readable media of claim 10 , wherein the operations further comprise:

generating a list of nodes that includes the plurality of branch nodes and the plurality of leaf nodes that include a subset of records associated therewith that includes more records that include third target values that do not meet the target criteria than records of the first set of records that include the first target values that meet the target criteria;

removing a first child node from the list of nodes based on the first child node having a first parent node in the list of nodes;

removing a second child node from the list of nodes; and

adding a second parent node to the list of nodes, the second parent node being a parent node of the second child node.

15. The non-transitory computer-readable media of claim 10 , wherein the rule includes a precondition based on a parent nodes of the rule node and a post-condition based on the branch criteria of the rule node.

16. The non-transitory computer-readable media of claim 10 , wherein the operations further comprise:

receiving from the graphical user interface an indication of a first percentage of records that have values of the target category that meet the target criteria;

selecting a second set of records from the data set, the second set of records having a second percentage of records that meet the target criteria, wherein the second percentage is within a threshold distance of the first percentage; and

presenting on the display in the graphical user interface one or more values of the second set of records.

17. The non-transitory computer-readable media of claim 10 , wherein the operations further comprise:

receiving a plurality of values of an additional record not included in the data set, each of the plurality of values corresponding to a different one of a subset of the categories of the plurality of records, the subset of categories not including the target category;

predicting whether the additional record is likely to include a third target value that meets the target criteria based on the plurality of values of the additional record; and

displaying at the graphical user interface results of the predicting whether the additional record is likely to meet the target criteria.

18. The non-transitory computer-readable media of claim 17 , wherein the operations further comprise:

receiving from the graphical user interface a new value for one of the subset of the categories, of the additional record;

predicting whether the additional record including the new value is likely to include a fourth target value that meets the target criteria based on the plurality of values of the additional record including the new value; and

displaying at the graphical user interface results of the predicting whether the additional record including the new value is likely to meet the target criteria.

19. A system comprising:

one or more computer-readable media configured to store one or more instructions; and

one or more processors coupled to the one or more computer-readable media, the one or more processors configured to execute the one or more instructions to cause or direct the system to perform operations comprising:

obtaining a data set that includes a plurality of records, each of the records including values in a plurality of categories, one of the plurality of categories being a target category;

obtaining an indication of a target criteria where a first set of records of the plurality of records each include a first target value of the target category that meets the target criteria;

generating a decision tree model using the data set, the decision tree model including a root node, a plurality of leaf nodes, and a plurality of branch nodes, each of the plurality of branch nodes representing a branch criteria of one of the plurality of categories, the branch criteria for each of the plurality of branch nodes selected based on relationships between the first target values that meet the target criteria and values of the one of the plurality of categories that meet the branch criteria;

pruning a branch node of the plurality of branch nodes, the pruned branch node selected for pruning based on a second set of records of the data set associated with the pruned branch node including more records that include second target values that do not meet the target criteria than records of the first set of records that include the first target values that meet the target criteria;

designating at least one of the remaining branch nodes as a rule node;

generating a rule based on the branch criteria of the rule node; and

presenting the rule on a display in a graphical user interface.

20. The system of claim 19 , wherein the operations further comprise:

generating a list of nodes that includes the plurality of branch nodes and the plurality of leaf nodes that include a subset of records associated therewith that includes more records that include third target values that do not meet the target criteria than records of the first set of records that include the first target values that meet the target criteria;

removing a first child node from the list of nodes based on the first child node having a first parent node in the list of nodes;

removing a second child node from the list of nodes; and

adding a second parent node to the list of nodes, the second parent node being a parent node of the second child node.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jun 1, 2018
From: DENG, LI; CHELIAN, SUHAS; CHANDER, AJAY
To: FUJITSU LIMITED
Reel/Frame 045966/0172 →
Continuity (1)
Related Publication 20190370600A1 · Dec 5, 2019