IP Library › Granted Patent US 11,928,156
Granted Patent B2
US 11,928,156 · App. 17/088,018 · Granted Mar 12, 2024

Learning-based automated machine learning code annotation with graph neural network

Inventors: Dakuo Wang (Cambridge, MA); Lingfei Wu (Elmsford, NY); Xuye Liu (Troy, NY); Yi Wang (Ann Arbor, MI); Chuang Gan (Cambridge, MA); Jing Xu (Xi'an, CN); Xue Ying Zhang (Xi'an, CN); Jun Wang (Xi'an, CN); Jing James Xu (Xi'an, CN)
Assignee: International Business Machines Corporation
G06F16/90332G06F16/9024G06F16/9558G06F40/211G06N3/08
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,928,156
App. No.
17/088,018
Granted
Mar 12, 2024
Kind
B2
Abstract

Obtain, at a computing device, a segment of computer code. With a classification module of a machine learning system executing on the computing device, determine a required annotation category for the segment of computer code. With an annotation generation module of the machine learning system executing on the computing device, generate a natural language annotation of the segment of computer code based on the segment of computer code and the required annotation category. Provide the natural language annotation to a user interface for display adjacent the segment of computer code.

Claims (51)

1. A method comprising:

obtaining, at a computing device, a segment of computer code;

with a classification module of a machine learning system executing on said computing device, determining a stage from a predefined set of stages of a data science workflow to which said segment of computer code corresponds and determining a required annotation category from a predefined set of categories for said segment of computer code based on said stage of said data science workflow;

with an annotation generation module of said machine learning system executing on said computing device, generating a natural language annotation of said segment of computer code based on said segment of computer code and said required annotation category;

generating an Abstract Syntax Tree (AST) structure of at least said segment of computer code, wherein said natural language annotation of said segment of computer code is further based on said Abstract Syntax Tree (AST) structure, wherein said annotation generation module includes a graph neural network; and

providing said natural language annotation to a user interface for display adjacent said segment of computer code.

2. The method of claim 1 , wherein:

said category is selected from a group consisting of process, result, education, reference, and reason; and

said stage is selected from a group consisting of environment configuration, data preparation and exploration, feature engineering and selection, and model building and selection.

3. The method of claim 2 , further comprising training said classification module by using documented notebooks, with extracted source code from said documented notebooks as input, and expert-labeled categories from said notebooks as output.

4. The method of claim 3 , further comprising:

obtaining a link or a reference to external application programming interface (API) documentation for a code function pasted in said documented notebooks;

collecting application programming interface (API) names from data science packages and short descriptions from external documentation sites related to said external application programming interface (API) documentation via a crawling script;

matching said application programming interface (API) names with said code functions and concatenating said corresponding descriptions; and

wherein said providing to said user interface for display adjacent said segment of computer code is based on said concatenated descriptions.

5. The method of claim 2 , further comprising training said annotation generation module by using documented notebooks, with extracted source code from said documented notebooks as input, and raw annotation sentences as output.

6. The method of claim 1 , further comprising obtaining user input responsive to said display and retaining said natural language annotation responsive thereto.

7. The method of claim 1 , further comprising obtaining user input responsive to said display and modifying said natural language annotation responsive thereto.

8. The method of claim 1 , further comprising correcting an error in said segment of computer code responsive to said display.

9. The method of claim 1 , wherein said user interface comprises a computational notebook implemented on a client computer device, further comprising, with said user interface, displaying said natural language annotation to a user adjacent said segment of computer code.

10. An apparatus comprising:

a memory;

a non-transitory computer readable medium including computer executable instructions; and

at least one processor, coupled to the memory and the non-transitory computer readable medium, and operative to execute the instructions to be operative to:

instantiate a classification module and an annotation generation module;

obtain a segment of computer code;

with said classification module, determine a stage from a predefined set of stages of a predefined data science workflow to which said segment of computer code corresponds and determine a required annotation category from a predefined set of categories for said segment of computer code based on said stage of said data science workflow;

with said annotation generation module, generate a natural language annotation of said segment of computer code based on said segment of computer code and said required annotation category;

generate an Abstract Syntax Tree (AST) structure of at least said segment of computer code, wherein said natural language annotation of said segment of computer code is further based on said Abstract Syntax Tree (AST) structure, wherein said annotation generation module includes a graph neural network; and

provide said natural language annotation to a user interface for display adjacent said segment of computer code.

11. The apparatus of claim 10 , wherein:

said category is selected from a group consisting of process, result, education, reference, and reason; and

said stage is selected from a group consisting of environment configuration, data preparation and exploration, feature engineering and selection, and model building and selection.

12. The apparatus of claim 11 , wherein said at least one processor is further operative to train said classification module by using documented notebooks, with extracted source code from said documented notebooks as input, and expert-labeled categories from said notebooks as output.

13. The apparatus of claim 11 , wherein said at least one processor is further operative to train said annotation generation module by using documented notebooks, with extracted source code from said documented notebooks as input, and raw annotation sentences as output.

14. The apparatus of claim 10 , further comprising a computational notebook implemented on a client computer device interconnected with said at least one processor, wherein said natural language annotation is displayed to a user adjacent said segment of computer code.

15. A non-transitory computer readable medium comprising computer executable instructions which when executed by a computer cause the computer to perform a method of:

instantiating a classification module and an annotation generation module;

obtaining, at said computer, a segment of computer code;

with said classification module, determining a stage from a predefined set of stages of a predefined data science workflow to which said segment of computer code corresponds and determining a required annotation category from a predefined set of categories for said segment of computer code based on said stage of said data science workflow;

with said annotation generation module, generating a natural language annotation of said segment of computer code based on said segment of computer code and said required annotation category;

generating an Abstract Syntax Tree (AST) structure of at least said segment of computer code, wherein said natural language annotation of said segment of computer code is further based on said Abstract Syntax Tree (AST) structure, wherein said annotation generation module includes a graph neural network;

providing said natural language annotation to a user interface for display adjacent said segment of computer code;

training said classification module by using documented notebooks, with extracted source code from said documented notebooks as input, and expert-labeled categories from said notebooks as output;

obtaining a link or a reference to external application programming interface (API) documentation for a code function pasted in said documented notebooks;

collecting application programming interface (API) names from data science packages and short descriptions from external documentation sites related to said external application programming interface (API) documentation via a crawling script;

matching said application programming interface (API) names with said code functions and concatenating said corresponding descriptions; and

wherein said providing to said user interface for display adjacent said segment of computer code is based on said concatenated descriptions;

wherein:

said category is selected from a group consisting of process, result, education, reference, and reason; and

said stage is selected from a group consisting of environment configuration, data preparation and exploration, feature engineering and selection, and model building and selection.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jan 25, 2024
From: WANG, DAKUO; WU, LINGFEI; LIU, XUYE; WANG, YI; GAN, CHUANG; XU, JING; ZHANG, XUE YING; WANG, JUN; XU, JING JAMES
To: INTERNATIONAL BUSINESS MACHINES CORPORATION
Reel/Frame 066252/0391 →
Continuity (1)
Related Publication 20220138266A1 · May 5, 2022
Cited By (1)
US 12,717,577