IP Library Granted Patent US 11,521,075
Granted Patent B2
US 11,521,075 · App. 16/917,267 · Granted Dec 6, 2022

Transfer learning system for automated software engineering tasks

Inventors: Colin Bruce Clement (Seattle, WA); Dawn Drain (Bellevue, WA); Neelakantan Sundaresan (Bellevue, WA); Alexey Svyatkovskiy (Bellevue, WA)
Assignee: MICROSOFT TECHNOLOGY LICENSING, LLC
G06N3/088G06F8/30G06F40/40G06N3/0454G06N20/00
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,521,075
App. No.
16/917,267
Granted
Dec 6, 2022
Kind
B2
Abstract

A transfer learning system is used for the development of neural transformer models pertaining to software engineering tasks. The transfer learning system trains source code domain neural transformer models with attention in various configurations on a large corpus of unsupervised training dataset of source code programs and/or source code-related natural language text. A web service provides the trained models for use in developing a model that may be fine-tuned on a supervised training dataset associated with a software engineering task thereby generating a tool to perform the software engineering task.

Claims (60)

1. A system comprising:

one or more processors; and

a memory that stores one or more programs that are configured to be executed by the one or more processors, the one or more programs including instructions to perform actions that:

train a plurality of source code domain neural transformer models with attention on an unsupervised training dataset of source code, the plurality of source code domain neural transformer models with attention including an encoder-only neural transformer model with attention, a decoder-only neural transformer model with attention, or an encoder-decoder neural transformer model with attention;

obtain a supervised training dataset for a specific software engineering task;

select one of the plurality of source code domain neural transformer models with attention; and

fine-tune the selected source code domain neural transformer model with attention with the supervised training dataset to perform the specific software engineering task.

2. The system of claim 1 , wherein the one or more programs include further instructions to perform actions that:

associate one or more software engineering tasks with a particular one of the plurality of source code domain neural transformer models; and

choose the selected source code domain neural transformer model with attention based on the associated software engineering task.

3. The system of claim 1 , wherein the one or more programs include further instructions to perform actions that:

train at least one of the plurality of source code domain neural transformer models with attention in a plurality of standard model sizes.

4. The system of claim 3 , wherein the one or more programs include further instructions to perform actions that:

obtain a requested model size;

choose a standard model size closest to the requested model size; and

alter one or more blocks of the selected source code domain neural transformer model with attention in the standard model size to meet the requested model size.

5. The system of claim 4 , wherein the one or more programs include further instructions to perform actions that:

perform knowledge distillation on unaltered blocks of the selected source code domain neural transformer model with attention.

6. The system of claim 3 ,

wherein a standard model-sized source code doamin neural transformer model with attention has a memory size; and

wherein the one or more programs include further instructions to perform actions that:

obtain a requested memory size; and

alter the selected source code domain neural transformer model with attention to meet the requested memory size.

7. The system of claim 1 , wherein the unsupervised training dataset includes natural language text of source code summaries.

8. A computer-implemented method, comprising:

providing a plurality of neural transformer models with attention having been trained on an unsupervised training dataset of source code, each model having a standard configuration of transformer blocks;

obtaining a request to train a second neural transformer model with attention with a requested configuration of transformer blocks that is less than the standard configuration to perform a particular software engineering task;

transferring a subset of the transformer blocks of a select one of the plurality of neural transformer models to configure the second neural transformer model with attention with the requested configuration of transformer blocks; and

training the second neural transformer model with a supervised training dataset to perform the particular software engineering task.

9. The computer-implemented method of claim 8 , further comprising:

configuring a first one of the plurality of neural transformer models with attention with encoder-only transformer blocks; and

associating a classification software engineering task with the first one of the plurality of neural transformer models with attention.

10. The computer-implemented method of claim 9 , further comprising:

replacing an output layer of the first one of the plurality of neural transformer models with attention with a classification layer configured for the supervised training dataset.

11. The computer-implemented method of claim 8 , further comprising:

configuring a second one of the plurality of neural transformer models with attention with decoder-only transformer blocks; and

associating an auto-regressive software engineering task with the second one of the plurality of neural transformer models with attention.

12. The computer-implemented method of claim 8 , further comprising:

configuring a third one of the plurality of neural transformer models with attention with encoder-decoder transformer blocks; and

associating a machine translation software engineering task with the third one of the plurality of neural transformer models with attention.

13. The computer-implemented method of claim 8 , further comprising:

employing a scaling function to determine the number of transformer blocks to transfer; and

applying knowledge distillation to the transferred transformer blocks.

14. The computer-implemented method of claim 8 , wherein the supervised training dataset includes source code snippets from different programming languages.

15. The computer-implemented method of claim 8 , wherein the supervised training dataset includes natural language code summaries.

16. A device, comprising:

one or more processors and a memory;

wherein the one or more processors are configured to perform actions that:

train a set of neural transformer models with attention on an unsupervised training dataset of source code snippets, the set including a neural transformer model with attention having encoder-only blocks, a neural transformer model with attention having decoder-only blocks, and a neural transformer model with attention having encoder-decoder blocks;

obtain a supervised training dataset of a software engineering task;

select one of the neural transformer models with attention;

transfer the blocks of the selected neural transformer model with attention to a second neural transformer model with attention; and

fine-tune the second neural transformer model with attention with the supervised training dataset to generate a tool that performs the software engineering task.

17. The device of claim 16 , wherein the unsupervised training dataset of source code snippets includes source code snippets in different programming languages.

18. The device of claim 16 , wherein the one or more processors are configured to perform actions that:

apply knowledge distillation to the transferred blocks.

19. The device of claim 16 , wherein the one or more processors are configured to perform actions that:

associate the software engineering task with a select one of the neural transformer models with attention.

20. The device of claim 16 , wherein the one or more processors are configured to perform actions that:

deploy the tool in an integrated development environment.

Assignments (3)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jul 2, 2022
From: DRAIN, DAWN
To: MICROSOFT TECHNOLOGY LICENSING, LLC.
Reel/Frame 060393/0520 →
CORRECTIVE ASSIGNMENT TO CORRECT THE NAME OF THE INVENTOR ALEX SVYATKOVSKIY PREVIOUSLY RECORDED AT REEL: 053110 FRAME: 0837. ASSIGNOR(S) HEREBY CONFIRMS THE ASSIGNMENT. Recorded Feb 19, 2021
From: SVYATKOVSKIY, ALEXEY
To: MICROSOFT TECHNOLOGY LICENSING, LLC.
Reel/Frame 055341/0932 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jul 2, 2020
From: CLEMENT, COLIN BRUCE; DRAIN, JAMES; SUNDARESAN, NEELAKANTAN; SVYAKOVSKIY, ALEXEY
To: MICROSOFT TECHNOLOGY LICENSING, LLC.
Reel/Frame 053110/0837 →
Continuity (2)
Provisional Application 63025529 · May 15, 2020
Related Publication 20210357762A1 · Nov 18, 2021
Cited By (2)
US 12,461,993 US 12,566,244