IP Library › Granted Patent US 12,333,277
Granted Patent B2
US 12,333,277 · App. 18/618,371 · Granted Jun 17, 2025

Machine-learned models for generating code snippets with predicted placeholders for optimizing software development

Inventors: Daniel Dun-ning Woo Johnson (Toronto, CA); Daniel Stefan Tarlow (Montréal, CA); Maxim Tabachnyk (Munich, DE); Marc Hatcher Rasi (Sunnyvale, CA); Jacob Austin (New York, NY); Hassan Abolhassani (Palo Alto, CA); Jacob Hanson Hegna (Minneapolis, MN)
Assignee: GOOGLE LLC
G06F8/33
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,333,277
App. No.
18/618,371
Granted
Jun 17, 2025
Kind
B2
Abstract

Systems and methods of the present disclosure are directed to a method for machine-learned code segment prediction for optimizing software development. The method includes obtaining an incomplete segment of code. The method includes processing the incomplete segment of code with a machine-learned code prediction model to obtain a sampled set of segment completion predictions that include code that completes the incomplete segment of code. The method includes determining an aggregated segment completion prediction from the sampled set of segment completion predictions. The method includes replacing a portion of the aggregated segment completion prediction with an input field, wherein the portion of the aggregated segment completion prediction is associated with a degree of certainty less than a threshold degree of certainty.

Claims (40)

1. A computer-implemented method for machine-learned code segment prediction for optimizing software development, comprising:

obtaining, by a computing system comprising one or more computing devices, an incomplete segment of code;

processing, by the computing system, the incomplete segment of code with a machine-learned code prediction model to obtain a segment completion prediction, wherein the segment completion prediction comprises code that completes the incomplete segment of code and an input field, wherein the input field replaces a portion of the segment completion prediction associated with a degree of certainty less than a threshold degree of certainty; and

providing, by the computing system, the segment completion prediction for display to a user.

2. The computer-implemented method of claim 1 , wherein processing the incomplete segment of code comprises:

processing, by the computing system, the incomplete segment of code with a machine-learned code prediction distillation model to obtain the segment completion prediction, wherein the machine-learned code prediction distillation model is trained based on a teacher model, wherein the teacher model is trained to generate a sampled set of segment completion predictions that are aggregated to obtain an aggregated segment completion prediction, and wherein a portion of the aggregated segment completion prediction associated with the degree of certainty less than the threshold degree of certainty is replaced with an input field.

3. The computer-implemented method of claim 2 , wherein the method further comprises:

evaluating, by the computing system, a distillation loss function that evaluates a difference between the segment completion prediction and a ground truth completion prediction; and

modifying, by the computing system, one or more values of one or more parameters of the machine-learned code prediction distillation model based at least in part on the distillation loss function.

4. The computer-implemented method of claim 3 , wherein the ground truth completion prediction comprises an aggregated segment completion prediction aggregated from a sampled set of completion predictions generated by the teacher model based on the incomplete segment of code, wherein a portion of the aggregated segment completion associated with the degree of certainty less than the threshold degree of certainty replaced with an input field.

5. The computer-implemented method of claim 1 , wherein the input field comprises a suggested portion of code that is selectable by a user.

6. The computer-implemented method of claim 1 , wherein the incomplete segment of code corresponds to a location of a cursor of a user within a development environment.

7. The computer-implemented method of claim 1 , wherein the input field comprises an input field for a machine-learned language model.

8. The computer-implemented method of claim 1 , wherein the machine-learned code prediction model comprises a machine-learned language model, and wherein the segment completion prediction comprises a language output from the machine-learned language model.

9. A computing system for machine-learned code segment prediction for optimizing software development, comprising:

one or more processors;

one or more non-transitory computer-readable media that store instructions that, when executed by the one or more processors, cause the computing system to perform operations, the operations comprising:

obtaining an incomplete segment of code;

processing the incomplete segment of code with a machine-learned code prediction model to obtain a segment completion prediction, wherein the segment completion prediction comprises code that completes the incomplete segment of code and an input field, wherein the input field replaces a portion of the segment completion prediction associated with a degree of certainty less than a threshold degree of certainty; and

providing the segment completion prediction.

10. The computing system of claim 9 , wherein processing the incomplete segment of code comprises:

processing the incomplete segment of code with a machine-learned code prediction distillation model to obtain the segment completion prediction, wherein the machine-learned code prediction distillation model is trained based on a teacher model, wherein the teacher model is trained to generate a sampled set of segment completion predictions that are aggregated to obtain an aggregated segment completion prediction, and wherein a portion of the aggregated segment completion prediction associated with the degree of certainty less than the threshold degree of certainty is replaced with an input field.

11. The computing system of claim 10 , wherein the operations further comprise:

evaluating a distillation loss function that evaluates a difference between the segment completion prediction and a ground truth completion prediction; and

modifying one or more values of one or more parameters of the machine-learned code prediction distillation model based at least in part on the distillation loss function.

12. The computing system of claim 11 , wherein the ground truth completion prediction comprises an aggregated segment completion prediction aggregated from a sampled set of completion predictions generated by the teacher model based on the incomplete segment of code, wherein a portion of the aggregated segment completion associated with the degree of certainty less than the threshold degree of certainty replaced with an input field.

13. The computing system of claim 9 , wherein the input field comprises a suggested portion of code that is selectable by a user.

14. The computing system of claim 9 , wherein the incomplete segment of code corresponds to a location of a cursor of a user within a development environment.

15. The computing system of claim 9 , wherein the input field comprises an input field for a machine-learned language model.

16. The computing system of claim 9 , wherein the machine-learned code prediction model comprises a machine-learned language model, and wherein the segment completion prediction comprises a language output from the machine-learned language model.

17. One or more non-transitory computer-readable media that store instructions that, when executed by one or more processors of a computing system, cause the computing system to perform operations, the operations comprising:

obtaining an incomplete segment of code;

processing the incomplete segment of code with a machine-learned code prediction model to obtain a segment completion prediction, wherein the segment completion prediction comprises code that completes the incomplete segment of code and an input field, wherein the input field replaces a portion of the segment completion prediction associated with a degree of certainty less than a threshold degree of certainty; and

providing the segment completion prediction as a suggestion within an Integrated Development Environment (IDE).

18. The one or more non-transitory computer-readable media of claim 17 , wherein processing the incomplete segment of code comprises:

processing the incomplete segment of code with a machine-learned code prediction distillation model to obtain the segment completion prediction, wherein the machine-learned code prediction distillation model is trained based on a teacher model, wherein the teacher model is trained to generate a sampled set of segment completion predictions that are aggregated to obtain an aggregated segment completion prediction, and wherein a portion of the aggregated segment completion prediction associated with the degree of certainty less than the threshold degree of certainty is replaced with an input field.

19. The one or more non-transitory computer-readable media of claim 18 , wherein the operations further comprise:

evaluating a distillation loss function that evaluates a difference between the segment completion prediction and a ground truth completion prediction; and

modifying one or more values of one or more parameters of the machine-learned code prediction distillation model based at least in part on the distillation loss function.

20. The one or more non-transitory computer-readable media of claim 19 , wherein the ground truth completion prediction comprises an aggregated segment completion prediction aggregated from a sampled set of completion predictions generated by the teacher model based on the incomplete segment of code, wherein a portion of the aggregated segment completion associated with the degree of certainty less than the threshold degree of certainty replaced with an input field.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Apr 10, 2024
From: JOHNSON, DANIEL DUN-NING WOO; TARLOW, DANIEL STEFAN; TABACHNYK, MAXIM; RASI, MARC HATCHER; AUSTIN, JACOB; ABOLHASSANI, HASSAN; HEGNA, JACOB HANSON
To: GOOGLE LLC
Reel/Frame 067061/0938 →
Continuity (2)
Continuation 17832199 · Jun 3, 2022
Related Publication 20240231765A1 · Jul 11, 2024
References Cited (16)
US 10048945B1 · Makkar · 2018 [cited by examiner]
US 20190079753A1 · Makkar · 2019 [cited by examiner]
US 20190079754A1 · Makkar · 2019 [cited by examiner]
US 20190303108A1 · Fu · 2019 [cited by examiner]
US 20200097261A1 · Smith · 2020 [cited by examiner]
US 20200410390A1 · Fu · 2020 [cited by examiner]
US 20210279042A1 · Allamanis · 2021 [cited by examiner]
US 20220107802A1 · Rao · 2022 [cited by examiner]
US 20220374208A1 · Allamanis · 2022 [cited by examiner]
Chakraborty et al., “CODIT: Code Editing with Tree-Based Neural Models,” IEEE, 2019, 14pg. (Year: 2019). [cited by examiner]
Liu et al., “A Self-Attentional Neural Architecture for Code Completion with Multi-Task Learning,” ACM, 2020, 11pg. (Year: 2020). [cited by examiner]
Lu et al., “CodeXGLUE: A Machine Learning Benchmark Dataset for Code Understanding and Generation,” arXiv, 2021, 14pg. (Year: 2021). [cited by examiner]
Allamanis et al., “Mining Idioms from Source Code”, arXiv:1404.0417v3, Jun. 2, 2014, 13 pages. [cited by applicant]
Guo et al., “Learning to Complete Code with Sketches”, arXiv:2106.10158v2, Jan. 23, 2022, 23 pages. [cited by applicant]
Sontag et al., “Introduction to Dual Decomposition for Inference”, 37 pages. [cited by applicant]
Visual Studio Code, “Snippets in Visual Studio Code”, https://code.visualstudio.com/docs/editor/userdefinedsnippets, retrieved on Aug. 11, 2022, 7 pages. [cited by applicant]
Cited By (1)
US 12,592,158