IP Library Granted Patent US 11,169,786
Granted Patent B2
US 11,169,786 · App. 16/781,344 · Granted Nov 9, 2021

Generating and using joint representations of source code

Inventors: Rohan Badlani (Stanford, CA); Owen Lewis (Stanford, CA); Georgios Evangelopoulos (Venice, CA); Olivia Hatalsky (San Jose, CA); Bin Ni (Fremont, CA)
Assignee: X DEVELOPMENT LLC
G06F8/42G06F8/38G06F8/73G06N3/0454G06N20/00
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,169,786
App. No.
16/781,344
Granted
Nov 9, 2021
Kind
B2
Abstract

Implementations are described herein for generating embeddings of source code using both the language and graph domains, and leveraging combinations of these semantically-rich and structurally-informative embeddings for various purposes. In various implementations, tokens of a source code snippet may be applied as input across a sequence-processing machine learning model to generate a plurality of token embeddings. A graph may also be generated based on the source code snippet. A joint representation may be generated based on the graph and the incorporated token embeddings. The joint representation generated from the source code snippet may be compared to one or more other joint representations generated from one or more other source code snippets to make a determination about the source code snippet.

Claims (30)

1. A method implemented using one or more processors, comprising:

applying tokens of a source code snippet as input across a sequence-processing machine learning model to generate a plurality of token embeddings;

generating a graph based on the source code snippet;

augmenting the graph by infusing nodes of the graph that correspond to the tokens of the source code snippet with corresponding token embeddings of the generated plurality of token embeddings;

processing the augmented graph using a graph neural network to generate a joint representation of the source code snippet, the joint representation comprising an aggregate node embedding based on the augmented graph infused with the corresponding token embeddings, wherein the aggregate node embedding is obtained by further infusing the nodes of the augmented graph with data indicative of the corresponding token embeddings from neighboring nodes for a number of iterations by propagating the date from each node along one or more edges to one or more immediate neighboring nodes;

calculating one or more distances in latent space between the joint representation generated from the source code snippet and one or more other joint representations generated from one or more other source code snippets to make a determination about the source code snippet; and

automatically generating a summary, review, or comment associated with the determination about the source code snippet.

2. The method of claim 1 , wherein the joint representation comprises a mean over the aggregate node embeddings generated from the augmented graph.

3. The method of claim 1 , wherein the calculating is performed as part of a search for other source code similar to the source code snippet.

4. The method of claim 1 , wherein the determination about the source code snippet comprises a source code quality score.

5. The method of claim 1 , wherein the graph comprises an abstract syntax tree.

6. The method of claim 1 , wherein the sequence-processing machine learning model comprises a neural network that includes a self attention mechanism.

7. The method of claim 6 , wherein the neural network comprises a transformer network.

8. A system comprising one or more processors and memory storing instructions that, in response to execution of the instructions by the one or more processors, cause the one or more processors to:

apply tokens of a source code snippet as input across a sequence-processing machine learning model to generate a plurality of token embeddings;

generate a graph based on the source code snippet;

augment the graph by infusing nodes of the graph that correspond to the tokens of the source code snippet with corresponding token embeddings of the generated plurality of token embeddings;

process the augmented graph using a graph neural network to generate a joint representation of the source code snippet, the joint representation comprising an aggregate node embedding based on the augmented graph infused with the corresponding token embeddings, wherein the aggregate node embedding is obtained by further infusing the nodes of the augmented graph with data indicative of the corresponding token embeddings from neighboring nodes for a number of iterations by propagating the data from each node along one or more edges to one or more immediate neighboring nodes;

calculate one or more distances in latent space between the joint representation generated from the source code snippet and one or more other joint representations generated from one or more other source code snippets to make a determination about the source code snippet; and

automatically generate a summary, review, or comment associated with the determination about the source code snippet.

9. The system of claim 8 , wherein the determination comprises a similarity between the source code snippet and other source code.

10. The system of claim 8 , wherein the determination about the source code snippet comprises a source code quality score.

11. The system of claim 8 , wherein the sequence-processing machine learning model comprises a transformer network.

12. At least one non-transitory computer-readable medium comprising instructions that, in response to execution of the instructions by one or more processors, cause the one or more processors to:

apply tokens of a source code snippet as input across a sequence-processing machine learning model to generate a plurality of token embeddings;

generate a graph based on the source code snippet;

augment the graph by infusing nodes of the graph that correspond to the tokens of the source code snippet with corresponding token embeddings of the generated plurality of token embeddings;

process the augmented graph using a graph neural network to generate a joint representation of the source code snippet, the joint representation comprising an aggregate node embedding based on the augmented graph infused with the corresponding token embeddings, wherein the aggregate node embedding is obtained by further infusing the nodes of the augmented graph with data indicative of the corresponding token embeddings from neighboring nodes for a number of iterations by propagating the data from each node along one or more edges to one or more immediate neighboring nodes;

calculate one or more distances in latent space between the joint representation generated from the source code snippet and one or more other joint representations generated from one or more other source code snippets to make a determination about the source code snippet; and

automatically generate a summary, review, or comment associated with the determination about the source code snippet.

Assignments (2)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Feb 2, 2023
From: X DEVELOPMENT LLC
To: GOOGLE LLC
Reel/Frame 062572/0565 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Feb 4, 2020
From: BADLANI, ROHAN; LEWIS, OWEN; EVANGELOPOULOS, GEORGIOS; HATALSKY, OLIVIA; NI, BIN
To: X DEVELOPMENT LLC
Reel/Frame 051713/0423 →
Continuity (1)
Related Publication 20210240453A1 · Aug 5, 2021
Cited By (3)
US 12,217,033 US 12,461,723 US 12,548,315