IP Library › Granted Patent US 10,452,367
Granted Patent B2
US 10,452,367 · App. 15/942,290 · Granted Oct 22, 2019

Variable analysis using code context

Inventors: Miltiadis Allamanis (Cambridge, GB); Marc Manuel Johannes Brockschmidt (Cambridge, GB)
Assignee: Microsoft Technology Licensing, LLC
G06F8/433G06F8/436G06N20/00
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 10,452,367
App. No.
15/942,290
Granted
Oct 22, 2019
Kind
B2
Abstract

Improving how a codebase is developed by analyzing the variables in the codebase's source code. Learned characteristics of a codebase are derived by obtaining context for some of the source code's variables. This context represents semantics and/or patterns associated with those variables. Once the learned characteristics are derived, they are then modified, or rather tuned, by incorporating context from second source code. Particular context for a particular variable used within the second source code is then obtained. This particular context represents semantics and/or patterns associated with the particular variable. This particular context is then analyzed using the learned characteristics to generate zero, one or more anticipated variables. Later, a notification regarding these anticipated variables is displayed. In some situations, conducting the analysis is a part of a variable renaming analysis while in other situation the analysis is a part of a variable misuse analysis.

Claims (43)

1. A computer system comprising:

one or more processors; and

one or more computer-readable hardware storage devices having stored thereon computer-executable instructions that are structured to be executable by the one or more processors to thereby cause the computer system to perform a method for analyzing variables in a codebase, the method comprising:

for each of at least some variables in a first codebase that includes first source code, obtaining a respective context that at least represents one or more semantics and/or patterns associated with each respective variable;

using those obtained contexts to derive a set of learned characteristics of the first codebase;

for second source code that differs from the first source code, using context obtained from the different second source code to modify the set of learned characteristics of the first codebase to generate an updated set of learned characteristics;

obtaining particular context for a particular variable that is used within the second source code, the particular variable's context at least representing one or more semantics and/or patterns associated with the particular variable;

analyzing the particular variable's context using the updated set of learned characteristics, including performing a variable renaming analysis;

based on determining, using the updated set of learned characteristics, that the particular variable's context indicates that the particular variable has a likelihood of being incorrect, generating one or more anticipated variables based on the updated set of learned characteristics and the particular variable's context, including generating at least one variable that is not currently included in the second source code,

wherein each of the one or more anticipated variables are included based on a determination that the respective anticipated variable should be used to replace the particular variable within the second source code; and

causing a notification to be displayed, the notification including at least one of the one or more anticipated variables.

2. The computer system of claim 1 , wherein analyzing the particular variable's context using the learned characteristics constitutes a variable misuse analysis, and wherein generating the one or more anticipated variables includes analyzing only variables that already exist in the second source code and selecting one or more of the already existing variables for inclusion among the one or more anticipated variables.

3. The computer system of claim 1 wherein the particular context for the particular variable is extracted from a graph of elements that corresponds to the particular variable.

4. The computer system of claim 3 , wherein the elements include one or more tokens, syntax tree nodes, declarations, names, or operators.

5. The computer system of claim 1 , wherein modifying the learned characteristics using the context obtained from the second source code is a part of a tuning phase that occurs to refine the learned characteristics to be specifically operable on the second source code.

6. The computer system of claim 1 , wherein generating the one or more anticipated variables includes generating a probability distribution for each of the one or more anticipated variables.

7. The computer system of claim 6 , wherein the probability distribution for the one or more anticipated variables represents a likelihood that each of the one or more anticipated variables should have been used in place of the particular variable.

8. The computer system of claim 1 , wherein analyzing the particular variable's context using the learned characteristics is performed as a part of either a variable renaming analysis or a variable misuse analysis.

9. A method for analyzing variables in a codebase, the method being performed by a computer system that includes one or more processors, the method comprising:

for each of at least some variables in a first codebase that includes first source code, obtaining a respective context that at least represents one or more semantics and/or patterns associated with each respective variable;

using those obtained contexts to derive a set of learned characteristics of the first codebase;

for second source code that differs from the first source code, using context obtained from the different second source code to modify the set of learned characteristics of the first codebase to generate an updated set of learned characteristics of the first codebase;

obtaining particular context for a particular variable that is used within the second source code, the particular variable's context at least representing one or more semantics and/or patterns associated with the particular variable;

analyzing the particular variable's context using the updated set of learned characteristics of the first code base, including performing a variable renaming analysis;

based on determining, using the updated set of learned characteristics characteristics of the first code base, that the particular variable's context indicates that the particular variable has a likelihood of being incorrect, generating one or more anticipated variables based on the updated set of learned characteristics for the first codebase and the particular variable's context, including generating at least one variable that is not currently included in the second source code,

wherein each of the one or more anticipated variables are included based on a determination that the respective anticipated variable should be used to replace the particular variable within the second source code; and

causing a notification to be displayed, the notification including at least one of the one or more anticipated variables.

10. The method of claim 9 , wherein analyzing the particular variable's context using the learned characteristics constitutes a variable misuse analysis, and wherein generating the one or more anticipated variables includes analyzing only variables that already exist in the second source code and selecting one or more of the already existing variables for inclusion among the one or more anticipated variables.

11. The method of claim 9 wherein the particular context for the particular variable includes a selected number of usage elements that are associated with the particular variable.

12. The method of claim 9 , wherein the particular context for the particular variable is extracted from a graph of elements that corresponds to the particular variable.

13. The method of claim 12 , wherein the elements include one or more tokens, syntax nodes, declarations, names, or operators.

14. The method of claim 9 , wherein generating the one or more anticipated variables includes generating a probability distribution for each of the one or more anticipated variables.

15. The method of claim 14 , wherein the probability distribution for the one or more anticipated variables represents a likelihood that each of the one or more anticipated variables should have been used in place of the particular variable.

16. The method of claim 9 , wherein analyzing the particular variable's context using the learned characteristics is performed as a part of either a variable renaming analysis or a variable misuse analysis.

17. One or more hardware storage devices having stored thereon computer-executable instructions that are structured to be executable by one or more processors of a computer system to thereby cause the computer system to perform a method for analyzing variables in a codebase, the method comprising:

for each of at least some variables in a first codebase that includes first source code, obtaining a respective context that at least represents one or more semantics and/or patterns associated with each respective variable;

using those obtained contexts to derive a set of learned characteristics of the first codebase;

for second source code that differs from the first source code, using context obtained from the different second source code to modify the set of learned characteristics of the first codebase to generate an updated set of learned characteristics of the first codebase;

obtaining particular context for a particular variable that is used within the second source code, the particular variable's context at least representing one or more semantics and/or patterns associated with the particular variable;

analyzing the particular variable's context using the updated set of learned characteristics of the first code base, including performing a variable renaming analysis;

based on determining, using the updated set of learned characteristics characteristics of the first code base, that the particular variable's context indicates that the particular variable has a likelihood of being incorrect, generating one or more anticipated variables based on the updated set of learned characteristics for the first codebase and the particular variable's context, including generating at least one variable that is not currently included in the second source code,

wherein each of the one or more anticipated variables are included based on a determination that the respective anticipated variable should be used to replace the particular variable within the second source code; and

causing a notification to be displayed, the notification including at least one of the one or more anticipated variables.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Apr 3, 2018
From: ALLAMANIS, MILTIADIS; BROCKSCHMIDT, MARC MANUEL JOHANNES
To: MICROSOFT TECHNOLOGY LICENSING, LLC
Reel/Frame 045428/0195 →
Continuity (2)
Provisional Application 62627604 · Feb 7, 2018
Related Publication 20190243622A1 · Aug 8, 2019
Cited By (1)
US 12,505,710