IP Library Granted Patent US 11,429,352
Granted Patent B2
US 11,429,352 · App. 16/917,967 · Granted Aug 30, 2022

Building pre-trained contextual embeddings for programming languages using specialized vocabulary

Inventors: Saurabh Pujar (White Plains, NY); Luca Buratti (White Plains, NY); Alessandro Morari (New York, NY); Jim Alain Laredo (Katonah, NY); Alfio Massimiliano Gliozzo (Brooklyn, NY); Gaetano Rossiello (Brooklyn, NY)
Assignee: International Business Machines Corporation
G06F8/31G06F8/40G06F9/445G06F40/40G06N20/00
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,429,352
App. No.
16/917,967
Granted
Aug 30, 2022
Kind
B2
Abstract

A method, a computer system, and a computer program product for building pre-trained contextual embeddings is provided. Embodiments of the present invention may include collecting programming code. Embodiments of the present invention may include loading and preparing the programming code using a specialized programming language keywords-based vocabulary. Embodiments of the present invention may include creating contextual embeddings for the programming code. Embodiments of the present invention may include storing the contextual embeddings.

Claims (43)

1. A method for building pre-trained contextual embeddings, the method comprising:

collecting programming code;

loading and preparing the programming code using a specialized programming language keywords-based vocabulary;

creating contextual embeddings for the programming code using the specialized programming language keywords-based vocabulary;

determining a context for the programming code based on the contextual embeddings, wherein the contextual embeddings are associated with one or more vectors;

using natural language processing (NLP) to perform language modeling to initialize the one or more vectors based off the words in the programming code;

extracting one or more tokens in the programming code to identify word contexts in the programming code; and

storing the contextual embeddings,

wherein the contextual embeddings are stored as pre-trained contextual embeddings that are built to use pre-trained models, fine tuning models or machine learning models.

2. The method of claim 1 , wherein the contextual embeddings are stored as pre-trained contextual embeddings that are built to use with programming languages in conjunction with configuration files and vocabulary files.

3. The method of claim 1 , wherein the programming code consists of source code.

4. The method of claim 1 , wherein the loading and preparing the programming code includes loading the programming code into a data loader and converting the programming code into a required form.

5. The method of claim 1 , wherein the creating contextual embeddings includes creating character tokens for the programming code using byte pair encoding (BPE).

6. The method of claim 1 , wherein the contextual embeddings are pre-trained for the programming code.

7. A computer system for building pre-trained contextual embeddings, comprising:

one or more processors, one or more computer-readable memories, one or more computer-readable tangible storage media, and program instructions stored on at least one of the one or more computer-readable tangible storage media for execution by at least one of the one or more processors via at least one of the one or more computer-readable memories, wherein the computer system is capable of performing a method comprising:

collecting programming code;

loading and preparing the programming code using a specialized programming language keywords-based vocabulary;

creating contextual embeddings for the programming code using the specialized programming language keywords-based vocabulary;

determining a context for the programming code based on the contextual embeddings, wherein the contextual embeddings are associated with one or more vectors;

using natural language processing (NLP) to perform language modeling to initialize the one or more vectors based off the words in the programming code;

extracting one or more tokens in the programming code to identify word contexts in the programming code; and

storing the contextual embeddings,

wherein the contextual embeddings are stored as pre-trained contextual embeddings that are built to use pre-trained models, fine tuning models or machine learning models.

8. The computer system of claim 7 , wherein the contextual embeddings are stored as pre-trained contextual embeddings that are built to use with programming languages in conjunction with configuration files and vocabulary files.

9. The computer system of claim 7 , wherein the programming code consists of source code.

10. The computer system of claim 7 , wherein the loading and preparing the programming code includes loading the programming code into a data loader and converting the programming code into a required form.

11. The computer system of claim 7 , wherein the creating contextual embeddings includes creating character tokens for the programming code using byte pair encoding (BPE).

12. The computer system of claim 7 , wherein the contextual embeddings are pre-trained for the programming code.

13. A computer program product for building pre-trained contextual embeddings, comprising:

one or more computer-readable tangible storage media and program instructions stored on at least one of the one or more computer-readable tangible storage media, the program instructions executable by a processor to cause the processor to perform a method comprising:

collecting programming code;

loading and preparing the programming code using a specialized programming language keywords-based vocabulary;

creating contextual embeddings for the programming code using the specialized programming language keywords-based vocabulary;

determining a context for the programming code based on the contextual embeddings, wherein the contextual embeddings are associated with one or more vectors;

using natural language processing (NLP) to perform language modeling to initialize the one or more vectors based off the words in the programming code;

extracting one or more tokens in the programming code to identify word contexts in the programming code; and

storing the contextual embeddings,

wherein the contextual embeddings are stored as pre-trained contextual embeddings that are built to use pre-trained models, fine tuning models or machine learning models.

14. The computer program product of claim 13 , wherein the contextual embeddings are stored as pre-trained contextual embeddings that are built to use with programming languages in conjunction with configuration files and vocabulary files.

15. The computer program product of claim 13 , wherein the programming code consists of source code.

16. The computer program product of claim 13 , wherein the loading and preparing the programming code includes loading the programming code into a data loader and converting the programming code into a required form.

17. The computer program product of claim 13 , wherein the creating contextual embeddings includes creating character tokens for the programming code using byte pair encoding (BPE).

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jul 1, 2020
From: PUJAR, SAURABH; BURATTI, LUCA; MORARI, ALESSANDRO; LAREDO, JIM ALAIN; GLIOZZO, ALFIO MASSIMILIANO; ROSSIELLO, GAETANO
To: INTERNATIONAL BUSINESS MACHINES CORPORATION
Reel/Frame 053096/0046 →
Continuity (1)
Related Publication 20220004365A1 · Jan 6, 2022
Cited By (3)
US 12,307,247 US 12,436,745 US 12,664,157