IP Library › Granted Patent US 12,682,229
Granted Patent B2
US 12,682,229 · App. 17/215,021 · Granted Jul 14, 2026

Machine learning in a continuous integration and deployment environment for compliance and security of infrastructure as code

Inventor: Fady Copty (Nazareth, IL)
Assignee: International Business Machines Corporation
G06N3/08G06F8/77G06F21/577G06N3/04G06F2221/033
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,682,229
App. No.
17/215,021
Filed
Mar 29, 2021
Granted
Jul 14, 2026
Kind
B2
Examiner
MANG, VAN C
Art Unit
2126
USPC
706/20
Abstract

In an approach for policy security shifting left of infrastructure as code compliance, a processor trains a neural network model to classify a code per policy and provide a policy vector score for the code associated with one or more policies. A processor enables the neural network model to scan and score a new code during a continuous integration and continuous deployment pipeline. A processor outputs a scanned score of the new code to a user. A processor retrains the neural network model by capturing a continuous integration and continuous deployment change and run-time compliance posture that occurs as a response by the user.

Claims (50)

1 . A computer-implemented method comprising:

training, by one or more processors using training data, a neural network model to classify a code, included in the training data, per policy of a plurality of policies and provide a vector of scores for the code,

wherein each score of the vector of scores indicates a level of possibility of complying with a respective policy of the plurality of policies,

wherein the training the neural network model includes adjusting weights of interconnections between intermediate layers of the neural network model,

wherein the training data includes one or more modifications of the code, and

wherein the one or more modifications of the code are selected from a group consisting of an addition of random keywords to the code, a scrambling of the code, blocks of the code re-ordered, lines within a block of the code re-ordered, tokens of the code renamed, and combinations thereof, and

wherein the training data is labeled using a vector of labels indicating a compliance check result of deploying the code and a cost-estimate or a performance-estimate of deploying the code;

generating a blocking alert based at least in part on a determination of a magnitude of the vector of scores for the code; and

retraining, by one or more processors, the neural network model by capturing a continuous integration and continuous deployment change and run-time compliance posture that occurs as a response by a user, and

wherein the neural network model is a convolutional neural network model, a recurrent neural network model, a transformer neural network model, or a combination of the convolutional neural network model and the recurrent neural network model.

2 . The computer-implemented method of claim 1 , further comprising:

enabling the neural network model,

wherein the enabling includes scanning a new code while the user types the new code.

3 . The computer-implemented method of claim 1 ,

wherein the code is an infrastructure as code.

4 . The computer-implemented method of claim 1 ,

wherein the plurality of policies comprises an industry-specific methodology requirement and a cost estimate check.

5 . A non-transitory computer-readable medium storing a set of instructions for wireless communication, the set of instructions comprising:

one or more instructions that, when executed by one or more processors of a device, cause the device to:

train, using training data, a neural network model to classify a code, included in the training data, per policy of a plurality of policies and provide a vector of scores for the code, wherein each score of the vector of scores indicates a level of possibility of complying with a respective policy of the plurality of policies, wherein the one or more instructions, to cause the device to train the neural network model, cause the device to adjust weights of interconnections between intermediate layers of the neural network model, wherein the training data includes one or more modifications of the code, and wherein the one or more modifications of the code included in the training data are selected from a group consisting of an addition of random keywords to the code, a scrambling of the code, blocks of the code re-ordered, lines within a block of the code re-ordered, tokens of the code renamed, and combinations thereof, and wherein the training data is labeled using a vector of labels indicating a compliance check result of deploying the code and a cost-estimate or a performance-estimate of deploying the code;

generate a blocking alert based at least in part on a determination of a magnitude of the vector of scores for the code; and

retrain the neural network model by capturing a continuous integration and continuous deployment change and run-time compliance posture that occurs as a response by a user, and

wherein the neural network model is a convolutional neural network model, a recurrent neural network model, a transformer neural network model, or a combination of the convolutional neural network model and the recurrent neural network model.

6 . The non-transitory computer-readable medium of claim 5 ,

wherein the one or more instructions cause the device to:

enable the neural network model which includes program instructions to scan a new code while the user types the new code.

7 . The non-transitory computer-readable medium of claim 5 ,

wherein the code is an infrastructure as code.

8 . The non-transitory computer-readable medium of claim 5 ,

wherein the plurality of policies comprises an industry-specific methodology requirement and a cost estimate check.

9 . A computer system, comprising:

one or more processors; and

one or more memory devices coupled to the one or more processors, wherein the one or more processors are configured to:

train a neural network model using training data to classify a code, included in the training data, per policy of a plurality of policies and provide a vector of scores for the code, wherein each score of the vector of scores indicates a level of possibility of complying with a respective policy of the plurality of policies, wherein the one or more processors, to train the neural network model, are configured to adjust weights of interconnections between intermediate layers of the neural network model, wherein the training data includes one or more modifications of the code, and wherein the one or more modifications of the code are selected from a group consisting of an addition of random keywords to the code, a scrambling of the code, blocks of the code re-ordered, lines within a block of the code re-ordered, tokens of the code renamed, and combinations thereof, and wherein the training data is labeled using a vector of labels indicating a compliance check result of deploying the code and a cost-estimate or a performance-estimate of deploying the code;

generate a blocking alert based at least in part on a determination of a magnitude of the vector of scores for the code; and

retrain the neural network model by capturing a continuous integration and continuous deployment change and run-time compliance posture that occurs as a response by a user, and

wherein the neural network model is a convolutional neural network model, a recurrent neural network model, a transformer neural network model, or a combination of the convolutional neural network model and the recurrent neural network model.

10 . The computer system of claim 9 ,

wherein the one or more processors are configured to:

enable the neural network model which includes program instructions to scan a new code while the user types the new code.

11 . The computer system of claim 9 ,

wherein the code is an infrastructure as code.

12 . The computer system of claim 9 ,

wherein the plurality of policies comprises an industry-specific methodology requirement and a cost estimate check.

13 . The computer-implemented method of claim 1 ,

wherein the training data comprises a vector of labels that present a result of deploying the code and running post-deployment checks on a deployed account.

14 . The computer-implemented method of claim 1 , further comprising:

labeling another code by deploying the code to a cloud account and running a configuration scan on the code.

15 . The computer-implemented method of claim 1 ,

wherein each score in the vector of scores has a value that increases as the code has a higher probability of complying with the respective policy.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Mar 29, 2021
From: COPTY, FADY
To: INTERNATIONAL BUSINESS MACHINES CORPORATION
Reel/Frame 055746/0946 →
Continuity (1)
Related Publication 20220309337A1 · Sep 29, 2022
References Cited (32)
US 10565093B1 · Herrin · 2020 [cited by examiner]
US 10832150B2 · Barry · 2020 [cited by examiner]
US 10872029B1 · Bawcom · 2020 [cited by examiner]
US 20190026085A1 · Bijani et al. · 2019 [cited by applicant]
US 20190109872A1 · Dhakshinamoorthy · 2019 [cited by examiner]
US 20190171966A1 · Rangasamy · 2019 [cited by applicant]
US 20190318366A1 · Carranza · 2019 [cited by examiner]
US 20190370473A1 · Matrosov · 2019 [cited by examiner]
US 20200341876A1 · Gandhi · 2020 [cited by examiner]
US 20200379879A1 · Plotnik et al. · 2020 [cited by applicant]
US 20210149788A1 · Downie · 2021 [cited by examiner]
US 20210319174A1 · Raj · 2021 [cited by examiner]
US 20220156641A1 · Fujii · 2022 [cited by examiner]
Fontana et al., “Code smell severity classification using machine learning techniques”, 2017, Knowledge-Based Systems, vol. 128, pp. 43-58 (Year: 2017). [cited by examiner]
Veggalam, “IFuzzer: An Evolutionary Interpreter Fuzzer Using Genetic Programming”, 2016, Computer Security—ESORICS 2016, vol. 2016, pp. 581-601 (Year: 2016). [cited by examiner]
Nembhard et al., “A hybrid approach to improving program security”, 2017, 2017 IEEE Symposium Series on Computational Intelligence (SSCI), vol. 2017, pp. 1-8 (Year: 2017). [cited by examiner]
Chen et al., “An Approach to Identifying Error Patterns for Infrastructure as Code”, 2018, Proceedings—29th IEEE International Symposium on Software Reliability Engineering Workshops, ISSREW 2018, vol. 29 (2018), pp. 12… [cited by examiner]
Logicworks, “Logicworks Launches Infrastructure-as-Code Platform Pulse to Deliver Security & Automated Best Practices”, 2018, retrieved from https://www.logicworks.com/blog/2018/02/logicworks-launches-infrastructure-cod… [cited by examiner]
Logicworks, “Security Risks in Public Cloud: The Tightrope Walk of Digital Transformation”, 2020, retrieved from https://www.logicworks.com/blog/2020/02/security-risks-in-public-cloud/ on Mar. 8, 2024 (Year: 2020). [cited by examiner]
Chernis et al., “Machine Learning Methods for Software Vulnerability Detection”, 2018, IWSPA '18: Proceedings of the Fourth ACM International Workshop on Security and Privacy Analytics, vol. 4 (2018), pp. 31-39 (Year: 2… [cited by examiner]
Mosolygo et al., “Towards a Prototype Based Explainable JavaScript Vulnerability Prediction Model”, Mar. 27, 2021, 2021 International Conference on Code Quality (ICCQ), vol. 2021, pp. 15-25 (Year: 2021). [cited by examiner]
Mirakhorli et al., “Detecting, Tracing, and Monitoring Architectural Tactics in Code”, 2016, IEEE Transactions on Software Engineering, vol. 42 No. 3, pp. 205-220 (Year: 2016). [cited by examiner]
Lemieux et al., “FairFuzz: a targeted mutation strategy for increasing greybox fuzz testing coverage”, 2018, ASE '18: Proceedings of the 33rd ACM/IEEE International Conference on Automated Software Engineering, vol. 33 … [cited by examiner]
Gao et al., “Fuzz Testing based Data Augmentation to Improve Robustness of Deep Neural Networks”, 2020, 2020 IEEE/ACM 42nd International Conference on Software Engineering (ICSE), vol. 42 (2020), pp. 1147-1158 (Year: 20… [cited by examiner]
Shahin et al., “Beyond Continuous Delivery: An Empirical Investigation of Continuous Deployment Challenges”, 2017, 2017 ACM/IEEE International Symposium on Empirical Software Engineering and Measurement (ESEM), vol. 201… [cited by examiner]
Atlassian, “What is Continuous Integration?”, Feb. 22, 2021, retrieved from https://web.archive.org/web/20210222160537/https://www.atlassian.com/continuous-delivery/continuous-integration on Nov. 1, 2024 (Year: 2021). [cited by examiner]
Andrzejak et al., “Detection of Memory Leaks in C/C++ Code via Machine Learning”, 2017, 2017 IEEE International Symposium on Software Reliability Engineering Workshops (ISSREW), vol. 2017, pp. 252-258 (Year: 2017). [cited by examiner]
Hulette, “Predicting Fault Locations from Failures Using a Machine Learning Classifier”, 2007, University of California, San Diego (Year: 2007). [cited by examiner]
Gupta et al., “DeepFix: Fixing Common C Language Errors by Deep Learning”, 2017, Proceedings of the Thirty-First AAAI Conference on Artificial Intelligence, vol. 31, pp. 1345-1351 (Year: 2017). [cited by examiner]
Lenarduzzi et al., “Towards Surgically-Precise Technical Debt Estimation: Early Results and Research Roadmap”, 2019, Proceedings of the 3rd ACM SIGSOFT International Workshop on Machine Learning Techniques for Software … [cited by examiner]
“Shifting Cloud Security Left with Infrastructure as Code”, DivvyCloud, Apr. 2020, 11 pages, <www.divvycloud.com/get-started>. [cited by applicant]
Mell et al., “The NIST Definition of Cloud Computing”, Recommendations of the National Institute of Standards and Technology, Special Publication 800-145, Sep. 2011, 7 pages. [cited by applicant]