Method and system for code security evaluation and rule generation using graph-based analysis
Aspects of the disclosure relate to evaluating source code repositories to generate secure coding rules. An application implemented by a computer system receives a list of repositories for multiple software agents, performs code analysis using threat modeling techniques to identify flaws, and generates evaluation reports. Data from the reports is stored in a graph database as nodes representing repositories, flaws, reports, and contextual information, with portions optionally converted into vector embeddings. The graph database is aggregated to identify flaws recurring across repositories. Based on these aggregated results, the application generates secure coding rules, which are provided to downstream systems. The downstream systems may apply the secure coding rules to influence generation of software code or configuration of agent blueprints, thereby enabling consistent remediation and improved software security across codebases.
1 . A method, comprising:
receiving, at an application implemented by a computer system, a list of code repositories for a plurality of software agents;
for each respective code repository in the list:
performing, by the application, a code analysis of the respective code repository using one or more threat modeling techniques to identify one or more flaws;
generating, by the application, an evaluation report based on the one or more flaws; and
storing, by the application, data from the evaluation report in a graph database as a plurality of nodes, wherein:
the plurality of nodes represent the respective code repository, the evaluation report, and contextual information associated with the respective code repository,
the plurality of nodes are connected by edges that link the one or more flaws to at least one or more corresponding code improvements or one or more secure coding rules, and
storing comprises converting at least a portion of the evaluation report into vector embeddings;
aggregating, by the application, the plurality of nodes in the graph database to identify one or more flaws that are present in more than one of the code repositories in the list;
generating, by the application, one or more secure coding rules based on the one or more flaws; and
providing, by the application, the one or more secure coding rules to a downstream system, wherein the downstream system comprises a blueprint management system and is configured to apply the one or more secure coding rules to influence configuration of agent blueprints.
2 . The method of claim 1 , wherein performing the code analysis comprises applying a threat modeling framework to the respective code repository.
3 . The method of claim 2 , wherein the threat modeling framework comprises a Multi-Agent Environment, Security, Threat, Risk, and Outcome (MAESTRO) threat modeling framework.
4 . The method of claim 1 , wherein the plurality of nodes further represent the one or more flaws and one or more code improvements identified in the evaluation report.
5 . The method of claim 1 , wherein the contextual information comprises an architecture diagram associated with the respective code repository.
6 . The method of claim 5 , wherein the architecture diagram is processed using an image recognition model to extract one or more features for inclusion in the graph database.
7 . The method of claim 1 , wherein generating the one or more secure coding rules comprises expressing the one or more secure coding rules in a structured format organized by programming language.
8 . The method of claim 1 , wherein the one or more secure coding rules comprise a starter-kit patch including one or more secure template source files.
9 . The method of claim 1 , wherein the one or more secure coding rules comprise one or more prompt modifications for a language model used to generate code.
10 . The method of claim 1 , wherein the one or more flaws comprise at least one of a code-level vulnerability, a configuration-level vulnerability, an architectural weakness associated with an agent workflow, or a combination thereof.
11 . A processing system, comprising:
one or more memories comprising computer-executable instructions; and
one or more processors configured to execute the computer-executable instructions and cause the processing system to:
receive, at an application implemented by a computer system, a list of code repositories for a plurality of software agents;
for each respective code repository in the list:
perform, by the application, a code analysis of the respective code repository using one or more threat modeling techniques to identify one or more flaws;
generate, by the application, an evaluation report based on the one or more flaws; and
store, by the application, data from the evaluation report in a graph database as a plurality of nodes, wherein;
the plurality of nodes represent the respective code repository, the evaluation report, and contextual information associated with the respective code repository,
the plurality of nodes are connected by edges that link the one or more flaws to at least one or more corresponding code improvements or one or more secure coding rules, and
storing comprises converting at least a portion of the evaluation report into vector embeddings;
aggregate, by the application, the plurality of nodes in the graph database to identify one or more flaws that are present in more than one of the code repositories in the list;
generate, by the application, one or more secure coding rules based on the one or more flaws; and
provide, by the application, the one or more secure coding rules to a downstream system, wherein the downstream system comprises a blueprint management system and is configured to apply the one or more secure coding rules to influence configuration of agent blueprints.
12 . The processing system of claim 11 , wherein performing the code analysis comprises applying a threat modeling framework to the respective code repository.
13 . The processing system of claim 12 , wherein the threat modeling framework comprises a Multi-Agent Environment, Security, Threat, Risk, and Outcome (MAESTRO) threat modeling framework.
14 . The processing system of claim 11 , wherein the plurality of nodes further represent the one or more flaws and one or more code improvements identified in the evaluation report.
15 . The processing system of claim 11 , wherein the contextual information comprises an architecture diagram associated with the respective code repository.
16 . The processing system of claim 15 , wherein the architecture diagram is processed using an image recognition model to extract one or more features for inclusion in the graph database.
17 . The processing system of claim 11 , wherein generating the one or more secure coding rules comprises expressing the one or more secure coding rules in a structured format organized by programming language.
18 . A method, comprising:
accessing, by a computer-implemented application, evaluation data associated with a plurality of code repositories, the evaluation data comprising flaws identified by applying one or more threat modeling techniques to source code of the plurality of code repositories;
storing, by the computer-implemented application, the evaluation data in a graph database as nodes representing the code repositories, the flaws, and contextual information describing the code repositories, wherein the nodes are connected by edges that link the flaws to one or more secure coding rules;
aggregating, by the computer-implemented application, the nodes in the graph database to detect recurring flaw patterns across the plurality of code repositories;
normalizing, by the computer-implemented application, the recurring flaw patterns into common flaw entries;
generating, by the computer-implemented application, one or more secure coding rules based on the recurring flaw patterns and the common flaw entries; and
outputting, by the computer-implemented application, the one or more secure coding rules to a downstream system, wherein the downstream system comprises a blueprint management system configured to influence configuration of agent blueprints.
19 . The method of claim 1 , wherein:
aggregating, by the application, the nodes in the graph database to detect one or more flaws comprises aggregating, by the application, the nodes in the graph database to detect recurring flaw patterns across the list of code repositories,
the method further comprises normalizing, by the application, the recurring flaw patterns into common flaw entries, and
generating, by the application, the one or more secure coding rules based on the one or more flaws comprises generating, by the application, the one or more secure coding rules based on the recurring flaw patterns and the common flaw entries.
20 . The method of claim 1 , wherein the one or more secure coding rules comprise one or more configuration policies for use by the blueprint management system.