IP Library › Granted Patent US 12,536,290
Granted Patent B2
US 12,536,290 · App. 18/731,845 · Granted Jan 27, 2026

Detecting artificial intelligence generated computer code

Inventors: Wei Cheng (Princeton Junction, NJ); Xianjun Yang (Santa Barbara, CA); Haifeng Chen (West Windsor, NJ)
Assignee: NEC Corporation
G06F21/57G06F11/3624G06F2221/033
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,536,290
App. No.
18/731,845
Granted
Jan 27, 2026
Kind
B2
Abstract

Systems and methods for detecting artificial intelligence (AI) generated computer code. Lines of code can be masked from a candidate code to obtain perturbed codes. Missing code can be generated from the perturbed codes by employing an AI code generator model to obtain machine-filled codes. Probabilities of the candidate code probability and the machine-filled codes as AI-generated can be predicted by employing a surrogate model. The candidate code can be distinguished as AI-generated by comparing the probabilities against a detection threshold to obtain detection results.

Claims (34)

1 . A computer-implemented method for detecting artificial intelligence (AI) generated computer code, comprising:

masking lines of code from a candidate code to obtain perturbed codes;

generating missing code from the perturbed codes by employing an AI code generator model to obtain machine-filled codes;

predicting probabilities of the candidate code and the machine-filled codes as AI-generated by employing a surrogate model; and

distinguishing the candidate code as AI-generated by comparing the probabilities against a detection threshold to obtain detection results.

2 . The computer-implemented method of claim 1 , further comprising flagging the candidate code as AI-generated to detect malicious code for a decision-making entity to perform an action.

3 . The computer-implemented method of claim 2 , wherein the action is securing a healthcare management system handling patient vital data by patching the flagged candidate code for potential security risks.

4 . The computer-implemented method of claim 1 , wherein predicting probabilities further comprises predicting the candidate code probability by predicting the probability of generating a remainder code given a prefix code.

5 . The computer-implemented method of claim 4 , wherein distinguishing the candidate code further comprises comparing n-gram divergence of a prefix code and a remainder code and n-gram divergence of prefix filled codes and remainder filled codes.

6 . The computer-implemented method of claim 1 , wherein predicting probabilities further comprises predicting machine-filled codes probabilities by predicting the probability of generating a remainder filled code given a prefix filled code.

7 . The computer-implemented method of claim 1 , wherein distinguishing the candidate code further comprises computing a difference between the candidate code probability and the machine-filled codes probabilities.

8 . A system for detecting artificial intelligence (AI) generated computer code, comprising:

a memory; and

one or more processor devices in communication with the memory configured to:

mask lines of code from a candidate code to obtain perturbed codes;

generate missing code from the perturbed codes by employing an AI code generator model to obtain machine-filled codes;

predict probabilities of the candidate code and the machine-filled codes as AI-generated by employing a surrogate model; and

distinguish the candidate code as AI-generated by comparing the probabilities against a detection threshold to obtain detection results.

9 . The system of claim 8 , further comprising the processor device flagging the candidate code as AI-generated to detect malicious code for a decision-making entity to perform an action.

10 . The system of claim 9 , wherein the processor device performs the action securing a healthcare management system handling patient vital data by patching the flagged candidate code for potential security risks.

11 . The system of claim 8 , wherein predicting probabilities by the processor device further comprises predicting the candidate code probability by predicting the probability of generating a remainder code given a prefix code.

12 . The system of claim 11 , wherein distinguishing the candidate code by the processor device further comprises comparing n-gram divergence of a prefix code and a remainder code and n-gram divergence of prefix filled codes and remainder filled codes.

13 . The system of claim 8 , wherein predicting probabilities by the processor device further comprises predicting machine-filled codes probabilities by predicting the probability of generating a remainder filled code given a prefix filled code.

14 . The system of claim 8 , wherein distinguishing the candidate code by the processor device further comprises computing a difference between the candidate code probability and the machine-filled codes probabilities.

15 . A non-transitory computer program product comprising a computer-readable storage medium including program code for detecting artificial intelligence (AI) generated computer code, wherein the program code when executed on a computer causes the computer to perform:

masking lines of code from a candidate code to obtain perturbed codes;

generating missing code from the perturbed codes by employing an AI code generator model to obtain machine-filled codes;

predicting probabilities of the candidate code and the machine-filled codes as AI-generated by employing a surrogate model; and

distinguishing the candidate code as AI-generated by comparing the probabilities against a detection threshold to obtain detection results.

16 . The non-transitory computer program product of claim 15 , further comprising flagging the candidate code as AI-generated to detect malicious code for a decision-making entity to perform an action.

17 . The non-transitory computer program product of claim 16 , wherein the action is securing a healthcare management system handling patient vital data by patching the flagged candidate code for potential security risks.

18 . The non-transitory computer program product of claim 15 , wherein predicting probabilities further comprises predicting the candidate code probability by predicting the probability of generating a remainder code given a prefix code.

19 . The non-transitory computer program product of claim 15 , wherein predicting probabilities further comprises predicting machine-filled codes probabilities by predicting the probability of generating a remainder filled code given a prefix filled code.

20 . The non-transitory computer program product of claim 18 , wherein distinguishing the candidate code further comprises computing a difference between the candidate code probability and the machine-filled codes probabilities.

Assignments (2)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Dec 12, 2025
From: NEC LABORATORIES AMERICA, INC.
To: NEC CORPORATION
Reel/Frame 073203/0407 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jun 3, 2024
From: CHENG, WEI; YANG, XIANJUN; CHEN, HAIFENG
To: NEC LABORATORIES AMERICA, INC.
Reel/Frame 067599/0147 →
Continuity (2)
Provisional Application 63521191 · Jun 15, 2023
Related Publication 20240419801A1 · Dec 19, 2024
References Cited (55)
US 10311218B2 · Aharoni · 2019 [cited by examiner]
US 11489874B2 · Roth · 2022 [cited by examiner]
US 12229535B2 · Kaitha · 2025 [cited by examiner]
US 20190354676A1 · Willis · 2019 [cited by examiner]
US 20220206786A1 · Silva · 2022 [cited by examiner]
Allal, L. B., Li, R., Kocetkov, D., Mou, C., Akiki, C., Ferrandis, C. M., . . . & von Werra, L. (Jan. 9, 2023). SantaCoder: don't reach for the stars!. arXiv preprint arXiv:2301.03988. [cited by applicant]
Bavarian, M., Jun, H., Tezak, N., Schulman, J., McLeavey, C., Tworek, J., & Chen, M. (Jul. 28, 2022). Efficient training of language models to fill in the middle. arXiv preprint arXiv:2207.14255. [cited by applicant]
Bian, N., Liu, P., Han, X., Lin, H., Lu, Y., He, B., & Sun, L. (May 8, 2023). A drop of ink may make a million think: The spread of false information in large language models. arXiv preprint arXiv:2305.04812. [cited by applicant]
Bubeck, S., Chandrasekaran, V., Eldan, R., Gehrke, J., Horvitz, E., Kamar, E., . . . & Zhang, Y. (Mar. 2023). Sparks of artificial general intelligence: Early experiments with gpt-4. arXiv preprint arXiv:2303.12712. [cited by applicant]
Chai, Y., Wang, S., Pang, C., Sun, Y., Tian, H., & Wu, H. (Dec. 13, 2022). ERNIE-Code: Beyond english-centric cross-lingual pretraining for programming languages. arXiv preprint arXiv:2212.06742. [cited by applicant]
Chen, M., Tworek, J., Jun, H., Yuan, Q., Pinto, H. P. D. O., Kaplan, J., . . . & Zaremba, W. (Jul. 7, 2021). Evaluating large language models trained on code. arXiv preprint arXiv:2107.03374. [cited by applicant]
Chen, X., Lin, M., Scharli, N., & Zhou, D. (Apr. 11, 2023). Teaching large language models to self-debug. arXiv preprint arXiv:2304.05128. [cited by applicant]
Clement, C. B., Drain, D., Timcheck, J., Svyatkovskiy, A., & Sundaresan, N. (Oct. 7, 2020). PyMT5: multi-mode translation of natural language and Python code with transformers. arXiv preprint arXiv:2010.03150. [cited by applicant]
Dey, A., Bhattacharya, S., & Chaki, N. (Mar. 5, 2019). Software watermarking: Progress and challenges. INAE Letters, 4, 65-75. [cited by applicant]
Dugan, L., Ippolito, D., Kirubarajan, A., Shi, S., & Callison-Burch, C. (Jun. 26, 2023). Real or fake text ?: Investigating human ability to detect boundaries between human-written and machine-generated text. In Proceed… [cited by applicant]
Fried, D., Aghajanyan, A., Lin, J., Wang, S., Wallace, E., Shi, F., . . . & Lewis, M. (Apr. 12, 2022). Incoder: A generative model for code infilling and synthesis. arXiv preprint arXiv:2204.05999. [cited by applicant]
Gao, L., Biderman, S., Black, S., Golding, L., Hoppe, T., Foster, C., . . . & Leahy, C. (Dec. 31, 2020). The pile: An 800gb dataset of diverse text for language modeling. arXiv preprint arXiv:2101.00027. [cited by applicant]
Hamilton, J., & Danicic, S. (Feb. 21, 2011). A survey of static software watermarking. In 2011 World Congress on Internet Security (WorldCIS—2011) (pp. 100-107). IEEE. [cited by applicant]
Hanley, H. W., & Durumeric, Z. (May 16, 2023). Machine-made media: Monitoring the mobilization of machine-generated articles on misinformation and mainstream news websites. arXiv preprint arXiv:2305.09820. [cited by applicant]
Hendrycks, D., Basart, S., Kadavath, S., Mazeika, M., Arora, A., Guo, E., . . . & Steinhardt, J. (May 20, 2021). Measuring coding challenge competence with apps. arXiv preprint arXiv:2105.09938. [cited by applicant]
Christo, L. (Dec. 8, 2021) Training CodeParrot from Scratch. https://huggingface.co/blog/codeparrot. [cited by applicant]
Kirchenbauer, J., Geiping, J., Wen, Y., Katz, J., Miers, I., & Goldstein, T. (Jul. 3, 2023). A watermark for large language models. In International Conference on Machine Learning (pp. 17061-17084). PMLR. [cited by applicant]
Krishna, K., Song, Y., Karpinska, M., Wieting, J., & Iyyer, M. (Feb. 13, 2024). Paraphrasing evades detectors of ai-generated text, but retrieval is an effective defense. Advances in Neural Information Processing System… [cited by applicant]
Kumar, S., Balachandran, V., Njoo, L., Anastasopoulos, A., & Tsvetkov, Y. (Oct. 14, 2022). Language generation models can cause harm: So what can we do about it? an actionable survey. arXiv preprint arXiv:2210.07700. [cited by applicant]
Lee, T., Hong, S., Ahn, J., Hong, I., Lee, H., Yun, S., . . . & Kim, G. (May 24, 2023). Who wrote this code? watermarking for code generation. arXiv preprint arXiv:2305.15060. [cited by applicant]
Li, Y., Choi, D., Chung, J., Kushman, N., Schrittwieser, J., Leblond, R., . . . & Vinyals, O. (Dec. 9, 2022). Competition-level code generation with alphacode. Science, 378(6624), 1092-1097. [cited by applicant]
Liu, J., Xia, C. S., Wang, Y., & Zhang, L. (Feb. 13, 2024). Is your code generated by chatgpt really correct? rigorous evaluation of large language models for code generation. Advances in Neural Information Processing S… [cited by applicant]
Luo, Z., Xu, C., Zhao, P., Sun, Q., Geng, X., Hu, W., . . . & Jiang, D. (Jun. 14, 2023). Wizardcoder: Empowering code large language models with evol-instruct. arXiv preprint arXiv:2306.08568. [cited by applicant]
Ma, H., Jia, C., Li, S., Zheng, W., & Wu, D. (Mar. 29, 2019). Xmark: dynamic software watermarking using Collatz conjecture. IEEE Transactions on Information Forensics and Security, 14(11), 2859-2874. [cited by applicant]
Mireshghallah, N., Mattern, J., Gao, S., Shokri, R., & Berg-Kirkpatrick, T. (May 17, 2023). Smaller language models are better black-box machine-generated text detectors. arXiv preprint arXiv:2305.09859. [cited by applicant]
Mitchell, E., Lee, Y., Khazatsky, A., Manning, C. D., & Finn, C. (Jul. 3, 2023). Detectgpt: Zero-shot machine-generated text detection using probability curvature. In International Conference on Machine Learning (pp. 24… [cited by applicant]
Ni, A., Iyer, S., Radev, D., Stoyanov, V., Yih, W. T., Wang, S., & Lin, X. V. (Jul. 3, 2023). Lever: Learning to verify language-to-code generation with execution. In International Conference on Machine Learning (pp. 26… [cited by applicant]
Nijkamp, E., Pang, B., Hayashi, H., Tu, L., Wang, H., Zhou, Y., . . . & Xiong, C. (Mar. 25, 2022). Codegen: An open large language model for code with multi-turn program synthesis. arXiv preprint arXiv:2203.13474. [cited by applicant]
Oshikawa, R., Qian, J., & Wang, W. Y. (Nov. 2, 2018). A survey on natural language processing for fake news detection. arXiv preprint arXiv:1811.00770. [cited by applicant]
Pan, Y., Pan, L., Chen, W., Nakov, P., Kan, M. Y., & Wang, W. Y. (May 23, 2023). On the risk of misinformation pollution with large language models. arXiv preprint arXiv:2305.13661. [cited by applicant]
Perkins, M., Roe, J., Postma, D., McGaughran, J., & Hickerson, D. (May 29, 2023). Game of tones: faculty detection of GPT-4 generated content in university assessments. arXiv preprint arXiv:2305.18081. [cited by applicant]
Raffel, C., Shazeer, N., Roberts, A., Lee, K., Narang, S., Matena, M., . . . & Liu, P. J. (Jun. 20, 2020). Exploring the limits of transfer learning with a unified text-to-text transformer. Journal of machine learning r… [cited by applicant]
Sun, T., Gaut, A., Tang, S., Huang, Y., ElSherief, M., Zhao, J., . . . & Wang, W. Y. (Jun. 21, 2019). Mitigating gender bias in natural language processing: Literature review. arXiv preprint arXiv:1906.08976. [cited by applicant]
Touvron, H., Lavril, T., Izacard, G., Martinet, X., Lachaux, M. A., Lacroix, T., . . . & Lample, G. (Feb. 27, 2023). Llama: Open and efficient foundation language models. arXiv preprint arXiv:2302.13971. [cited by applicant]
Wang, J., Liu, S., Xie, X., & Li, Y. (Apr. 11, 2023). Evaluating AIGC detectors on code content. arXiv preprint arXiv:2304.05193. [cited by applicant]
Wang, Y., Gong, D., Lu, B., Xiang, F., & Liu, F. (Feb. 27, 2018). Exception handling-based dynamic software watermarking. IEEE Access, 6, 8882-8889. [cited by applicant]
Wang, Y., Wang, W., Joty, S., & Hoi, S. C. (Sep. 2, 2021). Codet5: Identifier-aware unified pre-trained encoder-decoder models for code understanding and generation. arXiv preprint arXiv:2109.00859. [cited by applicant]
Xu, F. F., Alon, U., Neubig, G., & Hellendoorn, V. J. (Jun. 13, 2022). A systematic evaluation of large language models of code. In Proceedings of the 6th ACM SIGPLAN International Symposium on Machine Programming (pp. … [cited by applicant]
Yang, X., Cheng, W., Petzold, L., Wang, W. Y., & Chen, H. (May 27, 2023). Dna-gpt: Divergent n-gram analysis for training-free detection of gpt-generated text. arXiv preprint arXiv:2305.17359. [cited by applicant]
Zan, D., Chen, B., Yang, D., Lin, Z., Kim, M., Guan, B., . . . & Lou, J. G. (Jun. 14, 2022). CERT: continual pre-training on sketches for library-oriented code generation. arXiv preprint arXiv:2206.06888. [cited by applicant]
Zellers, R., Holtzman, A., Rashkin, H., Bisk, Y., Farhadi, A., Roesner, F., & Choi, Y. (Dec. 8, 2019). Defending against neural fake news. Advances in neural information processing systems, 32. [cited by applicant]
Zhang, K., Wang, D., Xia, J., Wang, W. Y., & Li, L. (Feb. 13, 2024). Algo: Synthesizing algorithmic programs with generated oracle verifiers. Advances in Neural Information Processing Systems, 36. [cited by applicant]
Zhang, S., Roller, S., Goyal, N., Artetxe, M., Chen, M., Chen, S., . . . & Zettlemoyer, L. (May 2, 2022). Opt: Open pre-trained transformer language models. arXiv preprint arXiv:2205.01068. [cited by applicant]
Zheng, Q., Xia, X., Zou, X., Dong, Y., Wang, S., Xue, Y., . . . & Tang, J. (Mar. 30, 2023). Codegeex: A pre-trained model for code generation with multilingual evaluations on humaneval-x. arXiv preprint arXiv:2303.17568. [cited by applicant]
Li, L., Wang, P., Ren, K., Sun, T., & Qiu, X. (Apr. 27, 2023). Origin tracing and detecting of llms. arXiv preprint arXiv:2304.14072. [cited by applicant]
Liang, W., Yuksekgonul, M., Mao, Y., Wu, E., & Zou, J. (Jul. 14, 2023). GPT detectors are biased against non-native English writers. Patterns, 4(7). [cited by applicant]
Liu, Y., Zhang, Z., Zhang, W., Yue, S., Zhao, X., Cheng, X., . . . & Hu, H. (Apr. 16, 2023). Argugpt: evaluating, understanding and identifying argumentative essays generated by gpt models. arXiv preprint arXiv:2304.076… [cited by applicant]
Lu, N., Liu, S., He, R., Wang, Q., Ong, Y. S., & Tang, K. (May 18, 2023). Large language models can be guided to evade ai-generated text detection. arXiv preprint arXiv:2305.10847. [cited by applicant]
Shi, Z., Wang, Y., Yin, F., Chen, X., Chang, K. W., & Hsieh, C. J. (Feb. 23, 2024). Red teaming language model detectors with language models. Transactions of the Association for Computational Linguistics, 12, 174-189. [cited by applicant]
Weng, L., Liu, S., Zhu, H., Sun, J., Kam-Kwai, W., Han, D., . . . & Chen, W. (Apr. 7, 2024). Towards an understanding and explanation for mixed-initiative artificial scientific text detection. Information Visualization,… [cited by applicant]