IP Library › Granted Patent US 12,314,690
Granted Patent B2
US 12,314,690 · App. 17/133,238 · Granted May 27, 2025

Methods and apparatus for automatic detection of software bugs

Inventors: Fangke Ye (Atlanta, GA); Justin Gottschlich (Santa Clara, CA); Shengtian Zhou (Palo Alto, CA); Roshni Iyer (Fremont, CA); Jesmin Jahan Tithi (San Jose, CA)
Assignee: INTEL CORPORATION
G06F8/34G06F8/60G06F8/77
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,314,690
App. No.
17/133,238
Granted
May 27, 2025
Kind
B2
Abstract

Methods, systems, and apparatus for automatic detection of software bugs are disclosed. An example apparatus includes a comparator to compare reference code to input code to detect a source code error in the input code; a graph generator to generate a graphical representation of the reference code or the input code, the graphical representation to identify non-overlapping code regions; and a root cause determiner to determine a root cause of the source code error in the input code, the root cause based on the non-overlapping code regions.

Claims (49)

1. An apparatus comprising:

interface circuitry;

machine-readable instructions;

at least one processor circuit to be programmed by the machine-readable instructions to:

analyze execution of a code snippet to determine whether the code snippet invokes an application programming interface of a test framework;

after determining that the code snippet does not invoke the application programming interface of the test framework, map the code snippet into a computer-generated vector space in memory or storage based on a first code vector of the code snippet;

compare the first code vector from the computer-generated vector space to a second code vector of reference code;

generate a semantic similarity score based on a distance between the first code vector of the code snippet and the second code vector of the reference code;

select the code snippet as input code based on the semantic similarity score;

generate at least one graphical representation of at least one of the reference code or the input code, the at least one graphical representation to identify a non-overlapping code region; and

determine a root cause of a source code error in the input code, the root cause based on the non-overlapping code region.

2. The apparatus of claim 1 , wherein the at least one graphical representation is a program-derived semantic graph.

3. The apparatus of claim 1 , wherein one or more of the at least one processor circuit is to select the code snippet as the input code based on a code similarity system.

4. The apparatus of claim 3 , wherein one or more of the at least one processor circuit is to form a code cluster, the code cluster including the reference code and code snippet.

5. The apparatus of claim 4 , wherein one or more of the at least one processor circuit is to form the code cluster based on a vector-based representation of the code snippet.

6. The apparatus of claim 1 , wherein the machine-readable instructions are to cause one or more of the at least one processor circuit to extract the code snippet from a code repository.

7. The apparatus of claim 1 , wherein one or more of the at least one processor circuit is to:

determine a type of clustering based on at least one of computational resources or data availability;

assign the code snippet to a cluster based on the type of clustering; and

select the code snippet and the reference code from the cluster.

8. At least one non-transitory computer readable medium comprising instructions to cause at least one processor circuit to at least:

analyze execution of a code snippet to determine whether the code snippet includes a test that invokes an application programming interface of a test framework, the test to identify the code snippet as reference code;

after determining that the code snippet does not include the test, map the code snippet into a computer-generated vector space in memory or storage based on a first code vector of the code snippet;

compare the first code vector from the computer-generated vector space to a second code vector of the reference code;

generate a semantic similarity score based on a distance between the first code vector of the code snippet and the second code vector of the reference code;

select the code snippet as input code based on the semantic similarity score;

detect, based on the reference code, a source code error in the input code;

detect a non-overlapping code region based on at least one graphical representation of at least one of the reference code or the input code; and

determine a root cause of the source code error based on the non-overlapping code region.

9. The at least one non-transitory computer readable medium as defined in claim 8 , wherein the instructions are to cause one or more of the at least one processor circuit to generate a program-derived semantic graph.

10. The at least one non-transitory computer readable medium as defined in claim 8 , wherein the instructions are to cause one or more of the at least one processor circuit to identify the reference code based on a code similarity system, and select the code snippet using the code similarity system.

11. The at least one non-transitory computer readable medium as defined in claim 10 , wherein the instructions are to cause one or more of the at least one processor circuit to form a code cluster, the code cluster including the reference code and the code snippet.

12. The at least one non-transitory computer readable medium as defined in claim 11 , wherein the instructions are to cause one or more of the at least one processor circuit to form the code cluster using a vector-based representation of the code snippet.

13. The at least one non-transitory computer readable medium as defined in claim 8 , wherein the instructions are to cause one or more of the at least one processor circuit to extract the code snippet from a code repository.

14. An apparatus comprising:

identifier circuitry to determine whether a code snippet invokes an application programming interface of a test framework based on execution of the code snippet;

mapper circuitry to, after a determination that the code snippet does invoke the application programming interface of the test framework, map the code snippet into a computer-generated vector space in memory or storage based on a first code vector of the code snippet;

software bug detector circuitry to select the code snippet as input code based on a semantic similarity score, the semantic similarity score based on a distance between the first code vector of the code snippet and a second code vector of reference code;

graph generator circuitry to generate at least one graphical representation of at least one of the reference code or the input code, the at least one graphical representation to identify a non-overlapping code region; and

root cause determiner circuitry to determine a root cause of a source code error in the input code, the root cause based on the non-overlapping code region.

15. The apparatus of claim 14 , wherein the at least one graphical representation is a program-derived semantic graph.

16. The apparatus of claim 14 , wherein the software bug detector circuitry is to select the code snippet as the input code based on a code similarity system.

17. The apparatus of claim 14 , including clusterer circuitry to form a code cluster, the code cluster including the reference code and the code snippet.

18. The apparatus of claim 17 , wherein the clusterer circuitry is to form the code cluster based on a vector-based representation of the code snippet.

19. The apparatus of claim 17 , wherein the clusterer circuitry is to form the code cluster based on deep neural network mapping.

20. The apparatus of claim 14 , including clusterer circuitry to:

determine a type of clustering based on at least one of computational resources or data availability; and

assign the code snippet to a cluster based on the type of clustering, the software bug detector circuitry is to select the code snippet and the reference code from the cluster.

21. The apparatus of claim 14 , including extractor circuitry to extract the code snippet from a code repository.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Mar 8, 2021
From: YE, FANGKE; GOTTSCHLICH, JUSTIN; ZHOU, SHENGTIAN; IYER, ROSHNI; JAHAN TITHI, JESMIN
To: INTEL CORPORATION
Reel/Frame 055517/0602 →
Continuity (1)
Related Publication 20210182031A1 · Jun 17, 2021
References Cited (61)
US 8336030B1 · Boissy · 2012 [cited by examiner]
US 8499280B2 · Davies · 2013 [cited by examiner]
US 10511554B2 · Adam et al. · 2019 [cited by applicant]
US 10776106B2 · Sahu · 2020 [cited by examiner]
US 10809984B2 · Mizrahi et al. · 2020 [cited by applicant]
US 11379221B2 · Zhou · 2022 [cited by examiner]
US 11720600B1 · Dixit · 2023 [cited by examiner]
US 20070168946A1 · Drissi · 2007 [cited by examiner]
US 20160062765A1 · Ji · 2016 [cited by examiner]
US 20170177712A1 · Kopru · 2017 [cited by examiner]
US 20170371860A1 · McAteer et al. · 2017 [cited by applicant]
US 20180309636A1 · Strom et al. · 2018 [cited by applicant]
US 20180373701A1 · McAteer et al. · 2018 [cited by applicant]
US 20190243622A1 · Allamanis et al. · 2019 [cited by applicant]
US 20190317743A1 · Cremeans et al. · 2019 [cited by applicant]
US 20190324731A1 · Zhou et al. · 2019 [cited by applicant]
US 20200074322A1 · Chungapalli et al. · 2020 [cited by applicant]
US 20200159934A1 · Yamaguchi et al. · 2020 [cited by applicant]
US 20200401662A1 · Chen · 2020 [cited by examiner]
US 20210073632A1 · Iyer et al. · 2021 [cited by applicant]
US 20210117807A1 · Zhou et al. · 2021 [cited by applicant]
US 20210210183A1 · Niggemann · 2021 [cited by examiner]
US 20210255853A1 · Zhou · 2021 [cited by examiner]
US 20210319357A1 · He et al. · 2021 [cited by applicant]
US 20210342490A1 · Briancon · 2021 [cited by examiner]
US 20220107799A1 · Wu et al. · 2022 [cited by applicant]
US 20220148699A1 · Kogan et al. · 2022 [cited by applicant]
US 20230085500A1 · Hong · 2023 [cited by examiner]
EP 3862822A1 · 2021 [cited by applicant]
Ahmet Okutan, “Use of Source Code Similarity Metrics in Software Defect Prediction”, published by Cornell University, 2018, pp. 1-14 (Year: 2018). [cited by examiner]
Gottschlich et al., “The Three Pillars of Machine Programming,” 2018, Intel Labs, MIT, Retrieved from the Internet <https://arxiv.org/abs/1803.07244> 11 pages. [cited by applicant]
Iyer et al., “Software Language Comprehension using a Program-Derived Semantic Graph,” Dec. 11, 2020, 34th Conference on Neural Information Processing Systems (NeurIPS), Computer-Assisted Programming Workshop, Vancouver… [cited by applicant]
Dinella, et al., “Hoppity: Learning Graph Transformations to Detect and Fix Bugs in Programs,” last modified on May 22, 2020, Published as a conference paper at International Conference on Learning Representations (ICLR… [cited by applicant]
“CodeQLdocumentation: QL Tutorials,” Dec. 18, 2020, GitHub, Inc., Retrieved from the Internet <https://help.semmle.com/QL/learn-ql/beginner/ql-tutorials.html> 2 pages. [cited by applicant]
Alam et al., “A Zero-Positive Learning Approach for Diagnosing Software Performance Regressions,” last revised on Jan. 1, 2020, 33rd Conference on Neural Information Processing Systems (NurIPS), Vancouver, Canada, Retri… [cited by applicant]
Amazon Code Guru, “Find Your Most Expensive Lines of Code and Improve Code Quality,” Dec. 9, 2020, Amazon Web Services, Inc., Retrieved from the Internet: <https://aws.amazon.com/codeguru/> 14 pages. [cited by applicant]
Marginean et al., “SapFix: Automated End-to-End Repair at Scale,” May 25-31, 2019, IEEE/ACM 41st International Conference on Software Engineering: Software Engineering in Practice (ICSE-SEIP), Montreal, QC, Canada, 10 p… [cited by applicant]
Allamanis et al., “A Survey of Machine Learning for Big Code and Naturalness,” Jul. 2018, ACM Computing Surveys, vol. 51, No. 4, Article 81, 37 pages. [cited by applicant]
Pradel et al., “DeepBugs: A Learning Approach to Name-Based Detection,” Nov. 2018, Proc. ACM Program, Lang, vol. 2, OOPSLA, Article 147, 25 pages. [cited by applicant]
Ye et al., “MISIM: A Novel Code Similarity System,” last revised on Oct. 9, 2020, Retrieved from the Internet: <https://arxiv.org/pdf/2006.05265.pdf> 20 pages. [cited by applicant]
Alexander Breckel, “Error mining,” XP058057451, Publication date: Jun. 2, 2012, Conference Proceedings Article, Mining Software Repositories, IEEE Press, 4 pages. [cited by applicant]
Netherlands Intellectual Property Office, “Search Report National,” issued in connection with Netherlands Patent Application No. 2029881, dated Feb. 16, 2023, 8 Pages. [cited by applicant]
Gottschlich et al., “The Pillars of Machine Programming,” Jun. 2018, Intel Labs, MIT, 11 pages. [cited by applicant]
United States Patent and Trademark Office, “Non-Final Office Action,” issued Feb. 8, 2024 in connection with U.S. Appl. No. 17/133,168, 9 pages. [cited by applicant]
United States Patent and Trademark Office, “Final Office Action,” issued Jun. 21, 2024 in connection with U.S. Appl. No. 17/133,168, 11 pages. [cited by applicant]
Allamanis et al., “Learning to Represent Programs with Graphs,” International Conference on Learning Representations, May 2018, 17 pages. [cited by applicant]
Netherlands Patent Office, “Search Report,” issued on Feb. 27, 2023 in connection with NL Patent Application No. 2029883, 26 pages, including machine translation. [cited by applicant]
Alon et al., “code2vec: Learning Distributed Representations of Code,” Proceedings of the ACM Program on Programming Languages, Jan. 2019, Article 40, retrieved from <https://dl.acm.org/doi/pdf/10.1145/3290353>, 29 page… [cited by applicant]
Luan et al., “Aroma: Code Recommendation via Structural Code Search,” Oct. 2019, Article 152, retrieved from <https://dl.acm.org/doi/pdf/10.1145/3360578>, 28 pages. [cited by applicant]
Ben-Nun et al., “Neural Code Comprehension: A Learnable Representation of Code Semantics,” ArXiv, Nov. 29, 2018, retrieved from <https://arxiv.org/pdf/1806.07336.pdf>, 17 pages. [cited by applicant]
Wikipedia, “Multi-label Classification,” Dec. 8, 2020, retrieved from <https://en.wikipedia.org/w/index.php?title=Multi-label_classification&oldid=993057796>, 4 pages. [cited by applicant]
Github, “Tree-Sitter,” Nov. 26, 2020, retrieved from <https://github.com/tree-sitter/tree-sitter/tree/53949b09fdca24bb9d3c17df0404fa96c6ced2f1>, 2 pages. [cited by applicant]
Parr, “ANTLR,” 2014, retrieved from <https://www.antlr.org/>, 3 pages. [cited by applicant]
Bahdanau et al., “Neural Machine Translation By Jointly Learning to Align and Translate,” ArXiv, May 19, 2016, retrieved from <https://arxiv.org/pdf/1409.0473.pdf>, 15 pages. [cited by applicant]
Schlichtkrull et al., “Modeling Relational Data with Graph Convolutional Networks,” ArXiv, Oct. 26, 2017, retrieved from <https://arxiv.org/pdf/1703.06103.pdf>, 9 pages. [cited by applicant]
Cosentino et al., “A Systematic Mapping Study of Software Development with GitHub,” IEEE Access, vol. 5, pp. 7173-7192, Jun. 7, 2017, 20 pages. [cited by applicant]
Ellis et al., “Write, Execute, Assess: Program Synthesis with A REPL,” 33rd Conference on Neural Information Processing Systems, 2019, 10 pages. [cited by applicant]
Kipf et al., “Semi-Supervised Classification with Graph Convolutional Networks,” ICLR, Feb. 22, 2017, 14 pages. [cited by applicant]
Radoi, “Unhack all the Code!,” Dec. 4, 2019, 9 pages. [cited by applicant]
Ragan-Kelley et al., “Halide: A Language and Compiler for Optimizing Parallelism Locality, and Recomputation in Image Processing Pipelines,” PLDI, Jun. 2013, 12 pages. [cited by applicant]
Satish et al. “Can Traditional Programming Bridge the Ninja Performance Gap for Parallel Computing Applications,” Communication of the ACM, vol. 58, No. 5, May 2015, 10 pages. [cited by applicant]
Cited By (2)
US 12,524,539 US 12,670,163