IP Library › Granted Patent US 12,175,225
Granted Patent B2
US 12,175,225 · App. 18/087,379 · Granted Dec 24, 2024

System and method for binary code decompilation using machine learning

Inventors: Aleksandr Ševčenko (Vilnius, LT); Mantas Briliauskas (Vilnius, LT)
Assignee: UAB 360 IT
G06F8/53G06N20/00G06F8/41G06F8/51G06F21/563G06N3/08
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,175,225
App. No.
18/087,379
Granted
Dec 24, 2024
Kind
B2
Abstract

Systems and methods for decompiling binary code or executables are provided herein. In some embodiments, a method of training a machine learning algorithm for decompiling binary code into readable source code includes collecting a data set of source code and at least one element associated with the source code; providing binary code using the data set; training a model configured to decompile the binary code into source code using the data set by: decompiling the collected binary code into intermediate source code; comparing the source code in the data set with the intermediate source code; and updating the model and repeating the training if the source code in the data set differs from the intermediate source code by more than a threshold amount.

Claims (42)

1. A method of training a machine learning algorithm for decompiling binary code into readable source code, the method comprising:

collecting a data set including source code in a programming language and at least one element associated with the source code;

providing binary code using the data set;

training a model configured to decompile the binary code into source code using the data set by:

decompiling the collected binary code into intermediate source code in the programming language;

comparing the source code in the data set with the intermediate source code; and

updating the model and repeating the training if the source code in the data set differs from the intermediate source code by more than a threshold amount, and determining that the model is trained if the source code in the data set does not differ from the intermediate source code by more than a threshold amount.

2. The method of claim 1 , wherein the model is a sequence-to-sequence model.

3. The method of claim 1 , wherein the threshold amount includes an error rate.

4. The method of claim 1 , wherein comparing includes comparing a fuzzy hash of the source code in the data set to a fuzzy hash of the intermediate source code.

5. The method of claim 1 , wherein the at least one element includes the programming language associated with the source code, a compiler name and version associated with the source code, target triplets, a compiler type associated with the source code, or compiler flags.

6. The method of claim 1 , wherein providing binary code using the data set includes compiling the source code in the data set using the at least one element.

7. The method of claim 1 , wherein the source code in the data set is compared with the intermediate source code with a loss function.

8. The method of claim 1 , further comprising:

receiving an executable file or a binary code segment identified within a file;

receiving identifying information associated with the executable file or the binary code segment; and

decompiling the executable or the binary code segment using the identifying information

by processing the executable or the binary code segment using the trained machine learning algorithm.

9. A non-transitory computer-readable medium storing a computer program, which, when read and executed by a computer causes the computer to perform a method of training a machine learning algorithm for decompiling binary code into readable source code, the method comprising:

collecting a data set including source code in a programming language and at least one element associated with the source code;

providing binary code using the data set;

training a model configured to decompile the binary code into source code using the data set by:

decompiling the binary code into intermediate source code in the programming language;

comparing the source code in the data set with the intermediate source code; and

updating the model and repeating the training if the source code in the data set differs from the intermediate source code by more than a threshold amount, and determining that the model is trained if the source code in the data set does not differ from the intermediate source code by more than a threshold amount.

10. The non-transitory computer-readable medium of claim 9 , wherein the model is a sequence-to-sequence model.

11. The non-transitory computer-readable medium of claim 9 , wherein the threshold amount includes an error rate.

12. The non-transitory computer-readable medium of claim 9 , wherein comparing includes comparing a fuzzy hash of the source code of the data set to a fuzzy hash of the intermediate source code.

13. The non-transitory computer-readable medium of claim 9 , wherein the at least one element includes the programming language associated with the source code, a compiler name and version associated with the source code, target triplets, a compiler type associated with the source code, or compiler flags [ ].

14. The non-transitory computer-readable medium of claim 9 , wherein providing binary code using the data set includes compiling the source code in the data set using the at least one element.

15. The non-transitory computer-readable medium of claim 9 , wherein the source code in the data set is compared with the intermediate source code with a loss function.

16. A system for training a machine learning algorithm for decompiling binary code into readable source code, the system having one or more processors configured for:

collecting a data set including source code in the programming language and at least one element associated with the source code;

providing binary code using the data set;

training a model configured to decompile the binary code into source code using the data set by:

decompiling the binary code into intermediate source code in the programming language;

comparing the source code in the data set with the intermediate source code; and

updating the model and repeating the training if the source code in the data set differs from the intermediate source code by more than a threshold amount, and determining that the model is trained if the source code in the data set does not differ from the intermediate source code by more than a threshold amount.

17. The system of claim 16 , wherein the model is a sequence-to-sequence model.

18. The system of claim 16 , wherein the threshold amount includes an error rate.

19. The system of claim 16 , wherein comparing includes comparing a fuzzy hash of the source code of the data set to a fuzzy hash of the intermediate source code.

20. The system of claim 16 , wherein the at least one element includes the programming language associated with the source code, a compiler name and version associated with the source code, target triplets, a compiler type associated with the source code, or compiler flags.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jan 2, 2023
From: SEVCENKO, ALEKSANDR; BRILIAUSKAS, MANTAS
To: UAB 360 IT
Reel/Frame 062252/0567 →
Continuity (1)
Related Publication 20240211225A1 · Jun 27, 2024