IP Library › Granted Patent US 12,393,423
Granted Patent B2
US 12,393,423 · App. 18/434,426 · Granted Aug 19, 2025

Methods and apparatus for intentional programming for heterogeneous systems

Inventors: Adam Herr (Forest Grove, OR); Derek Gerstmann (Del Mar, CA); Justin Gottschlich (Santa Clara, CA); Mikael Bourges-Sevenier (Santa Clara, CA); Sridhar Sharma (Palo Alto, CA)
Assignee: Intel Corporation
G06F9/30174G06F8/31G06F8/52G06F8/75G06F8/76G06F9/3877G06N3/04
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,393,423
App. No.
18/434,426
Granted
Aug 19, 2025
Kind
B2
Abstract

Methods, apparatus, systems and articles of manufacture are disclosed for intentional programming for heterogeneous systems. An example apparatus includes a code lifter to identify annotated code corresponding to an algorithm to be executed on the heterogeneous system based on an identifier being associated with the annotated code, and convert the annotated code in the first representation to intermediate code in a second representation by identifying the intermediate code as having a first algorithmic intent that corresponds to a second algorithmic intent of the annotated code, a domain specific language (DSL) generator to translate the intermediate code in the second representation to DSL code in a third representation when the first algorithmic intent matches the second algorithmic intent, the third representation corresponding to a DSL representation, and a code replacer to invoke a compiler to generate an executable including variant binaries based on the DSL code.

Claims (51)

1. A method comprising:

identifying a first code block in an intermediate representation based on an input code block in an imperative programming language representation, the first code block having a first purpose, the input code block having a second purpose associated with the first purpose;

translating the first code block from the intermediate representation to a domain specific language (DSL) representation;

compiling, by executing an instruction with at least one processor circuit, the DSL representation of the first code block into an executable variant of the input code block, the executable variant to be executed by an accelerator; and

outputting the executable variant.

2. The method of claim 1 , wherein the accelerator includes at least one of a field programmable gate array, a graphics processing unit, a vision processing unit, an application specific integrated circuit, an application-specific instruction set processor, a physics processing unit, a digital signal processor, an image processor, a coprocessor, a floating-point unit, a network processor, or a front-end processor.

3. At least one non-transitory computer readable storage medium comprising instructions that cause at least one processor circuit to at least:

identify a first code block in an intermediate representation based on an input code block in an imperative programming language representation, the first code block having a first purpose, the input code block having a second purpose associated with the first purpose;

translate the first code block from the intermediate representation to a domain specific language (DSL) representation;

compile the DSL representation of the first code block into an executable variant of the input code block, the executable variant to be executed by an accelerator; and

output the executable variant.

4. The at least one non-transitory computer readable storage medium of claim 3 , wherein the accelerator includes at least one of a field programmable gate array, a graphics processing unit, a vision processing unit, an application specific integrated circuit, an application-specific instruction set processor, a physics processing unit, a digital signal processor, an image processor, a coprocessor, a floating-point unit, a network processor, or a front-end processor.

5. The at least one non-transitory computer readable storage medium of claim 3 , wherein the instructions cause one or more of the at least one processor circuit to determine that the second purpose corresponds to the first purpose by executing a machine learning model.

6. The at least one non-transitory computer readable storage medium of claim 3 , wherein the instructions cause one or more of the at least one processor circuit to:

determine that the input code block includes a loop;

determine that the first code block is a stencil code block; and

translate the first code block into the DSL representation based on a replacement of the loop with the stencil code block.

7. The at least one non-transitory computer readable storage medium of claim 3 , wherein the accelerator is a first accelerator, and the instructions cause one or more of the at least one processor circuit to:

generate a first schedule for the first accelerator to execute a first portion of a workload, the first schedule based on the DSL representation;

generate a second schedule for a second accelerator to execute a second portion of the workload, the second schedule based on the DSL representation;

cause the first accelerator to execute the first portion based on the first schedule; and

cause the second accelerator to execute the second portion based on the second schedule.

8. The at least one non-transitory computer readable storage medium of claim 7 , wherein the instructions cause one or more of the at least one processor circuit to determine that the first accelerator is to execute the first portion in parallel with the execution of the second portion by the second accelerator.

9. The at least one non-transitory computer readable storage medium of claim 7 , wherein the second accelerator is a graphics processing unit, and the instructions cause one or more of the at least one processor circuit to:

determine that the first accelerator is to execute the first portion and the second portion; and

based on a determination that utilization of the first accelerator satisfies a threshold, cause the first accelerator to offload the second portion to the graphics processing unit.

10. The at least one non-transitory computer readable storage medium of claim 3 , wherein the instructions cause one or more of the at least one processor circuit to output the executable variant to the accelerator.

11. An apparatus comprising:

at least one interface circuit;

machine-readable instructions; and

at least one processor circuit to utilize the machine-readable instructions to at least:

identify a first code block in an intermediate representation based on an input code block in an imperative programming language representation, the first code block having a first purpose, the input code block having a second purpose associated with the first purpose;

translate the first code block from the intermediate representation to a domain specific language (DSL) representation;

compile the DSL representation of the first code block into an executable variant of the input code block, the executable variant to be executed by an accelerator; and

output the executable variant.

12. The apparatus of claim 11 , wherein the accelerator includes at least one of a field programmable gate array, a graphics processing unit, a vision processing unit, an application specific integrated circuit, an application-specific instruction set processor, a physics processing unit, a digital signal processor, an image processor, a coprocessor, a floating-point unit, a network processor, or a front-end processor.

13. The apparatus of claim 11 , wherein one or more of the at least one processor circuit is to determine that the second purpose corresponds to the first purpose by executing a machine learning model.

14. The apparatus of claim 11 , wherein one or more of the at least one processor circuit is to:

determine that the input code block includes a loop;

determine that the first code block is a stencil code block; and

translate the first code block into the DSL representation based on a replacement of the loop with the stencil code block.

15. The apparatus of claim 11 , wherein the accelerator is a first accelerator, and one or more of the at least one processor circuit is to:

generate a first schedule for the first accelerator to execute a first portion of a workload, the first schedule based on the DSL representation;

generate a second schedule for a second accelerator to execute a second portion of the workload, the second schedule based on the DSL representation;

cause the first accelerator to execute the first portion based on the first schedule; and

cause the second accelerator to execute the second portion based on the second schedule.

16. The apparatus of claim 15 , wherein one or more of the at least one processor circuit is to determine that the first accelerator is to execute the first portion in parallel with the execution of the second portion by the second accelerator.

17. The apparatus of claim 15 , wherein the second accelerator is a graphics processing unit, and one or more of the at least one processor circuit is to:

determine that the first accelerator is to execute the first portion and the second portion; and

based on a determination that utilization of the first accelerator satisfies a threshold, cause the first accelerator to offload the second portion to the graphics processing unit.

18. The apparatus of claim 11 , wherein one or more of the at least one processor circuit is to output the executable variant to the accelerator.

Continuity (3)
Continuation 17672142 · Feb 15, 2022
Continuation 16455388 · Jun 27, 2019
Related Publication 20240329997A1 · Oct 3, 2024
References Cited (98)
US 8296743B2 · Linderman · 2012 [cited by examiner]
US 9483282B1 · Vandervennet · 2016 [cited by applicant]
US 9800466B1 · Rangole · 2017 [cited by applicant]
US 10007520B1 · Ross · 2018 [cited by applicant]
US 10187252B2 · Byers · 2019 [cited by applicant]
US 10289394B2 · Huang et al. · 2019 [cited by applicant]
US 10445118B2 · Guo · 2019 [cited by applicant]
US 10713213B2 · Tamir · 2020 [cited by applicant]
US 10831451B2 · Udupa et al. · 2020 [cited by applicant]
US 10908884B2 · Herr · 2021 [cited by applicant]
US 11036477B2 · Herr · 2021 [cited by applicant]
US 11269639B2 · Herr et al. · 2022 [cited by applicant]
US 11449347B1 · Kong · 2022 [cited by examiner]
US 20090158248A1 · Linderman · 2009 [cited by applicant]
US 20100153934A1 · Lachner · 2010 [cited by applicant]
US 20110289519A1 · Frost · 2011 [cited by applicant]
US 20130212365A1 · Chen · 2013 [cited by applicant]
US 20160210174A1 · Hsieh · 2016 [cited by applicant]
US 20160350088A1 · Ravishankar · 2016 [cited by applicant]
US 20170123775A1 · Xu · 2017 [cited by applicant]
US 20180082212A1 · Faivishevsky · 2018 [cited by applicant]
US 20180173675A1 · Tamir · 2018 [cited by applicant]
US 20180183660A1 · Byers · 2018 [cited by applicant]
US 20190103872A1 · Clark · 2019 [cited by examiner]
US 20190317740A1 · Herr · 2019 [cited by applicant]
US 20190317741A1 · Herr · 2019 [cited by applicant]
US 20190317880A1 · Herr · 2019 [cited by applicant]
US 20190324755A1 · Herr et al. · 2019 [cited by applicant]
US 20200334084A1 · Jacobson · 2020 [cited by examiner]
United States Patent and Trademark Office, “Corrected Notice of Allowance” issued in connection with U.S. Appl. No. 16/455,628, dated Feb. 26, 2021, 7 pages. [cited by applicant]
United States Patent and Trademark Office, “Non-Final Action” issued in U.S. Appl. No. 16/455,388 on Mar. 3, 2021 (21 pages). [cited by applicant]
United States Patent and Trademark Office, “Advisory Action” issued in connection with U.S. Appl. No. 16/455,486 dated Mar. 10, 2021, 4 pages. [cited by applicant]
United States Patent and Trademark Office, “Corrected Notice of Allowance” issued in connection with U.S. Appl. No. 16/455,628, dated May 7, 2021, 2 pages. [cited by applicant]
United States Patent and Trademark Office, “Final Action” issued in U.S. Appl. No. 16/455,388 on Jul. 7, 2021 (20 pages). [cited by applicant]
United States Patent and Trademark Office, “Advisory Action”, issued in connection with U.S. Appl. No. 16/455,388, filed Oct. 13, 2021, 3 pages. [cited by applicant]
United States Patent and Trademark Office, “Notice of Allowance” issued in U.S. Appl. No. 16/455,388 on Nov. 2, 2021 (7 pages). [cited by applicant]
European Patent Office, “Communication pursuant to Article 94(3) EPC,” issued in connection with European patent application No. 20165699.8-1203, Jul. 5, 2022, 8 pages. [cited by applicant]
United States Patent and Trademark Office, “Non-Final Office Action”, issued in connection with U.S. Appl. No. 17/672,142, Apr. 13, 2023, 19 pages. [cited by applicant]
European Patent Office, “Communication under Rule 71(3) EPC—Intention to Grant,” issued in connection with European Patent Application No. 20 165 699.8-1203, dated Oct. 20, 2023, 98 pages. [cited by applicant]
United States Patent and Trademark Office, “Notice of Allowance and Fee(s) Due,” issued in connection with U.S. Appl. No. 17/672,142, dated Nov. 7, 2023, 8 pages. [cited by applicant]
United States Patent and Trademark Office, “Notice of Allowance,” issued in connection with U.S. Appl. No. 17/672,142, dated Feb. 28, 2024, 3 pages. [cited by applicant]
European Patent Office, “Extended European Search Report,” issued in connection with European Patent Application No. 24155106.8-1203, dated Mar. 28, 2024, 16 pages. [cited by applicant]
P. Balaprakash et al., “Autotuning in High-Performance Computing Applications,” in Proceedings of the IEEE, vol. 106, No. 11, pp. 2068-2083, Nov. 2018, doi: 10.1109/JPROC.2018.2841200 (16 pages). [cited by applicant]
Ragan-Kelley, “Halide,” [Online]. last retrieved Jul. 10, 2019, Available: http://halide-lang.org/, 3 pages. [cited by applicant]
United States Patent and Trademark Office, “Corrected Notice of Allowance” issued in connection with U.S. Appl. No. 16/455,379, dated Dec. 11, 2020, 2 pages. [cited by applicant]
Chen et al., “Learning to Optimize Tensor Programs,” 32nd Conference on Neural Information Processing Systems (NeurIPS 2018), Montréal, Canada, 12 pages. [cited by applicant]
Mullapudi et al., “Automatically Scheduling Halide Image Processing Pipelines,” SIGGRAPH, Anaheim, Jul. 24-28, 2016, 11 pages. [cited by applicant]
Chen et al., “TVM: An Automated End-to-End Optimizing Compiler for Deep Learning,” in SysML 2018, Palo Alto, 2018, 17 pages. [cited by applicant]
Iyer, Srinivasan, et al., “Learning a Neural Semantic Parser from User Feedback,” Proceedings of the 55th Annual Meeting of the Association for Computational Linguistics, Vancouver, Canada, Jul. 30-Aug. 4, 2017, p. 963-… [cited by applicant]
Ragan-Kelley et al., “Halide: a language and compiler for optimizing parallelism, locality, and recomputation in image processing pipelines,” PLDI Seattle, Jun. 16-21, 2013, 12 pages. [cited by applicant]
Hoffmann et al., “Dynamic Knobs for Responsive Power-Aware Computing,” Mar. 2011, 14 pages. [cited by applicant]
Huang et al., “Programming and Runtime Support to Blaze FPGA Accelerator Deployment at Datacenter Scale,” Oct. 2016, 21 pages. [cited by applicant]
Bergeron et al., “Hardware JIT compilation for off-the-shelf dynamically reconfigurable FPGAs,” DIRO, Universite de Montreal GRM, Ecole Polytechnique de Montreal, Budapest, Hungary, Mar. 29-Apr. 6, 2008 , 16 pages. [cited by applicant]
Greskamp et al., “A Virtual Machine for Merit-Based Runtime Reconfiguration,” Proceedings of the 13th Annual IEEE Symposium on Field-Programmable Custom Computing Machines, 2005, 2 pages. [cited by applicant]
Ragan-Kelley, “Decoupling Algorithms from the Organization of Computation for High Performance Image Processing: The design and implementation of the Halide language and compiler,” MIT, Cambridge, MA, Jun. 2014, 187 pag… [cited by applicant]
Yaghmazadeh, Navid, et al., “SQLizer: Query Synthesis from Natural Language,” Proc. ACM Program. Lang., vol. 1, No. OOPSLA, Article 63, Oct. 2017, p. 63:1-63:26 (26 pages). [cited by applicant]
Altera, “FPGA Run-Time Reconfiguration: Two Approaches,” Mar. 2008, ver. 1.0, 6 pages. [cited by applicant]
Munshi et al., “OpenCL Programming Guide,” Addison-Wesley Professional, Upper Saddle River, 2012, 120 pages. [cited by applicant]
Van De Geijn et al., “SUMMA: Scalable Universal Matrix Multiplication Algorithm,” University of Texas at Austin, Austin, 1995, 19 pages. [cited by applicant]
Wang et al. “EXOCHI”, Proceedings of the 2007 ACM SIGPLAN Conference on Programming Language Design and Implementation—PLDI '07: San Diego, CA, USA, Jun. 10-13, 2007; doi:10.1145/1250734.1250753, 11 pages. [cited by applicant]
Dagum et al., “OpenMP: An Industry-Standad API for Shared-Memory Programming,” IEEE Computational Science and Engineering, vol. 5, No. 1, pp. 46-55, Jan.-Mar. 1998, 10 pages. [cited by applicant]
Nickolls et al., “Scalable parallel programming with CUDA,” ACM Queue-GPU Computing, vol. 6, No. 2, pp. 40-53, 2008, I. Buck, M. Garland and K. Skadron, “Scalable Parallel Programming with CUDA,” ACM Queue-GPU Computing… [cited by applicant]
Raman et al., “Parcae: A system for flexibble parallel execution,” Jun. 2012, 20 pages. [cited by applicant]
The Khronos Group, “OpenCL Specification,” Nov. 14, 2012, version 1.2, 380 pages. [cited by applicant]
Lei, Tao et al., “From Natural Language Specifications to Program Input Parsers,” Proceedings of the 51st Annual Meeting of the Association for Computational Linguistics (vol. 1: Long Papers), Sofia, Bulgaria, Aug. 2013… [cited by applicant]
Ansel et al., “OpenTuner: An Extensible Framework for Program Autotuning,” Computer Science and Artificial Intelligence Laboratory Technical Report, MIT, Cambridge, Nov. 1, 2013, 15 pages. [cited by applicant]
Kamil et al., “Verified Lifting of Stencil Computations,” PLDI, Santa Barbara, pp. 711-726, Jun. 13-17, 2016, 16 pages. [cited by applicant]
Intel, “Intel® FPGA SDK for OpenCL Best Practices Guide,” May 8, 2017, 135 pages. [cited by applicant]
Cray X1TM System, “Optimizing Processor-bound Code,” http://docs.cray.com/books/S-2315-52/html-S-2315-52/z1073673157.html, 12 pages, retrieved on Sep. 22, 2017. [cited by applicant]
Greaves, “Distributing C# Methods and Threads over Ethernet-connected FPGAs using Kiwi,” 2011, retrieved on Sep. 22, 2017, 13 pages. [cited by applicant]
Altera, “Machines Ensuring the Right Path,” retrieved on Sep. 22, 2017, 4 pages. [cited by applicant]
IBM Reasearch, “Liquid Metal,” retrieved on Sep. 22, 2017, http://researcher.watson.ibm.com/researcher/view_group.php?id=122, 4 pages. [cited by applicant]
Maaz Bin Safeer Ahmad et al: “Automatically Leveraging MapReduce Frameworks for Data-Intensive Applications”, arxiv.org, Cornell University Library, 201 Olin Library Cornell University Ithaca, NY 14853, Jan. 30, 2018 (J… [cited by applicant]
Justin Gottschlich et al., “The Three Pillars of Machine Programming,” Intel Labs, MIT, May 8, 2018 (11 pages). [cited by applicant]
United States Patent and Trademark Office, “Non-Final Office Action” issued in connection with U.S. Appl. No. 15/713,301, dated Jul. 20, 2018, 10 pages. [cited by applicant]
United States Patent and Trademark Office, “Final Office Action”, issued in connection with U.S. Appl. No. 15/713,301, on Jan. 28, 2019, 12 pages. [cited by applicant]
Barati Saeid et al., “Proteus: Language and Runtime Support for Self-Adaptive Software Development,” IEEE Software, Institute of Electrical and Electronics Engineers, US, vol. 36, No. 2, Mar. 1, 2019, pp. 73-82, XP01171… [cited by applicant]
Sami Lazreg et al., “Multifaceted Automated Analyses for Variability-intensive embedded systems,” 41st ACM/IEEE International Conference on Software Engineering, May 2019, Montreal, Canada, hal-02061251, retrieved onlin… [cited by applicant]
United States Patent and Trademark Office, “Notice of Allowance”, issued in connection with U.S. Appl. No. 15/713,301, on May 23, 2019, 7 pages. [cited by applicant]
Mejbah Alam et al., “A Zero-Positive Learning Approach for Diagnosing Software Performance Regressions,” May 31, 2019, XP055749022, Retrieved from the Internet: URL: https://arxiv.org/pdf/1709.07536v4.pdf, retrieved on … [cited by applicant]
Adams et al., “Learning to Optimize Halide with Tree Search and Random Programs,” ACM Trans. Graph, vol. 38, No. 4, pp. 121:1-121:12, Jul. 2019, 12 pages. [cited by applicant]
Ahmad, M. and Cheung, A., “Metalift,” [Online]. Last retrieved Jul. 10, 2019, Available: http://metalift.uwplse.org/., 4 pages. [cited by applicant]
“SYCL: C++ Single-source Heterogeneos Programming for OpenCL,” Khronos, [Online] last retrieved Jul. 10, 2019, available: https://www.khronos.org/sycl/, 7 pages. [cited by applicant]
The Khronos Group, “Vulkan 1.1 API Specification.” Khronos, Mar. 2018, [Online]. last retrieved Sep. 25, 2019, Available: https://www.khronos.org/registry/vulkan/, 15 pages. [cited by applicant]
Ahmad et al., “Automatically translating image processing libraries to Halide,” ACM Trans. Graph. vol. 38, No. 6, pp. 204:2-204:13, last retrieved Oct. 2, 2019. [cited by applicant]
United States Patent and Trademark Office, “Non-Final Office Action” issued in connection with U.S. Appl. No. 16/455,379, dated Apr. 13, 2020, 11 pages. [cited by applicant]
United States Patent and Trademark Office, “Non-Final Office Action”, issued in connection with U.S. Appl. No. 16/455,486, on Jun. 25, 2020, 16 pages. [cited by applicant]
United States Patent and Trademark Office, “Notice of Allowance”, issued in connection with U.S. Appl. No. 16/455,379, on Jul. 23, 2020, 10 pages. [cited by applicant]
United States Patent and Trademark Office, “Non-Final Office Action”, issued in connection with U.S. Appl. No. 16/455,628, on Aug. 7, 2020, 18 pages. [cited by applicant]
United States Patent and Trademark Office, “Corrected Notice of Allowance” issued in connection with U.S. Appl. No. 16/455,379, dated Aug. 28, 2020, 2 pages. [cited by applicant]
European Patent Office, “European Search Report,” issued in connection with European patent application No. 20165699.8-1224, Nov. 19, 2020, 11 pages. [cited by applicant]
United States Patent and Trademark Office, “Corrected Notice of Allowance”, issued in connection with U.S. Appl. No. 16/455,628, on Dec. 11, 2020, 2 pages. [cited by applicant]
United States Patent and Trademark Office, “Final Office Action” issued in connection with U.S. Appl. No. 16/455,486, dated Dec. 14, 2020, 19 pages. [cited by applicant]
United States Patent and Trademark Office, “Corrected Notice of Allowance”, issued in connection with U.S. Appl. No. 16/455,379, on Dec. 30, 2020, 5 pages. [cited by applicant]
United States Patent and Trademark Office, “Notice of Allowance” issued in connection with U.S. Appl. No. 16/455,628, dated Feb. 5, 2021, 9 pages. [cited by applicant]
Alam et al., “A Zero-Positive Learning Approach for Diagnosing Software Performance Regressions,” arXiv:1709.07536, dated May 31, 2019, 12 pages. [cited by applicant]
Barati et al., “Proteus: Language and Runtime Support for Self-Adaptive Software Development,” published in IEEE Software, vol. 36, Issue: 2, Mar.-Apr. 2019, published Feb. 22, 2019, 10 pages. [cited by applicant]
European Patent Office, “Communication pursuant to Article 94(3) EPC,” issued in connection with European Patent Application No. 24155106.8, dated Mar. 10, 2025, 6 pages. [cited by applicant]