IP Library › Granted Patent US 12,288,066
Granted Patent B1
US 12,288,066 · App. 18/320,036 · Granted Apr 29, 2025

Operation fusion for instructions bridging execution unit types

Inventors: Zhaoxiang Jin (Austin, TX); Francesco Spadini (Sunset Valley, TX); Skanda K. Srinivasa (Austin, TX); Milos Becvar (Cedar Park, TX)
Assignee: Apple Inc.
G06F9/30134G06F9/3001G06F9/30036G06F9/30043
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,288,066
App. No.
18/320,036
Granted
Apr 29, 2025
Kind
B1
Abstract

Techniques are disclosed that relate to fusing operations for execution of certain instructions. A processor may include a first execution circuit, of a first type, coupled to a first register file, a second execution circuit, of a second type, coupled to a second register file and a load/store circuit coupled to the first and second register files. The load/store circuit includes an issue port configured to receive an instruction operation for execution, a memory execution circuit configured to execute memory access operations, and a register transfer execution circuit. The register transfer execution circuit is configured to execute instruction operations specifying data transfer from the first register file to the second register file and an operation to be performed using the data, and the load/store circuit is configured to direct a given instruction operation from the issue port to one of the memory execution circuit or the register transfer execution circuit.

Claims (59)

1. A processor, comprising:

a first execution circuit of a first type;

a second execution circuit of a second type different from the first type;

a first register file coupled to the first execution circuit;

a second register file coupled to the second execution circuit; and

a load/store circuit coupled to the first register file and the second register file, the load/store circuit comprising:

an issue port configured to receive an instruction operation for execution;

a memory execution circuit configured to execute memory access instruction operations; and

a register transfer execution circuit configured to execute an instruction operation specifying a transfer of data from the first register file to the second register file and further specifying an additional operation to be performed using the data; and

wherein the load/store circuit is configured to direct a given instruction operation from the issue port to one of the memory execution circuit or the register transfer execution circuit.

2. The processor of claim 1 , wherein the first execution circuit is an integer execution circuit and the second execution circuit is a floating-point execution circuit.

3. The processor of claim 1 , wherein the first execution circuit is a scalar execution circuit and the second execution circuit is a vector execution circuit.

4. The processor of claim 1 , wherein the first register file comprises a general-purpose register file.

5. The processor of claim 1 , wherein:

the additional operation is an integer-to-floating-point conversion operation; and

the register transfer execution circuit includes an integer-to-floating-point conversion circuit.

6. The processor of claim 5 , wherein the second execution circuit includes an additional integer-to-floating point conversion circuit.

7. The processor of claim 1 , wherein:

the additional operation is a duplication operation; and

the register transfer execution circuit includes a duplication circuit configured to read a value from the first register file and copy the value to one or more vector elements stored in the second register file.

8. The processor of claim 1 , further comprising a decoder circuit coupled to the load/store circuit, wherein the decoder circuit is configured to:

receive a fetched transfer instruction involving transfer of data from the first register file to the second register file; and

decode the fetched transfer instruction into an instruction operation for execution by the register transfer execution circuit.

9. The processor of claim 1 , wherein the first register file is not directly accessible by the second execution circuit and the second register file is not directly accessible by the first execution circuit.

10. The processor of claim 1 , wherein the first register file is configured to store values of the first type and the second register file is configured to store values of the second type.

11. The processor of claim 1 , wherein the issue port comprises a reservation station.

12. The processor of claim 1 , further comprising a dispatch circuit configured to issue to the load/store circuit the instruction operation specifying a transfer of data from the first register file to the second register file and further specifying an additional operation to be performed using the data.

13. The processor of claim 12 , wherein the dispatch circuit is configured to issue the instruction operation to the issue port of the load/store circuit.

14. A method, comprising:

detecting, by a processor, an instruction specifying a transfer of data between first and second register files of the processor and further specifying an additional operation to be performed using the data, wherein the first and second register files are coupled to respective first and second execution circuits of the processor, and wherein the first and second execution circuits are of different types;

decoding, by the processor, the instruction into an instruction operation for execution by a register transfer execution circuit in a load/store circuit of the processor;

receiving, by the processor, the instruction operation at the load/store circuit; and

executing, by the processor and using the register transfer execution circuit, the instruction operation.

15. The method of claim 14 , wherein:

specifying the additional operation includes specifying conversion of an integer value from the first register file to a floating-point value; and

the register transfer execution circuit includes an integer-to-floating-point conversion circuit.

16. The method of claim 14 , wherein

specifying the additional operation includes specifying reading of a scalar value from the first register file and copying of the scalar value to one or more vector elements stored in the second register file.

17. The method of claim 14 , further comprising dispatching, by the processor, the instruction operation to an issue port of the load/store circuit.

18. A non-transitory computer readable medium having stored thereon design information that specifies, in a format recognized by a fabrication system that is configured to use the design information, a circuit design for a processor, the processor comprising:

a first execution circuit of a first type;

a second execution circuit of a second type different from the first type;

a first register file coupled to the first execution circuit;

a second register file coupled to the second execution circuit; and

a load/store circuit coupled to the first register file and the second register file, the load/store circuit comprising:

an issue port configured to receive an instruction operation for execution;

a memory execution circuit configured to execute memory access instruction operations; and

a register transfer execution circuit configured to execute an instruction operation specifying a transfer of data from the first register file to the second register file and further specifying an additional operation to be performed using the data; and

wherein the load/store circuit is configured to direct a given instruction operation from the issue port to one of the memory execution circuit or the register transfer execution circuit.

19. The computer readable medium of claim 18 , wherein:

the first execution circuit is an integer execution circuit;

the second execution circuit is a floating-point execution circuit;

the additional operation is an integer-to-floating point conversion operation; and

the register transfer execution circuit includes an integer-to-floating-point conversion circuit.

20. The computer readable medium of claim 18 , wherein:

the first execution circuit is a scalar execution circuit;

the second execution circuit is a vector execution circuit;

the additional operation is a duplication operation; and

the register transfer execution circuit includes a duplication circuit configured to read a value from the first register file and copy the value to one or more vector elements stored in the second register file.

Assignments (2)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Dec 20, 2024
From: BECVAR, MILOS
To: APPLE INC.
Reel/Frame 069654/0063 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded May 18, 2023
From: JIN, ZHAOXIANG; SPADINI, FRANCESCO; SRINIVASA, SKANDA K.
To: APPLE INC.
Reel/Frame 063688/0640 →
Continuity (1)
Provisional Application 63376865 · Sep 23, 2022
References Cited (72)
US 3793631A · Silverstein et al. · 1974 [cited by applicant]
US 5303356A · Vassiliadis · 1994 [cited by applicant]
US 5420992A · Killian · 1995 [cited by applicant]
US 5689695A · Read · 1997 [cited by applicant]
US 5774737A · Nakano · 1998 [cited by applicant]
US 5794063A · Favor · 1998 [cited by applicant]
US 5805486A · Sharangpani · 1998 [cited by examiner]
US 5889984A · Mills · 1999 [cited by examiner]
US 6292888B1 · Nemirovsky et al. · 2001 [cited by applicant]
US 6295599B1 · Hansen et al. · 2001 [cited by applicant]
US 6338136B1 · Col · 2002 [cited by applicant]
US 6560624B1 · Otani et al. · 2003 [cited by applicant]
US 6754810B2 · Elliott · 2004 [cited by examiner]
US 7055022B1 · Col · 2006 [cited by applicant]
US 7818550B2 · Vaden · 2010 [cited by applicant]
US 8078845B2 · Sheffer · 2011 [cited by examiner]
US 8713084B2 · Weinberg · 2014 [cited by applicant]
US 9501286B2 · Col · 2016 [cited by applicant]
US 9747101B2 · Ould-Ahmed-Vall · 2017 [cited by examiner]
US 10324724B2 · Lai et al. · 2019 [cited by applicant]
US 10579389B2 · Lai et al. · 2020 [cited by applicant]
US 20010052063A1 · Tremblay et al. · 2001 [cited by applicant]
US 20020087955A1 · Ronen et al. · 2002 [cited by applicant]
US 20030167460A1 · Desai et al. · 2003 [cited by applicant]
US 20030236966A1 · Samra · 2003 [cited by applicant]
US 20040034757A1 · Gochman · 2004 [cited by applicant]
US 20040128483A1 · Grochowski · 2004 [cited by applicant]
US 20050084099A1 · Montgomery · 2005 [cited by applicant]
US 20050289208A1 · Harrison · 2005 [cited by applicant]
US 20070038844A1 · Valentine · 2007 [cited by applicant]
US 20100115248A1 · OuZiel et al. · 2010 [cited by applicant]
US 20100299505A1 · Uesugi · 2010 [cited by applicant]
US 20110035570A1 · Col · 2011 [cited by applicant]
US 20110264896A1 · Parks · 2011 [cited by examiner]
US 20110264897A1 · Henry · 2011 [cited by applicant]
US 20120144174A1 · Talpes · 2012 [cited by applicant]
US 20130024937A1 · Glew et al. · 2013 [cited by applicant]
US 20130125097A1 · Ebcioglu et al. · 2013 [cited by applicant]
US 20130179664A1 · Olson et al. · 2013 [cited by applicant]
US 20130262841A1 · Gschwind · 2013 [cited by applicant]
US 20140047221A1 · Irwin · 2014 [cited by applicant]
US 20140208073A1 · Blasco-Allue · 2014 [cited by applicant]
US 20140281397A1 · Loktyukhn et al. · 2014 [cited by applicant]
US 20140351561A1 · Parks · 2014 [cited by applicant]
US 20150039851A1 · Uliel · 2015 [cited by applicant]
US 20150089145A1 · Steinmacher-Burow · 2015 [cited by applicant]
US 20160004504A1 · Elmer · 2016 [cited by applicant]
US 20160147290A1 · Williamson et al. · 2016 [cited by applicant]
US 20160179542A1 · Lai · 2016 [cited by applicant]
US 20160291974A1 · Srinivas et al. · 2016 [cited by applicant]
US 20160378487A1 · Ouziel · 2016 [cited by applicant]
US 20170102787A1 · Gu et al. · 2017 [cited by applicant]
US 20170123808A1 · Caulfield · 2017 [cited by applicant]
US 20170177343A1 · Lai · 2017 [cited by applicant]
US 20180129498A1 · Levison et al. · 2018 [cited by applicant]
US 20180129501A1 · Levison · 2018 [cited by applicant]
US 20180267775A1 · Gopal · 2018 [cited by examiner]
US 20180300131A1 · Tannenbaum et al. · 2018 [cited by applicant]
US 20190056943A1 · Gschwind et al. · 2019 [cited by applicant]
US 20190102197A1 · Kumar et al. · 2019 [cited by applicant]
US 20190108023A1 · Lloyd et al. · 2019 [cited by applicant]
US 20200042322A1 · Wang et al. · 2020 [cited by applicant]
US 20200402287A1 · Shah et al. · 2020 [cited by applicant]
US 20210124582A1 · Kerr et al. · 2021 [cited by applicant]
US 20220019436A1 · Lloyd et al. · 2022 [cited by applicant]
US 20220035634A1 · Lloyd · 2022 [cited by applicant]
WO 2019218896A1 · 2019 [cited by applicant]
Office Action in U.S. Appl. No. 17/652,501 mailed Nov. 1, 2023, 47 pages. [cited by applicant]
J. E. Smith, “Future Superscalar Processors Based on Instruction Compounding,” Published 2007, Computer Science, pp. 121-131. [cited by applicant]
Christopher Celio et al., “The Renewed Case for the Reduced Instruction Set Computer: Avoiding ISA Bloat with Macro-Op Fusion for RISC-V,” arXiv:1607.02318v1 [cs.AR] Jul. 8, 2016; 16 pages. [cited by applicant]
Abhishek Deb et al., “SoftHV : A HW/SW Co-designed Processor with Horizontal and Vertical Fusion,” CF'11, May 3-5, 2011, 10 pages. [cited by applicant]
Ian Lee, “Dynamic Instruction Fusion,” UC Santa Cruz Electronic Theses and Dissertations, publication date Dec. 2012, 59 pages. [cited by applicant]