IP Library › Granted Patent US 12,217,060
Granted Patent B1
US 12,217,060 · App. 18/176,457 · Granted Feb 4, 2025

Instruction fusion

Inventors: Francesco Spadini (Sunset Valley, TX); Skanda K. Srinivasa (Austin, TX); Reena Panda (Cedar Park, TX); Brian T. Mokrzycki (Portland, OR); Haoyan Jia (Ellicott City, MD); Zhaoxiang Jin (Austin, TX)
Assignee: Apple Inc.
G06F9/30145G06F9/3001G06F9/30181
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,217,060
App. No.
18/176,457
Granted
Feb 4, 2025
Kind
B1
Abstract

Techniques are disclosed that relate to executing pairs of instructions. A processor may include fusion detector circuitry configured to detect a pair of fetched instructions and fuse the pair of fetched instructions into a fused instruction operation, and execution circuitry coupled to the fusion detector circuitry and configured to execute the fused instruction operation. In some embodiments the pair of instructions is executable to generate a remainder of a division operation. In some embodiments the pair of instructions is executable to compare two operands and perform a write operation based on the comparison. In some embodiments the pair of instructions is executable to perform an operation and apply a mask bit sequence to the result. The fusion detector circuitry may also be configured to obtain first and second portions of a constant value from first and second instructions and store the first and second portions in a destination register.

Claims (85)

1. A processor, comprising:

fusion detector circuitry configured to:

receive fetched instructions;

detect a first pair of the fetched instructions, wherein the first pair includes:

a first instruction that is executable to:

perform a divide operation using a dividend from a first source register and a divisor from a second source register; and

write a quotient of the divide operation to a first destination register; and

a second instruction that is executable to:

read the quotient, the dividend and the divisor from the first destination register, the first source register and the second source register;

calculate a remainder of the divide operation; and

write the remainder to the first destination register, overwriting the quotient; and

fuse the first pair of the fetched instructions into a first fused instruction operation that is executable to use the dividend and the divisor to calculate the remainder and write the remainder instead of the quotient to the first destination register; and

execution circuitry coupled to the fusion detector circuitry and configured to execute the first fused instruction operation.

2. The processor of claim 1 , wherein the execution circuitry comprises:

a divider circuit configured to generate a set of residual values related to the remainder; and

a conversion circuit configured to convert the set of residual values into the remainder.

3. The processor of claim 1 , wherein the execution circuitry is configured to execute the first fused instruction operation without performing a multiplication operation to calculate the remainder.

4. The processor of claim 1 , wherein the second instruction is a multiply-subtract instruction that is executable to perform a multiplication of a pair of operands and to subtract a result of the multiplication from another operand, and wherein the first instruction is coded to supply the divisor and the quotient as the pair of operands to be multiplied and to supply the dividend as the other operand.

5. The processor of claim 1 , wherein

the fusion detector circuitry is further configured to:

detect a second pair of the fetched instructions, wherein the second pair is executable to write, to a second destination register, a specified portion of an arithmetic/logic operation result, and wherein the second pair includes:

a first instruction that is executable to perform an arithmetic/logic operation to produce the arithmetic/logic operation result and write the arithmetic/logic operation result to the second destination register; and

a second instruction that is executable to perform a logical AND operation of the arithmetic/logic operation result with a specified mask bit sequence and write a result of the logical AND operation to the second destination register; and

fuse the second pair of the fetched instructions into a second fused instruction operation that is executable to perform the arithmetic/logic operation and write to the second destination register the specified portion, corresponding to the specified mask bit sequence, of the arithmetic/logic operation result; and

the execution circuitry is further configured to execute the second fused instruction operation.

6. The processor of claim 5 , wherein the execution circuitry is further configured to generate the specified mask bit sequence.

7. The processor of claim 1 , wherein

the fusion detector circuitry is further configured to:

detect a second pair of the fetched instructions, wherein the second pair includes:

a first instruction that is executable to:

perform a comparison of a first operand to a second operand; and

write to one or more bits of a status register based on a result of the comparison; and

a second instruction that is executable to write a value to a second destination register based on the first operand, the second operand, and bit values of the one or more bits of the status register; and

fuse the second pair of the fetched instructions into a second fused instruction operation that is executable to perform the comparison of the first operand to the second operand and write to the second destination register based on the result of the comparison; and

the execution circuitry is further configured to execute the second fused instruction operation.

8. The processor of claim 7 , wherein the second instruction of the second pair is executable to store either the first operand or the second operand in the second destination register, based on the result of the comparison of the first operand and the second operand.

9. The processor of claim 7 , wherein the second instruction of the second pair is executable to store either a value of “0” or a value of “1” in the second destination register, based on the result of the comparison of the first operand and the second operand.

10. The processor of claim 1 , wherein the fusion detector circuitry is further configured to:

detect a second pair of the fetched instructions, wherein the second pair is executable to store into a second destination register a constant value having a bit length larger than a width of an immediate value field of a first instruction or a second instruction of the second pair;

perform a register storage operation executable to:

obtain a first portion of the constant value from the first instruction of the second pair and a second portion of the constant value from the second instruction of the second pair; and

store the first and second portions of the constant value in corresponding first and second portions of the second destination register; and

prevent instruction operations corresponding to the first instruction of the second pair and the second instruction of the second pair from being dispatched to an execution pipeline of the processor.

11. The processor of claim 10 , wherein:

the first instruction of the second pair is one of:

a move/zero instruction that is executable to write the first portion of the constant value to the first portion of the second destination register and write zeros to the second portion of the second destination register;

a move/negate instruction that is executable to write the first portion of the constant value to the first portion of the second destination register and write ones to the second portion of the second destination register;

a logical OR instruction that is executable to perform a bitwise OR operation of the first portion of the constant value with a source register filled with zeros and write a result of the bitwise OR operation to the first portion of the second destination register; or

a logical XOR instruction that is executable to perform a bitwise exclusive OR operation of the first portion of the constant value with a source register filled with zeros and write a result of the bitwise exclusive OR operation to the first portion of the second destination register; and

the second instruction of the second pair is a move/keep instruction that is executable to write the second portion of the constant value to the second portion of the second destination register without changing bit values in the first portion of the second destination register.

12. The processor of claim 10 , wherein:

the first instruction of the second pair is executable to calculate a first address of a target page in memory and write the first address to the second destination register; and

the second instruction of the second pair is executable to add an offset value to the first address to form a second address and write the second address to the second destination register.

13. A method, comprising:

detecting, by a processor, a first instruction of a first pair of instructions, wherein the first instruction of the first pair is executable by the processor to:

perform a divide operation using a dividend from a first source register and a divisor from a second source register, and

write a quotient of the divide operation to a first destination register;

detecting, by the processor, a second instruction of the first pair of instructions, wherein the second instruction of the first pair is executable by the processor to:

read the quotient, dividend and divisor from the first destination register, first source register and second source register, respectively;

calculate a remainder of the divide operation; and

write the remainder to the first destination register, overwriting the quotient;

fusing, by the processor, the first pair of instructions into a first fused instruction operation that is executable by the processor to:

use the dividend and the divisor to calculate the remainder; and

write the remainder instead of the quotient to the first destination register; and

executing, by the processor, the first fused instruction operation.

14. The method of claim 13 , further comprising:

detecting, by the processor, a first instruction of a second pair of instructions, wherein the first instruction of the second pair is executable by the processor to:

perform an arithmetic/logic operation to produce an arithmetic/logic operation result; and

write the arithmetic/logic operation result to a second destination register;

detecting, by the processor, a second instruction of the second pair of instructions, wherein the second instruction of the second pair is executable by the processor to:

perform a logical AND operation of the arithmetic/logic operation result with a specified mask bit sequence; and

write a result of the logical AND operation to the second destination register;

fusing, by the processor, the second pair of instructions into a second fused instruction operation that is executable to perform the arithmetic/logic operation and write to the second destination register a portion, corresponding to the specified mask bit sequence, of the arithmetic/logic operation result; and

executing, by the processor, the second fused instruction operation.

15. The method of claim 13 , further comprising:

detecting, by the processor, a first instruction of a second pair of instructions, wherein the first instruction of the second pair is executable by the processor to perform a comparison of a first operand to a second operand and write to one or more bits of a status register based on a result of the comparison;

detecting, by the processor, a second instruction of the second pair of instructions, wherein the second instruction of the second pair is executable by the processor to write a value to a second destination register based on the first operand, the second operand, and bit values of the one or more bits of the status register;

fusing, by the processor, the second pair of instructions into a second fused instruction operation that is executable by the processor to perform the comparison of the first operand to the second operand and write to the second destination register based on the result of the comparison; and

executing, by the processor, the second fused instruction operation.

16. The method of claim 13 , further comprising:

detecting, by the processor, a second pair of instructions executable to store into a second destination register a constant value having a bit length larger than a width of an immediate value field of a first instruction or a second instruction of the second pair of instructions;

obtaining a first portion of the constant value from the first instruction of the second pair of instructions;

obtaining a second portion of the constant value from the second instruction of the second pair of instructions;

storing the first and second portions of the constant value in corresponding first and second portions of the second destination register; and

preventing instruction operations corresponding to the first instruction and second instruction of the second pair of instructions from being dispatched to an execution pipeline of the processor.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Feb 28, 2023
From: SPADINI, FRANCESCO; SRINIVASA, SKANDA K.; PANDA, REENA; MOKRZYCKI, BRIAN T.; JIA, HAOYAN; JIN, ZHAOXIANG
To: APPLE INC.
Reel/Frame 062834/0244 →
Continuity (1)
Provisional Application 63376822 · Sep 23, 2022
References Cited (66)
US 5303356A · Vassiliadis · 1994 [cited by examiner]
US 5420992A · Killian · 1995 [cited by applicant]
US 5689695A · Read · 1997 [cited by applicant]
US 5774737A · Nakano · 1998 [cited by applicant]
US 5794063A · Favor · 1998 [cited by applicant]
US 5805486A · Sharangpani · 1998 [cited by applicant]
US 5889984A · Mills · 1999 [cited by applicant]
US 6292888B1 · Nemirovsky · 2001 [cited by applicant]
US 6295599B1 · Hansen et al. · 2001 [cited by applicant]
US 6338136B1 · Col · 2002 [cited by applicant]
US 6560624B1 · Otani et al. · 2003 [cited by applicant]
US 6754810B2 · Elliott et al. · 2004 [cited by applicant]
US 7055022B1 · Col · 2006 [cited by applicant]
US 7818550B2 · Vaden · 2010 [cited by examiner]
US 8078845B2 · Sheffer et al. · 2011 [cited by applicant]
US 8713084B2 · Weinberg · 2014 [cited by examiner]
US 9501286B2 · Col · 2016 [cited by applicant]
US 9747101B2 · Ould-Ahmed-Vall et al. · 2017 [cited by applicant]
US 10324724B2 · Lai · 2019 [cited by applicant]
US 10579389B2 · Caulfield · 2020 [cited by applicant]
US 20030236966A1 · Samra · 2003 [cited by applicant]
US 20040034757A1 · Gochman · 2004 [cited by applicant]
US 20040128483A1 · Grochowski · 2004 [cited by applicant]
US 20050084099A1 · Montgomery · 2005 [cited by applicant]
US 20050289208A1 · Harrison · 2005 [cited by examiner]
US 20070038844A1 · Valentine · 2007 [cited by applicant]
US 20100115248A1 · Ouziel · 2010 [cited by applicant]
US 20110035570A1 · Col · 2011 [cited by applicant]
US 20110264896A1 · Parks · 2011 [cited by applicant]
US 20110264897A1 · Henry · 2011 [cited by applicant]
US 20120144174A1 · Talpes · 2012 [cited by examiner]
US 20130024937A1 · Glew · 2013 [cited by applicant]
US 20130125097A1 · Ebcioglu et al. · 2013 [cited by applicant]
US 20130179664A1 · Olson et al. · 2013 [cited by applicant]
US 20130262841A1 · Gschwind · 2013 [cited by applicant]
US 20140047221A1 · Irwin · 2014 [cited by applicant]
US 20140208073A1 · Blasco-Allue · 2014 [cited by applicant]
US 20140281397A1 · Loktyukhin · 2014 [cited by applicant]
US 20140351561A1 · Parks · 2014 [cited by applicant]
US 20150039851A1 · Uliel · 2015 [cited by applicant]
US 20150089145A1 · Steinmacher-Burow · 2015 [cited by applicant]
US 20160004504A1 · Elmer · 2016 [cited by applicant]
US 20160179542A1 · Lai · 2016 [cited by applicant]
US 20160291974A1 · Lingam · 2016 [cited by applicant]
US 20160378487A1 · Ouziel · 2016 [cited by examiner]
US 20170102787A1 · Gu · 2017 [cited by applicant]
US 20170123808A1 · Caulfield · 2017 [cited by examiner]
US 20170177343A1 · Lai · 2017 [cited by applicant]
US 20180129498A1 · Levison · 2018 [cited by applicant]
US 20180129501A1 · Levison · 2018 [cited by applicant]
US 20180267775A1 · Gopal · 2018 [cited by applicant]
US 20180300131A1 · Tannenbaum · 2018 [cited by applicant]
US 20190056943A1 · Gschwind · 2019 [cited by applicant]
US 20190102197A1 · Kumar et al. · 2019 [cited by applicant]
US 20190108023A1 · Lloyd · 2019 [cited by applicant]
US 20200042322A1 · Wang · 2020 [cited by applicant]
US 20200402287A1 · Shah · 2020 [cited by applicant]
US 20210124582A1 · Kerr · 2021 [cited by applicant]
US 20220019436A1 · Lloyd · 2022 [cited by applicant]
US 20220035634A1 · Lloyd · 2022 [cited by applicant]
WO 2019218896A1 · 2019 [cited by applicant]
Office Action in U.S. Appl. No. 17/652,501 mailed Nov. 1, 2023, 47 pages. [cited by applicant]
J. E. Smith, “Future Superscalar Processors Based On Instruction Compounding,” Published 2007, Computer Science, pp. 121-131. [cited by applicant]
Christopher Celio et al., “The Renewed Case for the Reduced Instruction Set Computer: Avoiding ISA Bloat with Macro-Op Fusion for RISC-V,” arXiv:1607.02318v1 [cs.AR] Jul. 8, 2016; 16 pages. [cited by applicant]
Abhishek Deb et al., “SoftHV : A Hw/Sw Co-designed Processor with Horizontal and Vertical Fusion,” CF 11, May 3-5, 2011, 10 pages. [cited by applicant]
Ian Lee, “Dynamic Instruction Fusion,” UC Santa Cruz Electronic Theses and Dissertations, publication date Dec. 2012, 59 pages. [cited by applicant]