IP Library Granted Patent US 12,217,020
Granted Patent B2
US 12,217,020 · App. 18/395,836 · Granted Feb 4, 2025

Small multiplier after initial approximation for operations with increasing precision

Inventor: Leonard Rarick (San Diego, CA)
Assignee: Imagination Technologies Limited
G06F7/552G06F7/523G06F7/537G06F7/5525G06F2207/5351
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,217,020
App. No.
18/395,836
Granted
Feb 4, 2025
Kind
B2
Abstract

In an aspect, a processor includes circuitry for iterative refinement approaches, e.g., Newton-Raphson, to evaluating functions, such as square root, reciprocal, and for division. The circuitry includes circuitry for producing an initial approximation; which can include a LookUp Table (LUT). LUT may produce an output that (with implementation-dependent processing) forms an initial approximation of a value, with a number of bits of precision. A limited-precision multiplier multiplies that initial approximation with another value; an output of the limited precision multiplier goes to a full precision multiplier circuit that performs remaining multiplications required for iteration(s) in the particular refinement process being implemented. For example, in division, the output being calculated is for a reciprocal of the divisor. The full-precision multiplier circuit requires a first number of clock cycles to complete, and both the small multiplier and the initial approximation circuitry complete within the first number of clock cycles.

Claims (22)

1. An apparatus for performing an iterative arithmetic operation on an input value to obtain an output value that is used to execute an instruction requesting an arithmetic operation on an operand by said input value, comprising:

initial approximation circuitry configured to provide, based on at least a portion of said input value, an initial approximation of said output value;

limited precision multiplier circuitry configured to obtain a first multiplication result of a first iteration;

full-precision multiplier circuitry configured to obtain a second multiplication result of the first iteration; and

an output configured to output said output value.

2. The apparatus of claim 1 , wherein the initial approximation of the output value has a second number of bits of precision that is less than a first number of bits of precision to which said output value is to be produced.

3. The apparatus of claim 1 , wherein the second multiplication result has a number of bits of precision that is the same as or greater than a first number of bits of precision to which said output value is to be produced.

4. The apparatus of claim 1 , wherein the full-precision multiplier circuitry requires a first number of clock cycles to finish its multiplication, and a combined number of clock cycles required by the initial approximation circuitry to provide the initial approximation and the limited precision multiplier circuitry to complete a multiplication is equal to or less than the first number of clock cycles.

5. The apparatus of claim 1 , wherein a mantissa for the output value is obtained by conducting an iterative refinement of the initial approximation, wherein a first multiplication in the iterative refinement is conducted by the limited precision multiplier circuitry and subsequent multiplications in the iterative refinement are conducted by the full precision multiplier circuitry.

6. The apparatus of claim 1 , wherein the initial approximation circuitry comprises a LookUp Table (LUT) configured to receive at least a portion of bits of the input value and to output a set of bits from which the initial approximation can be constructed.

7. The apparatus of claim 1 , wherein the full-precision multiplier circuitry is configured to perform a double-precision multiplication between two mantissas.

8. The apparatus of claim 1 , wherein the apparatus is configured to provide a value for a division of a dividend a by a divisor b, and the initial approximation circuitry is configured to produce the initial approximation as an initial approximation of a reciprocal of the divisor b.

9. The apparatus of claim 1 , wherein the apparatus is configured to provide a value for a square root of a value b and the initial approximation circuitry is configured to produce the initial approximation as an initial approximation of a reciprocal of the square root of b.

10. The apparatus of claim 1 , wherein the full-precision multiplier circuitry is further configured to feed the second multiplication result back as an input to the full-precision multiplier circuitry to iteratively refine said output value.

11. The apparatus of claim 10 , wherein the initial approximation of the output value has a second number of bits of precision that is less than a first number of bits of precision to which said output value is to be produced.

12. The apparatus of claim 10 , wherein the second multiplication result has a number of bits of precision that is the same as or greater than a first number of bits of precision to which said output value is to be produced.

13. The apparatus of claim 10 , wherein the full-precision multiplier circuitry requires a first number of clock cycles to finish its multiplication, and a combined number of clock cycles required by the initial approximation circuitry to provide the initial approximation and the limited precision multiplier circuitry to complete a multiplication is equal to or less than the first number of clock cycles.

14. The apparatus of claim 10 , wherein a mantissa for the output value is obtained by conducting an iterative refinement of the initial approximation, wherein a first multiplication in the iterative refinement is conducted by the limited precision multiplier circuitry and subsequent multiplications in the iterative refinement are conducted by the full precision multiplier circuitry.

15. The apparatus of claim 10 , wherein the initial approximation circuitry comprises a LookUp Table (LUT) configured to receive at least a portion of bits of the input value and to output a set of bits from which the initial approximation can be constructed.

16. The apparatus of claim 10 , wherein the full-precision multiplier circuitry is configured to perform a double-precision multiplication between two mantissas.

17. The apparatus of claim 10 , wherein the apparatus is configured to provide a value for a division of a dividend a by a divisor b, and the initial approximation circuitry is configured to produce the initial approximation as an initial approximation of a reciprocal of the divisor b.

18. The apparatus of claim 10 , wherein the apparatus is configured to provide a value for a square root of a value b and the initial approximation circuitry is configured to produce the initial approximation as an initial approximation of a reciprocal of the square root of b.

Assignments (1)
SECURITY INTEREST Recorded Jul 31, 2024
From: IMAGINATION TECHNOLOGIES LIMITED
To: FORTRESS INVESTMENT GROUP (UK) LTD
Reel/Frame 068221/0001 →
Continuity (4)
Continuation 18097316 · Jan 16, 2023
Continuation 17209809 · Mar 23, 2021
Division 14516643 · Oct 17, 2014
Related Publication 20240241695A1 · Jul 18, 2024
References Cited (9)
US 5220524A · Hesson · 1993 [cited by applicant]
US 8990278B1 · Clegg · 2015 [cited by applicant]
US 9904512B1 · Pasca · 2018 [cited by applicant]
US 20030149712A1 · Rogenmoser et al. · 2003 [cited by applicant]
US 20040024806A1 · Jeong et al. · 2004 [cited by applicant]
Gallagher et al., “Error-Correcting Goldschmidt Dividers Using Time Shared TMR,” 1998 IEEE Int'l Symposium on Defect and Fault Tolerance in VLSI Systems, pp. 226-229, 1998. [cited by applicant]
Gallagher et al., “Fault Tolerant Newton-Raphson Dividers Using Time Shared TMR,” 1998 IEEE Int'l Symposium on Defect and Fault Tolerance in VLSI Systems, pp. 241-243, 1998. [cited by applicant]
Ito et al., “Efficient Initial Approximation for Multiplicative Division and Square Root by a Multiplication with Operand Modification,” IEEE Transactions on Computers, vol. 46, No. 4, pp. 405-498, 1997. [cited by applicant]
Wang et al., “A Decimal Floating-Point Divider Using Newton-Raphson Iteration,” The Journal of VLSI Signal Processing, vol. 49, pp. 3-8. [cited by applicant]