IP Library Granted Patent US 8,046,399
Granted Patent B1
US 8,046,399 · App. 12/020,486 · Granted Oct 25, 2011

Fused multiply-add rounding and unfused multiply-add rounding in a single multiply-add module

Assignee: Oracle America, Inc.
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 8,046,399
App. No.
12/020,486
Granted
Oct 25, 2011
Kind
B1
Abstract

A computer processor including a single fused-unfused floating point multiply-add (FMA) module computes the result of the operation A*B+C for floating point numbers for fused multiply-add rounding operations and unfused multiply-add rounding operations. In one embodiment, a fused multiply-add rounding implementation is augmented with additional hardware which calculates an unfused multiply-add rounding result without adding additional pipeline stages. In one embodiment, a computation by the fused-unfused floating point multiply-add (FMA) module is initiated using a single opcode which determines whether a fused multiply-add rounding result or unfused multiply-add rounding result is generated.

Claims (36)

1. A computer system comprising:

a memory; and

a processor coupled to said memory, said processor comprising:

a fused-unfused floating point multiply-add (FMA) module, said fused-unfused FMA module for receiving a first multiply term, a second multiply term, and an addition term,

wherein in response to said processor receiving a single unfused multiply-add opcode, said fused-unfused FMA module generates an unfused multiply-add rounding result,

wherein in response to said processor receiving a single fused multiply-add opcode, said fused-unfused FMA module generates a fused multiply-add rounding result,

wherein said unfused multiply-add rounding result is a rounded sum of said addition term with a rounded value of the product of said first multiply term and said second multiply term without obtaining the product of said first multiply term and said second multiply term,

wherein in generating said unfused multiply-add rounding result, instead of generating said product of said first multiply term and said second multiply term, a first terminal partial product and a second terminal partial product are generated from said first multiply term and said second multiply term, and

further wherein each said first terminal partial product and said second terminal partial product are truncated to produce, respectively, a truncated first terminal partial product and a truncated second terminal partial product, before being combined with said addition term.

2. The computer system of claim 1 wherein said fused multiply-add rounding result is a rounded sum of said addition term with a value of the product of said first multiply term and said second multiply term without obtaining the product of said first multiply term and said second multiply term.

3. The computer system of claim 1 wherein said addition term is aligned to generate an aligned addition term before being combined with said truncated first terminal partial product and said truncated second terminal partial product.

4. The computer system of claim 3 wherein said aligned addition term is combined with said truncated first terminal partial product and said truncated second terminal partial product to generate a carry look-ahead value.

5. The computer system of claim 4 wherein one or more replacement bit values are generated, wherein said one or more replacement bit values are obtained from one or more bit values of said first multiply term and said second multiply term and said aligned addition term.

6. The computer system of claim 5 wherein said aligned addition term is added with said truncated first terminal partial product and said truncated second terminal partial product, and rounded to produce said unfused multiply-add rounding result when selected bits of said sum are replaced with said replacement value.

7. A computer processor comprising:

a fused-unfused floating point multiply-add (FMA) module, said fused-unfused FMA module for receiving a first multiply term, a second multiply term, and an addition term,

wherein in response to said processor receiving a single unfused multiply-add opcode, said fused-unfused FMA module generates an unfused multiply-add rounding result,

wherein in response to said processor receiving a single fused multiply-add opcode, said fused-unfused FMA module generates a fused multiply-add rounding result,

wherein said unfused multiply-add rounding result is a rounded sum of said addition term with a rounded value of the product of said first multiply term and said second multiply term without obtaining the product of said first multiply term and said second multiply term,

wherein in generating said unfused multiply-add rounding result, instead of generating said product of said first multiply term and said second multiply term, a first terminal partial product and a second terminal partial product are generated from said first multiply term and said second multiply term, and

further wherein each said first terminal partial product and said second terminal partial product are truncated to produce respectively a truncated first terminal partial product and a truncated second terminal partial product before being combined with said addition term.

8. The computer processor of claim 7 wherein said fused multiply-add rounding result is a rounded sum of said addition term with a value of the product of said first multiply term and said second multiply term without obtaining the product of said first multiply term and said second multiply term.

9. The computer processor of claim 7 wherein said addition term is aligned to generate an aligned addition term before being combined with said truncated first terminal partial product and said truncated second terminal partial product.

10. The computer processor of claim 9 wherein said aligned addition term is combined with said truncated first terminal partial product and said truncated second terminal partial product to generate a carry look-ahead value.

11. The computer processor of claim 10 wherein one or more replacement bit values are generated, wherein said one or more replacement bit values are obtained from one or more bit values of said first multiply term and said second multiply term and said aligned addition term.

12. The computer processor of claim 11 wherein said aligned addition term is added with said truncated first terminal partial product and said truncated second terminal partial product, and rounded to produce said unfused multiply-add rounding result when selected bits of said sum are replaced with said replacement value.

13. A computer system comprising:

a memory; and

a processor coupled to said memory, said processor comprising:

a fused-unfused floating point multiply-add (FMA) module, said fused-unfused FMA module, said fused-unfused FMA module comprising:

means for receiving a first multiply term, a second multiply term, and an addition term; and

means for generating an unfused multiply-add rounding result of said first multiply term, said second multiply term, and said addition term, when in an unfused multiply-add rounding mode, and for generating a fused multiply-add result of said first multiply term, said second multiply term and said addition term, when in a fused multiply-add rounding mode,

wherein said unfused multiply-add rounding result is a rounded sum of said addition term with a rounded value of the product of said first multiply term and said second multiply term without obtaining the product of said first multiply term and said second multiply term,

wherein in generating said unfused multiply-add rounding result, instead of generating said product of said first multiply term and said second multiply term, a first terminal partial product and a second terminal partial product are generated from said first multiply term and said second multiply term, and

further wherein each said first terminal partial product and said second terminal partial product are truncated to produce, respectively, a truncated first terminal partial product and a truncated second terminal partial product, before being combined with said addition term.

14. The computer system of claim 13 wherein said fused multiply-add rounding result is a rounded sum of said addition term with a value of the product of said first multiply term and said second multiply term without obtaining the product of said first multiply term and said second multiply term.

Assignments (2)
CHANGE OF NAME Recorded Sep 1, 2011
From: SUN MICROSYSTEMS, INC.
To: ORACLE AMERICA, INC.
Reel/Frame 026843/0357 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jan 25, 2008
From: INAGANTI, MURALI K.; RARICK, LEONARD D.
To: SUN MICROSYSTEMS, INC.
Reel/Frame 020419/0491 →