IP Library Granted Patent US 12670121
Granted Patent B2
US 12670121 · App. 18/647,549 · Granted Jun 30, 2026

Graphics processors and graphics processing units having dot product accumulate instruction for hybrid floating point format

Inventors: Subramaniam Maiyuran (Gold River, CA); Shubra Marwaha (Folsom, CA); Ashutosh Garg (Folsom, CA); Supratim Pal (Bangalore, IN); Jorge Parra (El Dorado Hills, CA); Chandra Gurram (Folsom, CA); Varghese George (Folsom, CA); Darin Starkey (Roseville, CA); Guei-Yuan Lueh (San Jose, CA)
Assignee: INTEL CORPORATION
G06F15/7839G06F7/5443G06F7/575G06F7/588G06F9/3001G06F9/30014G06F9/30036G06F9/3004G06F9/30043G06F9/30047G06F9/30065G06F9/30079G06F9/3887G06F9/3888G06F9/5011G06F9/5077G06F12/0215G06F12/0238G06F12/0246G06F12/0607G06F12/0802G06F12/0804G06F12/0811G06F12/0862G06F12/0866G06F12/0871G06F12/0875G06F12/0882G06F12/0888G06F12/0891G06F12/0893G06F12/0895G06F12/0897G06F12/1009G06F12/128G06F13/1626G06F15/8046G06F17/16G06F17/18G06T1/20G06T1/60H03M7/46G06F9/3802G06F9/3818G06F9/3867G06F2212/1008G06F2212/1021G06F2212/1044G06F2212/302G06F2212/401G06F2212/455G06F2212/60G06N3/08G06T15/06
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12670121
App. No.
18/647,549
Filed
Apr 26, 2024
Granted
Jun 30, 2026
Kind
B2
Art Unit
2181
USPC
712/221
Abstract

Graphics processors and graphics processing units having dot product accumulate instructions for a hybrid floating point format are disclosed. In one embodiment, a graphics multiprocessor comprises an instruction unit to dispatch instructions and a processing resource coupled to the instruction unit. The processing resource is configured to receive a dot product accumulate instruction from the instruction unit and to process the dot product accumulate instruction using a bfloat16 number (BF16) format.

Claims (24)

1 . An apparatus comprising:

processing circuitry coupled to a memory, the processing circuitry comprising graphics processing circuitry comprising:

an instruction unit to dispatch instructions; and

a processing resource coupled to the instruction unit, the processing resource to receive an instruction of the dispatched instructions and process the instruction using a hybrid floating point data format.

2 . The apparatus of claim 1 , wherein the hybrid floating point data format is a cross between a first floating point number format and a second floating point number format, wherein the instruction includes a dot product accumulate instruction that causes a second source operand to multiply a third source operand while adding a first source operand with output from multiplying the second source operand and the third source operand.

3 . The apparatus of claim 2 , wherein the first source operand comprises a single-precision floating point format while at least one of the second source operand or the third source operand comprises a bfloat16 number (BF16) format.

4 . The apparatus of claim 2 , wherein at least one of the first source operand or the destination are based on one or more of a half-precision floating point format, a single-precision floating point format, or a bfloat16 number (BF16) format.

5 . The apparatus of claim 1 , wherein the processing resource comprises a floating point unit (FPU) to execute the dot product accumulate instruction based on the BF16 format, wherein the graphics processing circuitry comprises graphics multiprocessing circuitry.

6 . The apparatus of claim 1 , wherein the dispatched instructions comprise single instruction multiple data (SIMD) instructions.

7 . A method comprising:

dispatching, by an instruction unit of processing circuitry of a computing device, instructions; and

receiving, by a processing resource of the processing circuitry, an instruction of the dispatched instructions and processing the instruction using a hybrid floating point data format.

8 . The method of claim 7 , wherein the hybrid floating point data format is a cross between a first floating point number format and a second floating point number format, wherein the instruction includes a dot product accumulate instruction that causes a second source operand to multiply a third source operand while adding a first source operand with output from multiplying the second source operand and the third source operand.

9 . The method of claim 8 , wherein the first source operand comprises a single-precision floating point format while at least one of the second source operand or the third source operand comprises a bfloat16 number (BF16) format.

10 . The method of claim 8 , wherein at least one of the first source operand or the destination are based on one or more of a half-precision floating point format, a single-precision floating point format, or a bfloat16 number (BF16) format.

11 . The method of claim 8 , wherein the processing resource comprises a floating point unit (FPU) to execute the dot product accumulate instruction based on the BF16 format, wherein the processing circuitry is coupled to a memory, wherein the dispatched instructions comprise single instruction multiple data (SIMD) instructions the processing circuitry comprises graphics processing circuitry having graphics multiprocessing circuitry.

12 . A system comprising:

graphics multiprocessor coupled to a memory, the graphics multiprocessor having

an instruction unit to dispatch instructions; and

a processing resource coupled to the instruction unit, the processing resource to receive an instruction of the dispatched instructions and process the instruction using a hybrid floating point data format.

13 . The system of claim 12 , wherein the hybrid floating point data format is a cross between a first floating point number format and a second floating point number format, wherein the instruction includes a dot product accumulate instruction that causes a second source operand to multiply a third source operand while adding a first source operand with output from multiplying the second source operand and the third source operand.

14 . The system of claim 13 , wherein the first source operand comprises a single-precision floating point format while at least one of the second source operand or the third source operand comprises a bfloat16 number (BF16) format.

15 . The system of claim 14 , wherein at least one of the first source operand or the destination are based on one or more of a half-precision floating point format, a single-precision floating point format, or a bfloat16 number (BF16) format.

16 . The system of claim 13 , wherein the processing resource comprises a floating point unit (FPU) to execute the dot product accumulate instruction based on the BF16 format, wherein the dispatched instructions comprise single instruction multiple data (SIMD) instructions, wherein the graphics processing circuitry comprises graphics multiprocessing circuitry.