IP Library Granted Patent US 12,379,927
Granted Patent B2
US 12,379,927 · App. 17/463,382 · Granted Aug 5, 2025

BFLOAT16 scale and/or reduce instructions

Inventors: Menachem Adelman (Haifa, IL); Alexander Heinecke (San Jose, CA); Robert Valentine (Kiryat Tivon, IL); Zeev Sperber (Zikhron Yaakov, IL); Amit Gradstein (Binyamina, IL); Mark Charney (Lexington, MA); Evangelos Georganas (San Mateo, CA); Dhiraj Kalamkar (Bangalore, IN); Christopher Hughes (Santa Clara, CA); Cristina Anderson (Hillsboro, OR)
Assignee: Intel Corporation
G06F9/30145G06F9/30036G06F9/30038G06F9/30101G06F9/30014
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,379,927
App. No.
17/463,382
Granted
Aug 5, 2025
Kind
B2
Abstract

Techniques for scale and reduction of BF16 data elements are described. An exemplary instruction includes fields for an having fields for an opcode, an identification of a location of a first packed data source operand, an identification of a location of a second packed data source operand, and an identification of a packed data destination operand, wherein the opcode is to indicate that execution circuitry is to perform, for each data element position of the packed data source operands, a floating point scale operation of a BF16 data element of the first packed data source by multiplying the data element by a power of 2 value, wherein a value of the exponent of the power of 2 value is a floor value of a BF16 data element of the second packed data source, and store a result of the floating point scale operation into a corresponding data element position of the packed data destination operand.

Claims (48)

1. An apparatus comprising:

decode circuitry to decode an instance of a single instruction, the single instruction to include fields for an opcode, an identification of a location of a first packed data source operand, an identification of a location of a second packed data source operand, and an identification of a packed data destination operand, wherein the opcode is to indicate that execution circuitry is to perform, for each data element position of the packed data source operands, a floating point scale operation of a BF16 data element of the first packed data source by multiplying the data element by a power of 2 value, wherein a value of an exponent of the power of 2 value is a floor value of a BF16 data element of the second packed data source, and store a result of the floating point scale operation into a corresponding data element position of the packed data destination operand; and

the execution circuitry to execute the decoded instruction according to the opcode.

2. The apparatus of claim 1 , wherein the field for the identification of the first source operand is to identify a vector register.

3. The apparatus of claim 1 , wherein the field for the identification of the first source operand is to identify a memory location.

4. The apparatus of claim 1 , wherein the execution circuitry is to use a round to nearest even rounding mode during execution of the decoded instruction.

5. The apparatus of claim 1 , wherein the floor value is a zero when the data element of the second packed data source is a denormal.

6. The apparatus of claim 1 , wherein the data element of the first packed data source is a zero when the data element of the first packed data source is a denormal.

7. The apparatus of claim 1 , wherein the instruction is to further include one or more fields for a writemask register.

8. A system comprising:

memory to store an instance of a single instruction;

decode circuitry to decode the instance of the single instruction, the single instruction to include fields for an opcode, an identification of a location of a first packed data source operand, an identification of a location of a second packed data source operand, and an identification of a packed data destination operand, wherein the opcode is to indicate that execution circuitry is to perform, for each data element position of the packed data source operands, a floating point scale operation of a BF16 data element of the first packed data source by multiplying the data element by a power of 2 value, wherein a value of an exponent of the power of 2 value is a floor value of a BF16 data element of the second packed data source, and store a result of the floating point scale operation into a corresponding data element position of the packed data destination operand; and

the execution circuitry to execute the decoded instruction according to the opcode.

9. The system of claim 8 , wherein the field for the identification of the first source operand is to identify a vector register.

10. The system of claim 8 , wherein the field for the identification of the first source operand is to identify a memory location.

11. The system of claim 8 , wherein the execution circuitry is to use a round to nearest even rounding mode during execution of the decoded instruction.

12. The system of claim 8 , wherein the floor value is a zero when the data element of the second packed data source is a denormal.

13. The system of claim 8 , wherein the instruction is to further include one or more fields for a writemask register.

14. The system of claim 8 , wherein the data element of the first packed data source is a zero when the data element of the first packed data source is a denormal.

15. A non-transitory machine-readable medium storing at least an instance of a particular single instruction, wherein the instance of the particular single instruction is to be processed by a processor by performing a method comprising:

decoding the instance of the single instruction, the single instruction having fields for an opcode, an identification of a location of a first packed data source operand, an identification of a location of a second packed data source operand, and an identification of a packed data destination operand, wherein the opcode is to indicate that execution circuitry is to perform, for each data element position of the packed data source operands, a floating point scale operation of a BF16 data element of the first packed data source by multiplying the data element by a power of 2 value, wherein a value of an exponent of the power of 2 value is a floor value of a BF16 data element of the second packed data source, and store a result of the floating point scale operation into a corresponding data element position of the packed data destination operand; and

executing the decoded instruction according to the opcode.

16. The non-transitory machine-readable medium of claim 15 , wherein the field for the identification of the first source operand is to identify a vector register.

17. The non-transitory machine-readable medium of claim 15 , wherein the field for the identification of the first source operand is to identify a memory location.

18. The non-transitory machine-readable medium of claim 15 , wherein the executing is to use a round to nearest even rounding mode during execution of the decoded instruction.

19. The non-transitory machine-readable medium of claim 15 , wherein the floor value is a zero when the data element of the second packed data source is a denormal.

20. The non-transitory machine-readable medium of claim 15 , wherein the instruction is to further include one or more fields for a writemask register.

21. The non-transitory machine-readable medium of claim 15 , wherein the data element of the first packed data source is a zero when the data element of the first packed data source is a denormal.

22. A non-transitory machine-readable medium storing at least an instance of a particular single instruction, wherein the instance of the particular single instruction is to be processed by a processor by performing a method comprising:

translating the particular single instruction from a first instruction set architecture to one or more instructions of a second, different instruction set architecture, the particular single instruction having fields for an opcode, an identification of a location of a first packed data source operand, an identification of a location of a second packed data source operand, and an identification of a packed data destination operand, wherein the opcode is to indicate that execution circuitry is to perform, for each data element position of the packed data source operands, a floating point scale operation of a BF16 data element of the first packed data source by multiplying the data element by a power of 2 value, wherein a value of an exponent of the power of 2 value is a floor value of a BF16 data element of the second packed data source, and store a result of the floating point scale operation into a corresponding data element position of the packed data destination operand;

decoding the one or more instructions of a second, different instruction set architecture;

executing the decoded one or more instructions of a second, different instruction set architecture.

23. The non-transitory machine-readable medium of claim 22 , wherein the field for the identification of the first source operand is to identify a vector register.

24. The non-transitory machine-readable medium of claim 22 , wherein the field for the identification of the first source operand is to identify a memory location.

25. The non-transitory machine-readable medium of claim 22 , wherein the executing is to use a round to nearest even rounding mode during execution of the decoded instruction.

26. The non-transitory machine-readable medium of claim 22 , wherein the floor value is a zero when the data element of the second packed data source is a denormal.

27. The non-transitory machine-readable medium of claim 22 , wherein the data element of the first packed data source is a zero when the data element of the first packed data source is a denormal.

28. The non-transitory machine-readable medium of claim 22 , wherein the instruction is to further include one or more fields for a writemask register.

29. A method comprising:

translating a particular single instruction from a first instruction set architecture to one or more instructions of a second, different instruction set architecture, the particular single instruction to include fields for an having fields for an opcode, an identification of a location of a first packed data source operand, an identification of a location of a second packed data source operand, and an identification of a packed data destination operand, wherein the opcode is to indicate that execution circuitry is to perform, for each data element position of the packed data source operands, a floating point scale operation of a BF16 data element of the first packed data source by multiplying the data element by a power of 2 value, wherein a value of an exponent of the power of 2 value is a floor value of a BF16 data element of the second packed data source, and store a result of the floating point scale operation into a corresponding data element position of the packed data destination operand;

decoding the one or more instructions of a second, different instruction set architecture;

executing the decoded one or more instructions of a second, different instruction set architecture.

30. The method of claim 29 , wherein the field for the identification of the first source operand is to identify a vector register.

31. The method of claim 29 , wherein the field for the identification of the first source operand is to identify a memory location.

32. The method of claim 29 , wherein the executing is to use a round to nearest even rounding mode during execution of the decoded instruction.

33. The method of claim 29 , wherein the floor value is a zero when the data element of the second packed data source is a denormal.

34. The method of claim 29 , wherein the instruction is to further include one or more fields for a writemask register.

35. The method of claim 29 , wherein the data element of the first packed data source is a zero when the data element of the first packed data source is a denormal.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jan 6, 2022
From: ADELMAN, MENACHEM; HEINECKE, ALEXANDER; VALENTINE, ROBERT; SPERBER, ZEEV; GRADSTEIN, AMIT; CHARNEY, MARK; GEORGANAS, EVANGELOS; KALAMKAR, DHIRAJ; HUGHES, CHRISTOPHER; ANDERSON, CRISTINA
To: INTEL CORPORATION
Reel/Frame 058584/0172 →
Continuity (1)
Related Publication 20230068781A1 · Mar 2, 2023
References Cited (42)
US 4815021A · Steiner et al. · 1989 [cited by applicant]
US 9448765B2 · Anderson · 2016 [cited by examiner]
US 9606770B2 · Anderson · 2017 [cited by examiner]
US 11366663B2 · Heinecke · 2022 [cited by examiner]
US 11372643B2 · Heinecke · 2022 [cited by examiner]
US 20030018676A1 · Shaw · 2003 [cited by applicant]
US 20110047358A1 · Eichenberger et al. · 2011 [cited by applicant]
US 20140067894A1 · Plondke et al. · 2014 [cited by applicant]
US 20150088946A1 · Anderson · 2015 [cited by examiner]
US 20160224512A1 · Moudgill et al. · 2016 [cited by applicant]
US 20190079768A1 · Heinecke et al. · 2019 [cited by applicant]
US 20190220278A1 · Adelman · 2019 [cited by examiner]
US 20190384575A1 · Hickmann et al. · 2019 [cited by applicant]
US 20200184309A1 · Patel · 2020 [cited by applicant]
US 20200371794A1 · Zbiciak et al. · 2020 [cited by applicant]
US 20200371805A1 · Lutz · 2020 [cited by applicant]
US 20210117194A1 · Heinecke et al. · 2021 [cited by applicant]
US 20210157589A1 · Heinecke et al. · 2021 [cited by applicant]
US 20220121727A1 · Hong et al. · 2022 [cited by applicant]
US 20220413805A1 · Li · 2022 [cited by examiner]
US 20230068781A1 · Adelman · 2023 [cited by examiner]
US 20240045682A1 · Heinecke · 2024 [cited by examiner]
CN 112199119A · 2021 [cited by applicant]
GB 2186105A · 1987 [cited by applicant]
WO 2008036946A1 · 2008 [cited by applicant]
WO 2020190814A1 · 2020 [cited by applicant]
European Search Report and Search Opinion, EP App. No. 22185939.0, Jan. 20, 2023, 9 pages. [cited by applicant]
European Search Report and Search Opinion, EP App. No. 22185990.3, Jan. 20, 2023, 10 pages. [cited by applicant]
Extended European Search Report and Search Opinion for Application No. 22183762.8, Dec. 21, 2022, 9 pages. [cited by applicant]
Intel, “Intel (registered) 64 and IA-32 Architectures Software Developer's Manual”, vol. 2 (2A, 2B & 2C): Instruction Set Reference, A-Z, Order No. 325383-055US, Jun. 2015, 1011 pages. [cited by applicant]
Intel, “Intel® Architecture Instruction Set Extensions Programming Reference”, Order No. 319433-023, Aug. 1, 2015, 1178 pages. [cited by applicant]
Office Action , EP App. No. 22185939.0 , Aug. 22, 2023, 05 pages. [cited by applicant]
Office Action , EP App. No. 22185990.3, Aug. 22, 2023, 05 pages. [cited by applicant]
Wikipedia, “bfloat16 floating-point format”, available online at <https://en.wikipedia.org/w/index.php?title=Bfloat16_floating-point_format&oldid=1033549871>, Jul. 14, 2021, 4 pages. [cited by applicant]
European Search Report and Search Opinion, EP App. No. 22188067.7, Jan. 25, 2023, 12 pages. [cited by applicant]
European Search Report and Search Opinion, EP App. No. 22188069.3, Jan. 25, 2023, 10 pages. [cited by applicant]
European Search Report and Search Opinion, EP App. No. 22188079.2, Jan. 25, 2023, 10 pages. [cited by applicant]
Notification of Publication of Patent Application for Invention, CN App. No. 202210866252.7, Mar. 8, 2023, 3 pages (1 page of English Translation and 2 pages of Original Document). [cited by applicant]
Decision to grant, EP App. No. 22183762.8, Feb. 15, 2024, 2 pages. [cited by applicant]
Intention to grant, EP App. No. 22183762.8, Oct. 11, 2023, 7 pages. [cited by applicant]
Intel, “Intel® Architecture Instruction Set Extensions and Future Features Programming Reference”, Reference No. 319433-032, Jan. 2018, 137 pages. [cited by applicant]
Notification of Oral Proceeding, EP App. No. 22185990.3, Mar. 22, 2024, 9 pages. [cited by applicant]