IP Library › Granted Patent US 12,541,365
Granted Patent B2
US 12,541,365 · App. 18/925,482 · Granted Feb 3, 2026

Systems and methods for performing instructions to convert to 16-bit floating-point format

Inventors: Alexander F. Heinecke (San Jose, CA); Robert Valentine (Kiryat Tivon, IL); Mark J. Charney (Lexington, MA); Raanan Sade (Portland, OR); Menachem Adelman (Modi'in, IL); Zeev Sperber (Zichron Yackov, IL); Amit Gradstein (Binyamina, IL); Simon Rubanovich (Haifa, IL)
Assignee: Intel Corporation
G06F9/30025G06F9/30014G06F9/30018G06F9/30036G06F9/30038G06F9/30105G06F9/3802G06F9/3818G06F9/384G06F9/3887G06F9/3888
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,541,365
App. No.
18/925,482
Granted
Feb 3, 2026
Kind
B2
Abstract

Disclosed embodiments relate to systems and methods for performing instructions to convert to 16-bit floating-point format. In one example, a processor includes fetch circuitry to fetch an instruction having fields to specify an opcode and locations of a first source vector comprising N single-precision elements, and a destination vector comprising at least N 16-bit floating-point elements, the opcode to indicate execution circuitry is to convert each of the elements of the specified source vector to 16-bit floating-point, the conversion to include truncation and rounding, as necessary, and to store each converted element into a corresponding location of the specified destination vector, decode circuitry to decode the fetched instruction, and execution circuitry to respond to the decoded instruction as specified by the opcode.

Claims (40)

1 . A processor comprising:

a control register to specify a rounding mode;

fetch circuitry to fetch a format conversion instruction;

a decode unit to decode the format conversion instruction, the format conversion instruction having an opcode, a first field to specify a source vector register, a second field to specify a mask, and a third field to specify a destination vector register, the source vector register to store a source vector having a plurality of 32-bit single-precision floating point data elements, wherein the mask has a plurality of mask bits respectively corresponding to a plurality of data element positions of the destination vector register; and

execution circuitry coupled with the decode unit, the execution circuitry to perform operations corresponding to the format conversion instruction, including to:

for data element positions of the destination vector register for which the corresponding mask bit is one:

convert corresponding 32-bit single-precision floating point data elements of the source vector to corresponding 16-bit floating point data elements, according to the rounding mode specified by the control register, the 16-bit floating point data elements having a format, the format including a sign bit, an 8-bit exponent, seven explicit mantissa bits, and one implicit mantissa bit; and

store the 16-bit floating point data elements in the data element positions of the destination vector register; and

for data element positions of the destination vector register for which the corresponding mask bit is zero, not change values in the data element positions, wherein the destination vector register has a plurality of data element positions that do not correspond to the plurality of mask bits.

2 . The processor of claim 1 , wherein the plurality of data element positions that do not correspond to the plurality of mask bits is a same number as a number of data element positions that do correspond to the plurality of mask bits.

3 . The processor of claim 1 , wherein the mask is a first register in a set of registers having a second register, the processor not supporting using the second register as a mask.

4 . The processor of claim 1 , wherein the source vector is a 128-bit source vector.

5 . The processor of claim 1 , wherein the source vector is a 256-bit source vector.

6 . The processor of claim 1 , wherein the source vector is a 512-bit source vector.

7 . The processor of claim 1 , wherein the source vector is a 1024-bit source vector.

8 . The processor of claim 1 , wherein the format conversion instruction allows the source vector to be any one of 128-bits and 512-bits.

9 . The processor of claim 1 , wherein the format is a bfloat16 format.

10 . The processor of claim 1 , wherein the processor is a general-purpose CPU core.

11 . The processor of claim 1 , wherein the processor is a reduced instruction set computing (RISC) processor.

12 . The processor of claim 1 , wherein the plurality of data element positions that do not correspond to the plurality of mask bits is a same number as a number of data element positions that do correspond to the plurality of mask bits, wherein the mask is a first register in a set of registers having a second register, the processor not supporting using the second register as a mask, and wherein the format conversion instruction allows the source vector to be any one of 128-bits and 512-bits.

13 . The processor of claim 1 , wherein the plurality of data element positions that do not correspond to the plurality of mask bits is a same number as a number of data element positions that do correspond to the plurality of mask bits, wherein the format is a bfloat16 format, wherein the processor is a general-purpose CPU core, and wherein the format conversion instruction allows the source vector to be any one of 128-bits and 512-bits.

14 . A method comprising:

specifying a rounding mode in a control register;

fetching a format conversion instruction;

decoding the format conversion instruction, the format conversion instruction having an opcode, a first field specifying a source vector register, a second field specifying a mask, and a third field specifying a destination vector register, the source vector register storing a source vector having a plurality of 32-bit single-precision floating point data elements, wherein the mask has a plurality of mask bits respectively corresponding to a plurality of data element positions of the destination vector register; and

performing operations corresponding to the format conversion instruction, including:

for data element positions of the destination vector register for which the corresponding mask bit is one:

converting corresponding 32-bit single-precision floating point data elements of the source vector to corresponding 16-bit floating point data elements, according to the rounding mode specified by the control register, the 16-bit floating point data elements having a format, the format including a sign bit, an 8-bit exponent, seven explicit mantissa bits, and one implicit mantissa bit; and

storing the 16 -bit floating point data elements in the data element positions of the destination vector register; and

for data element positions of the destination vector register for which the corresponding mask bit is zero, not changing values in the data element positions, wherein the destination vector register has a plurality of data element positions that do not correspond to the plurality of mask bits.

15 . The method of claim 14 , wherein the plurality of data element positions that do not correspond to the plurality of mask bits is a same number as a number of data element positions that do correspond to the plurality of mask bits, wherein the mask is a first register in a set of registers having a second register, the second register not able to be used as a mask, and wherein the format conversion instruction allows the source vector to be any one of 128-bits and 512-bits.

16 . The method of claim 14 , wherein the format is a bfloat16 format, wherein the plurality of data element positions that do not correspond to the plurality of mask bits is a same number as a number of data element positions that do correspond to the plurality of mask bits, and wherein the format conversion instruction allows the source vector to be any one of 128-bits and 512-bits.

17 . A processor comprising:

a decode unit to decode a format conversion instruction, the format conversion instruction to indicate a location of a first source operand, a location of a second source operand, a destination register, a writemask register, and a type of masking, the first source operand to include a first plurality of 32-bit single-precision floating point data elements, the second source operand to include a second plurality of 32-bit single-precision floating point data elements, the writemask register to store a plurality of mask bits each corresponding to a data element position in the destination register, the type of masking to be either zeroing masking or merging masking; and

execution circuitry coupled to the decode unit, the execution circuitry to perform operations corresponding to the format conversion instruction, including to:

for each of the first plurality of 32-bit single-precision floating point data elements that is of a first type, convert the 32-bit single-precision floating point data element to a 16-bit floating point data element using round to nearest even rounding behavior and store a result data element in a corresponding data element position in a first half of a result in the destination register if the mask bit corresponding to the data element position in the plurality of mask bits is set, and otherwise include a masked data element in the data element position, and

for each of the second plurality of 32-bit single-precision floating point data elements that is of the first type, convert the 32-bit single-precision floating point data element to a 16-bit floating point data element using round to nearest even rounding behavior and store a result data element in a corresponding data element position in a second half of the result in the destination register if the mask bit corresponding to the data element position in the plurality of mask bits is set, and otherwise include a masked data element in the data element position, the result data elements stored in the destination register to have a format that includes one sign bit, eight exponent bits, and seven explicit mantissa bits, the masked data element to be a zero value if the type of masking is zeroing masking and to be a value stored in the data element position prior to performing the operations corresponding to the format conversion instruction if the type of masking is merging masking, wherein the first type is a normal number.

18 . The processor of claim 17 , wherein the format is a BF16 format, wherein the first half of the result is a lower order half of the result and the second half of the result is a higher order half of the result.

19 . The processor of claim 17 , wherein any denormal data elements in the first plurality of 32-bit single-precision floating point data elements and the second plurality of 32-bit single-precision floating point data elements are treated as zero values, and wherein the round to nearest even rounding behavior is used irrespective of a rounding behavior specified by a control register.

20 . The processor of claim 17 , wherein the location of the first source operand is a register location or a memory location, and wherein the first type excludes zero, denormal, infinity, and NaN, and wherein the first source operand and the second source operand consist of a same number of bits, wherein the same number of bits is 128, 256, or 512 bits.

Continuity (3)
Continuation 17851468 · Jun 28, 2022
Continuation 16186384 · Nov 9, 2018
Related Publication 20250117217A1 · Apr 10, 2025
References Cited (102)
US 5995122A · Hsieh et al. · 1999 [cited by applicant]
US 8412761B2 · Yoshida · 2013 [cited by applicant]
US 8667250B2 · Sprangle et al. · 2014 [cited by applicant]
US 10073695B2 · Anderson · 2018 [cited by examiner]
US 10705839B2 · Valentine et al. · 2020 [cited by applicant]
US 20020184282A1 · Yuval et al. · 2002 [cited by applicant]
US 20040268094A1 · Abdallah et al. · 2004 [cited by applicant]
US 20080077779A1 · Zohar et al. · 2008 [cited by applicant]
US 20120011348A1 · Eichenberger et al. · 2012 [cited by applicant]
US 20130073838A1 · Gschwind et al. · 2013 [cited by applicant]
US 20130290685A1 · Corbal et al. · 2013 [cited by applicant]
US 20140149724A1 · Valentine · 2014 [cited by examiner]
US 20140195580A1 · Anderson et al. · 2014 [cited by applicant]
US 20140208080A1 · Ould-Ahmed-Vall et al. · 2014 [cited by applicant]
US 20150039854A1 · Wick · 2015 [cited by applicant]
US 20150088946A1 · Anderson et al. · 2015 [cited by applicant]
US 20150286482A1 · Espasa et al. · 2015 [cited by applicant]
US 20170061279A1 · Kloss et al. · 2017 [cited by applicant]
US 20170068516A1 · Anderson et al. · 2017 [cited by applicant]
US 20170177350A1 · Ould-Ahmed-Vall · 2017 [cited by applicant]
US 20170286109A1 · Jha · 2017 [cited by applicant]
US 20180081685A1 · Bhuiyan et al. · 2018 [cited by applicant]
US 20180121199A1 · Uliel et al. · 2018 [cited by applicant]
US 20180157464A1 · Lutz et al. · 2018 [cited by applicant]
US 20180189065A1 · Raman et al. · 2018 [cited by applicant]
US 20180293078A1 · Gabrielli et al. · 2018 [cited by applicant]
US 20180321937A1 · Brown et al. · 2018 [cited by applicant]
US 20190042544A1 · Kashyap et al. · 2019 [cited by applicant]
US 20190354568A1 · Lindberg et al. · 2019 [cited by applicant]
CN 104145245A · 2014 [cited by applicant]
CN 106030510A · 2016 [cited by applicant]
CN 108369573A · 2018 [cited by applicant]
CN 108647044A · 2018 [cited by applicant]
WO 2013101233A1 · 2013 [cited by applicant]
WO 2015147895A1 · 2015 [cited by applicant]
WO 2017105715A1 · 2017 [cited by applicant]
AMD, “AMD64 Technology AMD64 Architecture Programmer's Manual vol. 1: Application Programming”, Rev. 3.20, May 2013, 386 pages. [cited by applicant]
Amd64, “Advanced Micro Devices AMD64 Technology AMD64 Architecture Programmer's Manual vol. 4: 128-Bit and 256-Bit Media Instructions”, Revision 3.17, Advanced Micro Devices, May 2013, 144 pages. [cited by applicant]
Bagnara et al., “Symbolic Path-Oriented Test Data Generation for Floating-Point Programs”, 2013 IEEE Sixth International Conference on Software Testing, Verification and Validation, IEEE, 2013, 10 pages. [cited by applicant]
Basu, S. K., “Parallel and Distributed Computing: Architectures and Algorithms”, PHI Learning, Jul. 30, 2016, pp. 36-37. [cited by applicant]
Communication pursuant to Article 94(3) EPC, EP App. No. 20216494.3, Sep. 21, 2021, 10 pages. [cited by applicant]
Decision to Grant, EP App. No. 20216494.3, Nov. 7, 2024, 2 pages. [cited by applicant]
Decision to Grant, EP App. No. 21169540.8, Dec. 14, 2023, 2 pages. [cited by applicant]
European Communication pursuant to Article 94(3) EPC, EP App. No. 20216494.3, Apr. 7, 2021, 9 pages. [cited by applicant]
European Communication pursuant to Article 94(3) EPC, EP App. No. 20216494.3, Jul. 18, 2022, 9 pages. [cited by applicant]
European Search Report and Search Opinion , EP App. No. 20207968.7, Feb. 10, 2021, 12 pages. [cited by applicant]
European Search Report and Search Opinion, EP App. No. 19201879.4, Jun. 24, 2020, 10 pages. [cited by applicant]
European Search Report and Search Opinion, EP App. No. 21169540.8, Jul. 27, 2021, 14 pages. [cited by applicant]
European Search Report, EP App. No. 20216494.3, Mar. 24, 2021, 5 pages. [cited by applicant]
Final Office Action, U.S. Appl. No. 16/186,384, Sep. 15, 2020, 9 pages. [cited by applicant]
Final Office Action, U.S. Appl. No. 17/851,468, Dec. 11, 2023, 10 pages. [cited by applicant]
IEEE, “IEEE Standard for Floating-Point Arithmetic”, IEEE Computer Society, IEEE Std 754(trademark)-2008 (Revision of IEEE Std 754-1985), Aug. 29, 2008, 70 pages. [cited by applicant]
Intel Corporation, “Intel® Architecture Instruction Set Extensions Programming Reference”, Order No. 319433-012, Feb. 2012, 596 pages. [cited by applicant]
Intel Corporation, “Intel® Architecture Instruction Set Extensions Programming Reference”, Order No. 319433-023, Aug. 2015, 12 pages. [cited by applicant]
Intel, “BFLOAT16—Hardware Numerics Definition”, White Paper, Revision 1.0, Document No. 338302-001US, Nov. 2018, 7 pages. [cited by applicant]
Intel, “Intel® 64 and IA-32 Architectures Software Developer's Manual”, vol. 1: Basic Architecture, Order No. 253665-066US, Mar. 2018, 4 pages (5-27, 14-18, 14-19). [cited by applicant]
Intention to grant , EP App. No. 21169540.8, Aug. 21, 2023, 07 pages. [cited by applicant]
Intention to Grant, EP App. No. 20216494.3, Jul. 3, 2024, 6 pages. [cited by applicant]
Kolli et al., “High-Performance Transactions for Persistent Memories”, ACM, 2016, 13 Pages. [cited by applicant]
Koster et al., “Flexpoint: An Adaptive Numerical Format for Efficient Training of Deep Neural Networks”, NIPS'17: Proceedings of the 31st International Conference on Neural Information Processing Systems, Dec. 2017, pp.… [cited by applicant]
Lacassagne et al., “16-bit floating point instructions for embedded multimedia applications”, Proceedings of the Seventh International Workshop on Computer Architecture for Machine Perception (CAMP'05), IEEE, 2005, 6 pa… [cited by applicant]
Non-Final Office Action, U.S. Appl. No. 16/186,384, Apr. 9, 2020, 9 pages. [cited by applicant]
Non-Final Office Action, U.S. Appl. No. 16/186,384, May 21, 2021, 10 pages. [cited by applicant]
Non-Final Office Action, U.S. Appl. No. 17/851,468, May 5, 2023, 10 pages. [cited by applicant]
Notice of Allowance, CN App. No. 202011497335.0, Feb. 29, 2024, 08 pages (04 pages of English Translation and 04 pages of Original Document). [cited by applicant]
Notice of Allowance, CN App. No. 202110484218.9, Jun. 25, 2024, 11 pages (05 page of English Translation and 06 pages of Original Document). [cited by applicant]
Notice of Allowance, U.S. App. No. 16/186,384, Feb. 2, 2022, 8 pages. [cited by applicant]
Notice of Allowance, U.S. Appl. No. 16/186,384, Jun. 1, 2022, 3 pages. [cited by applicant]
Notice of Allowance, U.S. Appl. No. 16/186,384, Mar. 17, 2022, 3 pages. [cited by applicant]
Notice of Allowance, U.S. Appl. No. 17/133,078, Mar. 22, 2021, 8 pages. [cited by applicant]
Notice of Allowance, U.S. Appl. No. 17/133,255, Mar. 22, 2021, 8 pages. [cited by applicant]
Notice of Allowance, U.S. Appl. No. 17/851,468, Jul. 11, 2024, 7 pages. [cited by applicant]
Notice of Allowance, U.S. Appl. No. 17/851,468, Sep. 13, 2024, 2 pages. [cited by applicant]
NVIDIA, “NVIDIA Tesla P100”, Whitepaper, WP-08019-001_v01.1, pp. 1-45. [cited by applicant]
NVIDIA, “Parallel Thread Execution ISA”, Application Guide, v6.3, Oct. 2018, 87 pages. [cited by applicant]
Office Action , EP App. No. 19201879.4, Jun. 7, 2024, 05 pages. [cited by applicant]
Office Action , EP App. No. 19201879.4, Sep. 5, 2023, 08 pages. [cited by applicant]
Office Action, CN App. No. 202011497335.0, Sep. 26, 2023, 3 pages. [cited by applicant]
Office Action, CN App. No. 202110484218.9, Jan. 21, 2024, 05 pages of Original Document Only. [cited by applicant]
Office Action, EP App. No. 19201879.4, Apr. 21, 2022, 9 pages. [cited by applicant]
Office Action, EP App. No. 19201879.4, Feb. 9, 2021, 8 pages. [cited by applicant]
Office Action, EP App. No. 20207968.7, Aug. 22, 2023, 10 pages. [cited by applicant]
Office Action, EP App. No. 20207968.7, Dec. 6, 2021, 8 pages. [cited by applicant]
Office Action, EP App. No. 20207968.7, Jul. 5, 2023, 3 pages. [cited by applicant]
Office Action, EP App. No. 20207968.7, Jun. 23, 2022, 8 pages. [cited by applicant]
Office Action, EP App. No. 20207968.7, Nov. 13, 2023, 7 pages. [cited by applicant]
Office Action, EP App. No. 20216494.3, Apr. 18, 2023, 9 pages. [cited by applicant]
Office Action, EP App. No. 20216494.3, Feb. 2, 2024 17 pages. [cited by applicant]
Office Action, EP App. No. 20216494.3, Mar. 2, 2022, 7 pages. [cited by applicant]
Office Action, EP App. No. 20216494.3, Nov. 22, 2022, 10 pages. [cited by applicant]
Office Action, EP App. No. 21169540.8, Apr. 13, 2022, 5 pages. [cited by applicant]
Office Action, EP App. No. 21169540.8, Nov. 22, 2022, 11 pages. [cited by applicant]
Second Office Action, CN App. No. 202011497335.0, Jan. 4, 2024, 03 pages of Original Document Only. [cited by applicant]
Stephens et al., The ARM Scalable Vector Extension, IEEE Micro, vol. 37, No. 2, Mar.-Apr. 2017, 8 pages. [cited by applicant]
Tagliavini et al., “A Transprecision Floating-Point Platform for Ultra-Low Power Computing”, EDAA, 2018, 6 pages. [cited by applicant]
Tagliavini et al., “A Transprecision Floating-Point Platform for Ultra-Low Power Computing”, EDAA, Design, Automation And Test in Europe, 2018, pp. 1051-1056. [cited by applicant]
Waterman et al., “The RISC-V Instruction Set Manual, vol. I: Base User-Level ISA”, Electrical Engineering and Computer Sciences, Technical Report No. UCB/EECS-2011-62, May 13, 2011, 34 pages. [cited by applicant]
Wikimedia Foundation, Inc., “BFloat16 Floating-Point Format”, available online at <https://en.wikipedia.org/w/index.php?title=Bfloat16_floating%ADpoint_format&oldid=856298357>, Aug. 24, 2018, 3 pages. [cited by applicant]
Wikimedia Foundation, Inc., “Half-Precision Floating-Point Format”, available online at <https://en.wikipedia.org/w/index.php?title=Half-precision_floating%ADpoint_format&oldid=849804123>, Jul. 11, 2018, 6 pages. [cited by applicant]
Wikipedia, “Bfloat16 Floating-Point Format”, Available Online at <https://en.wikipedia.org/w/index.php?title=Bfloat16_floating-point_format&oldid=862115029>, Oct. 2, 2018, 5 pages. [cited by applicant]
Wikipedia, “Half-Precision Floating-Point Format”, Available Online at <https://en.wikipedia.org/w/index.php?title=Half-precision_floating-point_format&oldid=866403892>, Oct. 30, 2018, 5 pages. [cited by applicant]
Extended European Search Report, EP App. No. 25170319.5, Jul. 18, 2025, 12 pages. [cited by applicant]