IP Library › Granted Patent US 12,399,743
Granted Patent B2
US 12,399,743 · App. 17/652,109 · Granted Aug 26, 2025

Padding input data for artificial intelligence accelerators

Inventors: Cedric Lichtenau (Stuttgart, DE); Vijayalakshmi Srinivasan (New York, NY); Sunil K Shukla (Scarsdale, NY); Swagath Venkataramani (White Plains, NY); Kailash Gopalakrishnan (New York, NY); Holger Horbach (Aidlingen, DE); Razvan Peter Figuli (Remchingen, DE); Wei Wang (Yorktown Heights, NY); Yulong Li (Hartsdale, NY); Martin A Lutz (Peekskill, NY)
Assignee: International Business Machines Corporation
G06F9/5016G06F9/3887
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,399,743
App. No.
17/652,109
Filed
Feb 23, 2022
Granted
Aug 26, 2025
Kind
B2
Art Unit
2182
USPC
712/204
Abstract

Processing input data for transmittal to a data consumer such as an artificial intelligence engine is performed by arranging the input data into a uniform structure made up of sticks of data combined to form pages of sticks. A stick is any well-sized set of input data elements whereby the size of the stick is fixed. A masking pattern is established for sticks of data having certain ranges of invalid data for consumption of partial sticks while maintaining validity of the input data being transferred. The mask pattern is derived based on set-active-mask-and-value (SAMV) instructions. The derived mask pattern is carried forward for subsequent load instructions to the data consumer.

Claims (66)

1. A method for padding data while loading data from a memory to a data consumer, the method comprising:

receiving a load instruction including a padding instruction and a read instruction to read input data arranged in a stick layout from a memory to a data consumer, the padding instruction including a replacement value for masked elements and padding parameters;

deriving a mask pattern from parameters specified in the padding instruction;

padding the input data by masking invalid elements of the stick layout; and

generating a set of data sticks including the padded data for the data consumer.

2. The method of claim 1 , wherein the parameters specified in the padding instructions include layout parameters and pad range parameters.

3. The method of claim 1 , further comprising:

responsive to the load instruction, transmitting the set of data sticks to the data consumer.

4. The method of claim 1 , further comprising:

decoding the padding instructions to determine the replacement value and the mask pattern.

5. The method of claim 1 , wherein:

the stick layout is separable into a predefined number of slices, and

masking the invalid elements includes:

identifying invalid slices in a stick, the invalid slices including an invalid element, and

determining a starting slice having a portion of valid elements and a portion of invalid elements.

6. The method of claim 1 , wherein the padding instruction provides a specified data format in a 6-bit instruction including data format, cross-slice dimension type, and a key dimension of a stick.

7. The method of claim 1 , wherein the stick layout in which the input data is arranged is a pre-determined size matching an SIMD (single instruction, multiple data) capacity of an accelerator performing the deriving and padding steps.

8. A computer program product comprising a computer-readable storage medium having a set of instructions stored therein which, when executed by a processor, causes the processor to perform a method comprising:

receiving a load instruction including a padding instruction, the load instruction including a read instruction to read input data arranged in a stick layout from a memory to a data consumer, the padding instruction including a replacement value for masked elements and padding parameters;

deriving a mask pattern from parameters specified in the padding instruction;

padding the input data by masking invalid elements of the stick layout; and

generating a set of data sticks including the padded data for the data consumer.

9. The computer program product of claim 8 , wherein the parameters specified in the padding instructions include layout parameters and pad range parameters.

10. The computer program product of claim 8 , further causing the processor to perform a method comprising:

responsive to the load instruction, transmitting the set of data sticks to the data consumer.

11. The computer program product of claim 8 , further causing the processor to perform a method comprising:

decoding the padding instructions to determine the replacement value and the mask pattern.

12. The computer program product of claim 8 , wherein:

the stick layout is separable into a predefined number of slices, and

masking the invalid elements includes:

identifying invalid slices in a stick, the invalid slices including an invalid element, and

determining a starting slice having a portion of valid elements and a portion of invalid elements.

13. The computer program product of claim 8 , wherein the padding instruction provides a specified data format in a 6-bit instruction including data format, cross-slice dimension type, and a key dimension of a stick.

14. A computer system for padding data while loading data from a memory to a data consumer, the computer system comprising:

a processor set; and

a computer readable storage medium;

wherein:

the processor set is structured, located, connected, and/or programmed to run program instructions stored on the computer readable storage medium; and

the program instructions which, when executed by the processor set, cause the processor set to perform a method comprising:

receiving a load instruction including a padding instruction, the load instruction including a read instruction to read input data arranged in a stick layout from a memory to a data consumer, the padding instruction including a replacement value for masked elements and padding parameters;

deriving a mask pattern from parameters specified in the padding instruction;

padding the input data by masking invalid elements of the stick layout; and

generating a set of data sticks including the padded data for the data consumer.

15. The computer system of claim 14 , wherein the parameters specified in the padding instructions include layout parameters and pad range parameters.

16. The computer system of claim 14 , further causing the processor set to perform a method comprising:

responsive to the load instruction, transmitting the set of data sticks to the data consumer.

17. The computer system of claim 14 , further causing the processor set to perform a method comprising:

decoding the padding instructions to determine the replacement value and the mask pattern.

18. The computer system of claim 14 , wherein:

the stick layout is separable into a predefined number of slices, and

masking the invalid elements includes:

identifying invalid slices in a stick, the invalid slices including an invalid element, and

determining a starting slice having a portion of valid elements and a portion of invalid elements.

19. The computer system of claim 14 , wherein the padding instruction provides a specified data format in a 6-bit instruction including data format, cross-slice dimension type, and a key dimension of a stick.

20. A computer-implemented method comprising:

decoding a padding instruction to determine a replacement pad value and padding parameters of a mask pattern for a set of stickified input data, the padding instruction embedded in a first load instruction for the stickified input data;

deriving a mask pattern from the padding parameters; and

applying the derived mask pattern and the associated replacement pad value to a subsequent load instruction to load subsequent input data, the subsequent load instruction equivalent to the first load instruction, the subsequent input data is padded during transmission to a requesting consumer.

21. The method of claim 20 , wherein the set of stickified input data is arranged as sticks, each stick being separable into pre-defined number of slices.

22. The method of claim 21 , further comprising:

padding the set of stickified input data to generate the set of data sticks by masking invalid elements of each stick, the masking invalid elements of a stick including:

identifying invalid slices in the stick, the invalid slices including an invalid element, and

determining a starting slice having a portion of valid elements and a portion of invalid elements.

23. The method of claim 20 , wherein the replacement pad value is determined with reference to a pad value register file.

24. The method of claim 20 , further comprising:

generating a set of data sticks from the set of stickified input data for a requesting consumer, the set of data sticks padded according to the derived mask pattern including the replacement pad value.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Feb 23, 2022
From: LICHTENAU, CEDRIC; SRINIVASAN, VIJAYALAKSHMI; SHUKLA, SUNIL K; VENKATARAMANI, SWAGATH; GOPALAKRISHNAN, KAILASH; HORBACH, HOLGER; FIGULI, RAZVAN PETER; WANG, WEI; LI, YULONG; LUTZ, MARTIN A
To: INTERNATIONAL BUSINESS MACHINES CORPORATION
Reel/Frame 059072/0410 →
Continuity (1)
Related Publication 20230267003A1 · Aug 24, 2023
References Cited (38)
US 12198221B2 · Surti et al. · 2025 [cited by applicant]
US 12206552B2 · Guim Bernat et al. · 2025 [cited by applicant]
US 20020010793A1 · Noll · 2002 [cited by examiner]
US 20170139709A1 · Gschwind · 2017 [cited by applicant]
US 20180088853A1 · Kotra et al. · 2018 [cited by applicant]
US 20180165574A1 · Young · 2018 [cited by applicant]
US 20190042250A1 · Anders · 2019 [cited by applicant]
US 20190079768A1 · Heinecke · 2019 [cited by applicant]
US 20200005128A1 · Temam · 2020 [cited by applicant]
US 20200081744A1 · Siegl · 2020 [cited by applicant]
US 20200104691A1 · Bai · 2020 [cited by applicant]
US 20210065005A1 · Zhu · 2021 [cited by applicant]
US 20210182059A1 · Akin · 2021 [cited by applicant]
US 20210192359A1 · Khish Ardestani Zadeh · 2021 [cited by applicant]
US 20210224125A1 · Liu · 2021 [cited by applicant]
US 20210264250A1 · Singh · 2021 [cited by applicant]
US 20220027546A1 · Ren · 2022 [cited by examiner]
US 20240135158A1 · Heyne et al. · 2024 [cited by applicant]
CN 1501259A · 2004 [cited by applicant]
CN 107924291A · 2018 [cited by applicant]
CN 113383309A · 2021 [cited by applicant]
CN 113836049A · 2021 [cited by applicant]
CN 119013653A · 2024 [cited by applicant]
DE 112023001068T5 · 2025 [cited by applicant]
GB 2630701A · 2024 [cited by applicant]
WO 2021119907A1 · 2021 [cited by applicant]
WO 2023161783A1 · 2023 [cited by applicant]
Anonymous. ““RDNA 2” Instruction Set Architecture.” Published Nov. 30, 20 by AMD. 291 pages. https://developer.amd.com/wp-content/resources/RDNA2_Shader_ISA_November2020.pdf. [cited by applicant]
Anonymous. “Parallel Thread Execution ISA.” Published Oct. 22 by NVIDIA. 516 pages. https://docs.nvidia.com/cuda/parallel-thread-execution/index.html#texture-instructions-tld4. [cited by applicant]
Anonymous. “Unit Description.” Printed Oct. 13, 2021. 43 pages. Published by NVDLA. <http://nvdla.org/hw/v1/ias/unit_description.html>. [cited by applicant]
Dave, et al., “Hardware Acceleration of Sparse and Irregular Tensor Computations of ML Models: A Survey and Insights.” Last edited Jul. 22, 2021. 44 pages. Published by ARXIV. https://arxiv.org/abs/2007.00864v2. [cited by applicant]
Heyne et al., “Processing Tensors ”, U.S. Appl. No. 18/046,322, filed Oct. 13, 2022, 47 pages. [cited by applicant]
IBM Appendix P, “List of patents and patent applications to be treated as related”, Filed Oct. 14, 2022, 2 pages. [cited by applicant]
“Patent Cooperation Treaty PCT International Search Report”, International application No. PCT/IB2023/051533, International filing date Feb. 20, 2023, Date of mailing Jun. 19, 2023, 6 pages. [cited by applicant]
Adiono et al., “Low Latency YOLOv3—Tiny Accelerator for Low-Cost FPGA Using General Matrix Multiplication Principle”, IEEE Access, date of current version Oct. 25, 2021, 24 pages. [cited by applicant]
Authors et al.: Disclosed Without Attribution. “System for Managing a Container Based Resource Pool for On-Demand Usage”, IP.com No. IPCOM000262948D, Jul. 16, 2020, 4 pages. [cited by applicant]
Miller Michael. “Virtual Accelerator Engines”, CTO, MoSys, Inc., 2020, 9 pages. [cited by applicant]
White Paper, “Achieving vMotion Acceleration Over Efficient Virtualized Network (EVN)”, Mellanox Technologies, Dec. 2014, 8 pages. [cited by applicant]