IP Library Granted Patent US 12,468,947
Granted Patent B2
US 12,468,947 · App. 17/814,782 · Granted Nov 11, 2025

Stickification using anywhere padding to accelerate data manipulation

Inventors: Swagath Venkataramani (White Plains, NY); Vijayalakshmi Srinivasan (New York, NY); Shubham Jain (Elmsford, NY); Sarada Krithivasan (Chennai, IN); Sanchari Sen (White Plains, NY)
Assignee: International Business Machines Corporation
G06N3/082G06F9/3004G06F9/3836
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,468,947
App. No.
17/814,782
Granted
Nov 11, 2025
Kind
B2
Abstract

Embodiments are provided for efficient realization of memory-bound operations in a computing system by a processor. Data may be read from and written to a memory at a granular level using a stickification operation. One or more regions of activation and weight tensor data on the memory may be annotated by coupling the stickification operation with padding.

Claims (36)

1 . A method for providing efficient realization of memory-bound operations a computing environment by one or more processors comprising:

reading or writing data from a memory at a granular level using a stickification operation; and

annotating regions of activation and weight tensor data on the memory by coupling the stickification operation with padding to accelerate data manipulation operations.

2 . The method of claim 1 , further including:

unstickifying one or more inputs of a tensor; and

restickifying the one or more inputs of the tensors.

3 . The method of claim 1 , further including dividing one or more tensors in one or more of a plurality of dimensions.

4 . The method of claim 1 , further including providing a padded region to be spread in a plurality of dimensions on a data stick.

5 . The method of claim 1 , further including transposing a tensor by shuffling a plurality of dimensions on a data stick, wherein the data stick format is unaltered.

6 . The method of claim 1 , further including reshaping one or more dimensions of the tensor into a single dimension or a plurality of dimensions.

7 . The method of claim 1 , further including squeezing a dimension of a tensor having a size equal to a defined value.

8 . A system for providing efficient realization of memory-bound operations in a computing environment, comprising:

one or more computers with executable instructions that when executed cause the system to:

read or writing data from a memory at a granular level using a stickification operation; and

annotate regions of activation and weight tensor data on the memory by coupling the stickification operation with padding.

9 . The system of claim 8 , wherein the executable instructions when executed cause the system to:

unstick one or more inputs of a tensor; and

restick the one or more inputs of the tensors.

10 . The system of claim 8 , wherein the executable instructions when executed cause the system to divide one or more tensors in one or more of a plurality of dimensions.

11 . The system of claim 8 , wherein the executable instructions when executed cause the system to provide a padded region to be spread in a plurality of dimensions on a data stick.

12 . The system of claim 8 , wherein the executable instructions when executed cause the system to transpose a tensor by shuffling a plurality of dimensions on a data stick, wherein the data stick format is unaltered.

13 . The system of claim 8 , wherein the executable instructions when executed cause the system to reshape one or more dimensions of the tensor into a single dimension or a plurality of dimensions.

14 . The system of claim 8 , wherein the executable instructions when executed cause the system to squeeze a dimension of a tensor having a size equal to a defined value.

15 . A computer program product for providing efficient realization of memory-bound operations in a computing environment, the computer program product comprising:

one or more computer readable storage media, and program instructions collectively stored on the one or more computer readable storage media, the program instruction comprising:

program instructions to read or writing data from a memory at a granular level using a stickification operation; and

program instructions to annotate regions of activation and weight tensor data on the memory by coupling the stickification operation with padding to accelerate data manipulation operations.

16 . The computer program product of claim 15 , further including program instructions to:

unstick one or more inputs of a tensor; and

restick the one or more inputs of the tensors.

17 . The computer program product of claim 15 , further including program instructions to divide one or more tensors in one or more of a plurality of dimensions.

18 . The computer program product of claim 15 , further including program instructions to provide a padded region to be spread in a plurality of dimensions on a data stick.

19 . The computer program product of claim 15 , further including program instructions to transpose a tensor by shuffling a plurality of dimensions on a data stick, wherein the data stick format is unaltered.

20 . The computer program product of claim 15 , further including program instructions to:

reshape one or more dimensions of the tensor into a single dimension or a plurality of dimensions; and

squeeze a dimension of a tensor having a size equal to a defined value.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jul 25, 2022
From: VENKATARAMANI, SWAGATH; SRINIVASAN, VIJAYALAKSHMI; JAIN, SHUBHAM; KRITHIVASAN, SARADA; SEN, SANCHARI
To: INTERNATIONAL BUSINESS MACHINES CORPORATION
Reel/Frame 060610/0701 →
Continuity (1)
Related Publication 20240028899A1 · Jan 25, 2024
References Cited (20)
US 9299124B2 · Chen et al. · 2016 [cited by applicant]
US 10133691B2 · Brewer et al. · 2018 [cited by applicant]
US 10644723B2 · Stainbrook et al. · 2020 [cited by applicant]
US 10812105B2 · Park et al. · 2020 [cited by applicant]
US 10904058B2 · Zhang et al. · 2021 [cited by applicant]
US 10915450B2 · Brown et al. · 2021 [cited by applicant]
US 20190370631A1 · Fais et al. · 2019 [cited by applicant]
US 20200192803A1 · Sun et al. · 2020 [cited by applicant]
US 20200218985A1 · Wei et al. · 2020 [cited by applicant]
US 20200250525A1 · Addepalli et al. · 2020 [cited by applicant]
US 20200382239A1 · Dikarev et al. · 2020 [cited by applicant]
US 20210056396A1 · Majnemer et al. · 2021 [cited by applicant]
US 20240220768A1 · Zhang · 2024 [cited by examiner]
US 20240370693A1 · Zhang · 2024 [cited by examiner]
Faleiro, et al., “High performance multi-core transaction processing via deterministic execution”, Ph.D. dissertation, Yale University, Dec. 2018 (Year: 2018). [cited by examiner]
Faleiro et al., “Lazy Evaluation of Transactions in Database Systems”, SIGMOD '14: Proceedings of the 2014 ACM SIGMOD International Conference on Management of Data, Jun. 2014, pp. 15-26, https://doi.org/10.1145/2588555… [cited by applicant]
Tahara, Daniel, “Scheduling Heuristics for Lazy Database Systems”, Yale University Department of Computer Science, YALEU/DCS/TR-1488, May 2014, (7 pages). [cited by applicant]
Wang et al., “Parameterized Hardware Accelerators for Lattice-Based Cryptography and Their Application to the HW / SW Co-Design of qTESLA”, IACR Transactions on Cryptographic Hardware and Embedded Systems, 2020(3), 269-… [cited by applicant]
Alon et al., “Tight Bounds for Shared Memory Systems Accessed by Byzantine Processes”, Distributed Computing, vol. 18, Issue 2, Dec. 2005, pp. 99-109, https://doi.org/10.1007/s00446-005-0125-8 (11 pages). [cited by applicant]
Sasaki et al., “Practical Byte-Granular Memory Blacklisting using Califorms”, MICRO '52: Proceedings of the 52nd Annual IEEE/ACM International Symposium on Microarchitecture, 2019, pp. 558-571, https://doi.org/10.1145/3… [cited by applicant]