IP Library Granted Patent US 12,450,056
Granted Patent B2
US 12,450,056 · App. 17/701,308 · Granted Oct 21, 2025

Efficient data layout and alignment for wide-vector accelerator systems

Inventors: Shubham Jain (Elmsford, NY); Geoffrey Burr (Cupertino, CA); Yasuteru Kohda (Yamato, JP)
Assignee: International Business Machines Corporation
G06F9/30038G06F9/30032G06F9/3877
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,450,056
App. No.
17/701,308
Granted
Oct 21, 2025
Kind
B2
Abstract

Efficient data layout and alignment techniques for effectively executing AI workloads in wide-vector accelerator systems are provided. In one aspect, a method for processing AI workloads includes: logically dividing a data vector into a hierarchy of segments and sub-segments with each of the segments including more than one of the sub-segments, wherein each of the sub-segments includes words, and each of the words includes data-bits; and physically mapping the data-bits such that the words belonging to a same given one of the sub-segments are mapped contiguously across all of the segments. An AI accelerator system is also provided.

Claims (45)

1. A method for processing artificial intelligence (AI) workloads, the method comprising:

logically dividing a data vector into a hierarchy of segments and sub-segments with each of the segments comprising more than one of the sub-segments, wherein the data vector is a row from memory, each of the sub-segments comprises words, and each of the words comprises data-bits;

physically mapping the data-bits such that the words belonging to a same given one of the sub-segments are mapped contiguously across all of the segments;

pulling the row from the memory; and

performing alignment operations on the segments, the sub-segments, or a combination thereof to create an aligned data-vector.

2. The method of claim 1 , wherein the data-bits are physically mapped such that a distance between the words of a particular one of the sub-segments is minimized across all of the segments.

3. The method of claim 1 , wherein the data-bits are physically mapped such that a distance between instances of a same given one of the words in each of the sub-segments is minimized across all of the segments.

4. The method of claim 1 , wherein the alignment operations comprise one or more of rotating the segments either clockwise or counter-clockwise by at least one of the segments, and shifting the sub-segments either left or right by at least one of the sub-segments.

5. The method of claim 4 , wherein the performing of the alignment operations comprises rotating the segments by at least one of the segments, and wherein at least one of the segments is rotated from one end of the row to an opposite end of the row.

6. The method of claim 1 , further comprising:

copying a portion of the aligned data-vector into a final output-vector, while selectively masking a remainder of the aligned data-vector.

7. The method of claim 6 , further comprising:

repeating the logically dividing, the mapping, the copying, the pulling and the performing iteratively over multiple other rows in the memory.

8. The method of claim 6 , further comprising:

transmitting the final output-vector to a compute engine.

9. The method of claim 8 , wherein the compute engine comprises an analog AI crossbar array, and wherein the method further comprises:

rearranging weights across rows or columns of the analog AI crossbar array using valid-zero interleaving mapping.

10. An artificial intelligence (AI) accelerator system, comprising:

a memory;

a compute engine connected to the memory; and

a hardware logic unit configured to:

pull a row from the memory, wherein the row is logically divided into a hierarchy of segments and sub-segments with each of the segments comprising more than one of the sub-segments, each of the sub-segments comprising words, and each of the words comprising data-bits, and wherein the data-bits are physically mapped such that the words belonging to a same given one of the sub-segments are mapped contiguously across all of the segments;

perform alignment operations on the segments, sub-segments or combinations thereof to create an aligned data vector;

copy a portion of the aligned data-vector into a final output-vector, while selectively masking a remainder of the aligned data-vector; and

transmit the final output-vector to the compute engine.

11. The AI accelerator system of claim 10 , wherein the compute engine comprises an analog AI crossbar array.

12. The AI accelerator system of claim 10 , wherein the hardware logic unit when performing the alignment operations is configured to perform one or more of: rotating the segments either clockwise or counter-clockwise by at least one of the segments, and shifting the sub-segments either left or right by at least one of the sub-segments.

13. The AI accelerator system of claim 12 , wherein the hardware logic unit comprises:

at least one set of first multiplexers configured to perform the rotating of the segments by at least one of the segments; and

a set of second multiplexers configured to perform the shifting of the sub-segments by at least one of the sub-segments.

14. The AI accelerator system of claim 10 , wherein the hardware logic unit comprises:

segment mask-registers configured to mask bits in the aligned data vector from the alignment operations performed on the segments; and

sub-segment mask-registers configured to mask bits in the aligned data vector from the alignment operations performed on the sub-segments.

15. A non-transitory computer program product for processing artificial intelligence (AI) workloads, the computer program product comprising a computer readable storage medium having program instructions embodied therewith, the program instructions executable by a computer to cause the computer to:

logically divide a data vector into a hierarchy of segments and sub-segments with each of the segments comprising more than one of the sub-segments, wherein the data vector is a row from memory each of the sub-segments comprises words, and each of the words comprises data-bits;

physically map the data-bits such that the words belonging to a same given one of the sub-segments are mapped contiguously across all of the segments;

pulling the row from the memory; and

performing alignment operations on the segments, the sub-segments, or a combination thereof to create an aligned data-vector.

16. The non-transitory computer program product of claim 15 , wherein the data-bits are physically mapped such that a distance between the words of a particular one of the sub-segments is minimized across all of the segments, and wherein the data-bits are physically mapped such that a distance between instances of a same given one of the words in each of the sub-segments is minimized across all of the segments.

17. The non-transitory computer program product of claim 15 , wherein the alignment operations comprise one or more of rotating the segments either clockwise or counter-clockwise by at least one of the segments, and shifting the sub-segments either left or right by at least one of the sub-segments.

18. The non-transitory computer program product of claim 15 , wherein the program instructions further cause the computer to:

copy a portion of the aligned data-vector into a final output-vector, while selectively masking a remainder of the aligned data-vector; and

transmit the final output-vector to a compute engine.

19. The method of claim 1 , wherein the alignment operations are not performed on data at a granularity finer than the sub-segments.

20. The method of claim 8 , wherein the compute engine comprises an analog AI crossbar array that includes an arrangement of uniquely addressable resistive processing units that embody a neural network.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Mar 22, 2022
From: JAIN, SHUBHAM; BURR, GEOFFREY; KOHDA, YASUTERU
To: INTERNATIONAL BUSINESS MACHINES CORPORATION
Reel/Frame 059343/0932 →
Continuity (1)
Related Publication 20230305841A1 · Sep 28, 2023
References Cited (21)
US 10467527B1 · Margaglia et al. · 2019 [cited by applicant]
US 10733039B2 · Torng et al. · 2020 [cited by applicant]
US 10877812B2 · Siegl et al. · 2020 [cited by applicant]
US 11017842B2 · Troia · 2021 [cited by applicant]
US 11055003B2 · Sun et al. · 2021 [cited by applicant]
US 11081149B1 · Park · 2021 [cited by applicant]
US 20200005127A1 · Baum et al. · 2020 [cited by applicant]
US 20200005902A1 · Mellen · 2020 [cited by examiner]
US 20200133990A1 · Mathuriya et al. · 2020 [cited by applicant]
US 20200242459A1 · Manipatruni et al. · 2020 [cited by applicant]
US 20210019633A1 · Venkatesh · 2021 [cited by applicant]
US 20210264257A1 · Hu et al. · 2021 [cited by applicant]
US 20210266000A1 · Ge · 2021 [cited by applicant]
US 20230305841A1 · Jain · 2023 [cited by examiner]
CN 110163478A · 2019 [cited by examiner]
Mishty et al., “Designing Efficient and High-Performance AI Accelerators with Customized STT-MRAM,” IEEE Transactions on Very Large Scale Integration (VLSI) Systems, pp. 1-13, Apr. 2021. [cited by applicant]
Hwang et al., “ASimOV: A Framework for Simulation and Optimization of an Embedded AI Accelerator,” Micromachines, vol. 12, Issue 7, Jul. 19, 2021 (12 pages). [cited by applicant]
Yan et al., “iCELIA: A Full-Stack Framework for STT-MRAM-Based Deep Learning Acceleration,” IEEE Transactions on Parallel and Distributed Systems, vol. 31, No. 2, pp. 408-422, Feb. 1, 2020. [cited by applicant]
Chi et al.; “PRIME: A Novel Processing-in-memory Architecture for Neural Network Computation in ReRAM-based Main Memory,” ISCA ACM/IEEE 43rd Annual International Symposium on, Jun. 18-22, 2016 (13 pages). [cited by applicant]
Hussain et al.; “PVMC: Programmable Vector Memory Controller,” ASAP IEEE 25th International Conference on Application-Specific Systems, Architectures and Processors, Jun. 18-20, 2014 (8 pages). [cited by applicant]
Mell et al., “The NIST Definition of Cloud Computing,” NIST Special Publication 800-145, Sep. 2011 (7 pages). [cited by applicant]