IP Library › Granted Patent US 12,705,056
Granted Patent B2
US 12,705,056 · App. 18/227,608 · Granted Aug 11, 2026

Computing system having an accumulator register having entries having a bit width larger than a main register file

Inventors: Brian W. Thompto (Austin, TX); Maarten J. Boersma (Holzgerlingen, DE); Andreas Wagner (Wildberg, DE); Jose E. Moreira (Irvington, NY); Hung Q. Le (Austin, TX); Silvia Melitta Mueller (St. Ingbert, DE); Dung Q. Nguyen (Austin, TX)
Assignee: International Business Machines Corporation
G06F9/30098G06F9/30036G06F9/3012G06F9/3001G06F9/30109G06F9/3013G06F9/384
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,705,056
App. No.
18/227,608
Filed
Jul 28, 2023
Granted
Aug 11, 2026
Kind
B2
Examiner
DOMAN, SHAWN
Art Unit
2183
USPC
712/221
Abstract

A computer system, processor, and method for processing information is disclosed that includes at least one computer processor; a main register file associated with the at least one processor, the main register file having a plurality of entries for storing data, one or more write ports to write data to the main register file entries, and one or more read ports to read data from the main register file entries; one or more execution units including a dense math execution unit; and at least one accumulator register file having a plurality of entries for storing data. The results of the dense math execution unit in an aspect are written to the accumulator register file, preferably to the same accumulator register file entry multiple times, and the data from the accumulator register file is written to the main register file.

Claims (44)

1 . A processor for processing electronic data, the processor comprising:

a main physical register file having a plurality of main physical register file entries for storing main register data, each main physical register file entry having a main register file bit width for storing the main register data;

a physical accumulator register file separate and distinct from the main physical register file and having a plurality of physical accumulator register file entries for storing accumulator register data, each physical accumulator register file entry having an accumulator register bit field width, wherein the accumulator register bit field width is wider than the main register file bit field width;

one or more execution units for performing operations on the electronic data,

wherein the processor is configured to:

map at least one of the plurality of main physical register file entries to at least one of the plurality of physical accumulator register file entries;

perform operations with the one or more execution units;

write results of the operations performed with the one or more execution units to the physical accumulator register file; and

write data in the at least one of the plurality of physical accumulator register file entries to the at least one of the main physical register file entries to which the at least one of the plurality of physical accumulator register entries is mapped.

2 . The processor of claim 1 , wherein the processor is configured to read and write the at least one of the plurality of physical accumulator register file entries that is mapped to the at least one of the plurality of main physical register file entries without writing the main physical register file.

3 . The processor of claim 1 , wherein the processor is further configured so that the physical accumulator register file is both a source and a target during the operations of the one or more execution unit operations.

4 . The processor of claim 1 , wherein the processor is further configured to write the at least one of the plurality of physical accumulator register file entries that is mapped to the at least one of the plurality of main physical register file entries several times during operations of the one or more execution units without writing results of the operations of the one or more execution units to the main physical register file.

5 . The processor of claim 1 , wherein the one or more execution units include a dense math execution unit and the at least one physical accumulator register file is local to the dense math execution unit.

6 . The processor of claim 5 , wherein the dense math execution unit is a matrix-multiply-accumulator (MMA) unit and the at least one physical accumulator register file is located in the MMA.

7 . A processor for processing instructions, the processor comprising:

a main physical register file having a plurality of main physical register file entries for storing main register data, each main physical register file entry having a main register bit field width for storing the main register data;

one or more execution units including a dense math execution unit;

at least one physical accumulator register file separate and distinct from the main physical register file and having a plurality of physical entries for storing accumulator register data, each physical accumulator register file entry of the at least one physical accumulator register file having an accumulator register bit field width that is wider than the main register bit field width of the plurality of main physical register file entries and wherein each physical entry in the at least one physical accumulator register file is mapped to a plurality of main physical register file entries,

the processor configured to:

write results of the dense math execution unit to the at least one physical accumulator register file; and

write data from the at least one physical accumulator register file to the main physical register file.

8 . The processor of claim 7 , wherein the processor is further configured to write results back to a same physical accumulator register file entry multiple times.

9 . The processor of claim 7 , wherein the processor is further configured to write data from the at least one physical accumulator register file to a plurality of main physical register file entries in response to an instruction accessing a main physical register file entry that is mapped to a physical accumulator register file entry.

10 . The processor of claim 7 , wherein the processor is further configured to prime the at least one physical accumulator register file to receive data.

11 . The processor of claim 10 , wherein the processor is further configured to mark, in response to priming a physical accumulator register file entry, the plurality of main physical register file entries mapped to the primed physical accumulator register file entry as busy.

12 . The processor of claim 10 , wherein the processor is further configured to prime the at least one physical accumulator register file in response to an instruction to store data to the at least one physical accumulator register file.

13 . The processor of claim 7 , wherein the at least one physical accumulator register file is local to the dense math execution unit.

14 . The processor of claim 13 , wherein the dense math execution unit is a matrix-multiply-accumulator (MMA) unit and the at least one physical accumulator register file is located in the MMA.

15 . The processor of claim 7 , wherein the processor further comprises a vector scalar execution unit (VSU) and the dense math execution unit is a matrix-multiply-accumulator (MMA) unit and the main physical register file is a VS register file located in the VSU and the at least one physical accumulator register file is mapped to a plurality of consecutive VS register file entries.

16 . A computing system for processing information, the computing system comprising:

a main register file having a plurality of entries for storing main register data;

one or more execution units including a dense math execution unit;

at least one accumulator register file having a plurality of accumulator register file entries for storing accumulator register data, wherein the at least one accumulator register file is associated with the dense math execution unit,

the computing system configured to:

prime at least one accumulator register file entry to receive data, wherein the at least one accumulator register file entry is at least one of the plurality of accumulator register file entries of the at least one accumulator register file associated with the dense math execution unit;

mark, in response to priming the at least one accumulator register file entry to receive data, one or more main register file entries mapped to the at least one primed accumulator register file entry as busy; and

process data in the dense math execution unit where results of the dense math execution unit are written to the at least one primed accumulator register file entry.

17 . The computing system of claim 16 , the computing system further configured to:

prime the at least one accumulator register file entry to receive data in response to an instruction to store data to the at least one accumulator register file; and

write results back to the at least one primed accumulator register file entry multiple times.

18 . The computing system of claim 17 , the computing system further configured to:

de-prime the at least one primed accumulator register file entry written to multiple times;

write the resulting data from the at least one primed accumulator register file entry written to multiple times to the main register file; and

deallocate the at least one de-primed accumulator register file entry.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jul 28, 2023
From: THOMPTO, BRIAN W.; BOERSMA, MAARTEN J.; WAGNER, ANDREAS; MOREIRA, JOSE E.; LE, HUNG Q.; MUELLER, SILVIA MELITTA; NGUYEN, DUNG Q.
To: INTERNATIONAL BUSINESS MACHINES CORPORATION
Reel/Frame 064421/0432 →
Continuity (3)
Continuation 17458717 · Aug 27, 2021
Continuation 16555640 · Aug 29, 2019
Related Publication 20230367597A1 · Nov 16, 2023
References Cited (47)
US 5838984A · Nguyen et al. · 1998 [cited by applicant]
US 5983256A · Peleg et al. · 1999 [cited by applicant]
US 6029006A · Alexander et al. · 2000 [cited by applicant]
US 7181484B2 · Stribaek et al. · 2007 [cited by applicant]
US 8572153B1 · New · 2013 [cited by examiner]
US 8719828B2 · Lewis et al. · 2014 [cited by applicant]
US 9575890B2 · Busaba et al. · 2017 [cited by applicant]
US 9727370B2 · Greiner et al. · 2017 [cited by applicant]
US 10387122B1 · Olsen · 2019 [cited by applicant]
US 11132198B2 · Thompto et al. · 2021 [cited by applicant]
US 20020108026A1 · Balmer et al. · 2002 [cited by applicant]
US 20040044716A1 · Colon-Bonet · 2004 [cited by applicant]
US 20040078554A1 · Glossner et al. · 2004 [cited by applicant]
US 20050138323A1 · Snyder · 2005 [cited by applicant]
US 20050172106A1 · Ford et al. · 2005 [cited by applicant]
US 20050240644A1 · Van Berkel et al. · 2005 [cited by applicant]
US 20060095729A1 · Hokenek et al. · 2006 [cited by applicant]
US 20090248780A1 · Symes · 2009 [cited by examiner]
US 20130205123A1 · Vorbach · 2013 [cited by applicant]
US 20130212354A1 · Mimar · 2013 [cited by examiner]
US 20190012170A1 · Qadeer et al. · 2019 [cited by applicant]
US 20190171448A1 · Chen et al. · 2019 [cited by applicant]
US 20190250915A1 · Yadavalli · 2019 [cited by applicant]
US 20190294412A1 · Loh · 2019 [cited by examiner]
US 20200320662A1 · Lueh et al. · 2020 [cited by applicant]
US 20210173662A1 · Leenstra et al. · 2021 [cited by applicant]
CN 101246435A · 2008 [cited by applicant]
CN 104969215A · 2015 [cited by applicant]
CN 105378651A · 2016 [cited by applicant]
CN 108874744A · 2018 [cited by applicant]
CN 114787772A · 2022 [cited by applicant]
DE 112020004071B4 · 2025 [cited by applicant]
EP 0876026A2 · 1998 [cited by applicant]
GB 2603653A · 2022 [cited by applicant]
JP 2022546034A · 2022 [cited by applicant]
WO 2021038337A1 · 2021 [cited by applicant]
English-language translation of German Office Action dated Jan. 2, 2024, received in a corresponding foreign application, namely German File No. 11 2020 004 071.2, 12 pages. [cited by applicant]
English-language translation of Notice of Reasons for Refusal dated Feb. 20, 2024 received in a corresponding foreign patent application, namely Japanese Patent Application No. 2022-513041, 9 pages. [cited by applicant]
Pedram et al., “On the Efficiency of Register File versus Broadcast Interconnect for Collective Communications in Data-Parallel Hardware Accelerators”, IEEE 24th International Symposium on Computer Architecture and High… [cited by applicant]
Kim et al., “An Instruction Set and Microarchitecture for Instruction Level Distributed Processing”, ACM SIGARCH Computer Architecture News 30(2) ⋅ Apr. 2002, 11 pages. [cited by applicant]
Anonymous, “Method for removing accumulator dependencies”, IP.com Prior Art Database Technical Disclosure, IP.com No. IPCOM000008848D, Jul. 17, 2001, 6 pages. [cited by applicant]
Anonymous, “Using a common Error Correcting Special Purpose Register for correcting errors in a regular file”, IP.com Prior Art Database Technical Disclosure, IP.com No. IPCOM000202463D, Dec. 16, 2010, 3 pages. [cited by applicant]
Anonymous, “Methods for Application Checkpointing using Application Dependence Analysis”, IP.com Prior Art Database Technical Disclosure, IP.com No. IPCOM000222538D, Oct. 16, 2012, 6 pages. [cited by applicant]
International Search Report and Written Opinion dated Nov. 3, 2020 received in a corresponding foreign application, 7 pages. [cited by applicant]
List of IBM Patents or Patent Applications Treated as Related dated Jul. 28, 2023, 2 pages. [cited by applicant]
G.R. Wilson, “Embedded Systems and Computer Architecture”, Reed Elsevier pie group, pp. 216-224 (Year: 2002). [cited by applicant]
Anonymous, “Method to Prime and De-prime the Accumulator Register for Dense Math Engine (MMA) Execution”, Jul. 27, 2020, IP.com, pp. 1-6. [cited by applicant]