IP Library Granted Patent US 12,061,666
Granted Patent B2
US 12,061,666 · App. 18/189,625 · Granted Aug 13, 2024

Distributing matrix multiplication processing among processing nodes

Inventor: Aaron M. Collier (Bloomington, MN)
Assignee: Hewlett Packard Enterprise Development LP
G06F17/16G06F9/5066G06F9/544
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,061,666
App. No.
18/189,625
Granted
Aug 13, 2024
Kind
B2
Abstract

Based on a predetermined number of available processor sockets, a plurality of candidate matrix decompositions are identified, which correspond to a multiplication of matrices. Based on a first comparative relationship of a variation of first sizes of the plurality of candidate matrix decompositions along a first dimension and a second comparative relationship of a variation of second sizes of the plurality of candidate matrix decomposition sizes along a second dimension, a given candidate matrix decomposition is selected. Processing of the multiplication among the processor sockets is distributed based on the given candidate matrix decomposition.

Claims (13)

1. An apparatus comprising:

at least one hardware processor; and

a memory to store instructions that, when executed by the at least one hardware processor, cause the at least one hardware processor to:

determine a plurality of candidate matrix decompositions associated with a matrix-matrix multiplication based on a predetermined number of available processor sockets;

based on first load balancing metrics, select a given candidate matrix decomposition of the plurality of candidate matrix decompositions and distribute processing of the matrix-matrix multiplication among the predetermined number of available processor sockets based on the selected given candidate matrix decomposition;

determine a plurality of candidate matrix sub-decompositions for partitioning the given candidate matrix block decomposition based on a predetermined number of processing nodes for each processor socket of the predetermined number of available processor sockets; and

based on second load balancing metrics, select a given candidate matrix sub-decomposition of the plurality of candidate matrix sub-decompositions, and for each processor socket of the predetermined number of available processor sockets, distribute processing of the matrix-matrix multiplication among the processing nodes of the processor socket based on the selected given candidate matrix sub-decomposition.

2. The apparatus of claim 1 , wherein the instructions, when executed by the at least one hardware processor, further cause the at least one hardware processor to, for each processing node:

determine a plurality of candidate processing thread-to-processing node assignments based on a predetermined number of threads for the processing node; and

based on third load balancing metrics, select a given candidate processing thread-to-processing node thread assignment and distribute processing of the matrix-matrix multiplication among the threads for the processing node based on the given candidate processing thread-to-processing node thread assignment.

3. The apparatus of claim 2 , wherein the third load balancing metrics comprise at least one of a cache block size and a processor core per last level cache number.

4. The apparatus of claim 2 , wherein the processing nodes comprises non-uniform memory access (NUMA) nodes.

5. The apparatus of claim 1 , wherein the first load balancing metrics comprise metrics to bias the selection of the given candidate decomposition to favor vertical partitioning.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Mar 24, 2023
From: COLLIER, AARON M.
To: HEWLETT PACKARD ENTERPRISE DEVELOPMENT LP
Reel/Frame 063094/0588 →
Continuity (2)
Division 16886189 · May 28, 2020
Related Publication 20230281271A1 · Sep 7, 2023
Cited By (1)
US 12,346,403