IP Library › Granted Patent US 12,731,013
Granted Patent B2
US 12,731,013 · App. 17/234,995 · Granted Sep 8, 2026

Weighted matrix for input data stream

Inventors: Jitendra Onkar Kolhe (Bangalore, IN); Soumitra Chatterjee (Bangalore, IN)
Assignee: Hewlett Packard Enterprise Development LP
G06N3/063G06F7/5443G06F9/5027G06F17/16
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,731,013
App. No.
17/234,995
Granted
Sep 8, 2026
Kind
B2
Abstract

Examples of performing convolution operations based on a weighted matrix are described. In an example, an input data stream vector is processed using a weighted matrix stored onto a processing unit of a neural network accelerator. The weighted matrix may correspond to a first convolution filter and a second convolution filter.

Claims (75)

1 . A system comprising:

a neural network accelerator;

a processor;

a machine-readable storage medium comprising machine-level instructions executable by the processor to:

receive an input data stream comprising a plurality of channels, wherein each channel of the plurality of channels comprises a respective matrix of image data of the input data stream;

process the image data of the input data stream, by:

flattening at least a portion of the respective matrix of each channel of the plurality of channels to obtain a respective plurality of single dimension input data vectors; and

combining the respective plurality of single dimension input data vectors to obtain a single dimension input data stream vector;

while processing the image data of the input data stream, generate a weighted matrix for convolutional processing of the image data of the input data stream, by:

obtaining a first convolution filter and a second convolution filter;

flattening the first convolution filter and the second convolution filter to obtain a first single dimensional vector and a second single dimensional vector;

combining the first single dimensional vector and the second single dimensional vector to obtain the weighted matrix;

storing the weighted matrix in a memory of the neural network accelerator; and

applying the weighted matrix to a memristor crossbar array of a processing unit of the neural network accelerator;

perform the convolutional processing, comprising at least one of: blurring, sharpening, embossing, or edge detection of the image data of the input data stream, by performing a matrix vector multiplication operation to generate an output element array, wherein the matrix vector multiplication operation comprises applying the weighted matrix applied to the memristor crossbar array of the processing unit to the single dimension input data stream vector; and

generate a layered output stream comprising a first layer output corresponding to the first convolution filter and a second layer output corresponding to the second convolution filter, by storing a first portion of the output element array associated with the first convolution filter in the first layer output and storing a second portion of the output element array associated with the second convolution filter in the second layer output.

2 . The system as claimed in claim 1 , wherein the machine-readable storage medium comprises machine-level instructions executable by the processor to:

obtain a subsequent single dimensional vector based on a subsequent convolution filter; and

merge the subsequent single dimensional vector with the first single dimensional vector and the second single dimensional vector to provide the weighted matrix.

3 . The system as claimed in claim 1 , wherein the plurality of channels comprises a predefined number of channels.

4 . The system as claimed in claim 1 , wherein the matrix vector multiplication operation comprises a dot product of the single dimension input data stream vector and the weighted matrix.

5 . The system as claimed in claim 1 , wherein the output element array comprises a first element and a second element, wherein the first element corresponds to the first convolution filter and the second element corresponds to the second convolution filter.

6 . The system as claimed in claim 5 , wherein the first element and the second element are implemented within a first output matrix and a second output matrix, respectively.

7 . The system as claimed in claim 1 , wherein the portion of the respective matrix of each channel of the plurality of channels is based on a filter window corresponding to a size of the first convolution filter and the second convolution filter.

8 . The system as claimed in claim 3 , wherein the predefined number of channels of the input data stream is equal to a number of channels of the first convolution filter and the second convolution filter.

9 . The system as claimed in claim 1 , comprising:

a neural network compiler that generates the machine-level instructions from programmable instructions expressed using a domain specific language (DSL);

wherein the processing unit is a memristor crossbar array-based processing unit.

10 . A method comprising:

receiving, via processing circuitry, an input data stream, wherein the input data stream comprises image data comprising a respective matrix for each of a plurality of channels, the plurality of channels comprising a predefined number of channels;

processing the image data, by:

flattening at least a portion of the respective matrix for each channel of the plurality of channels to obtain a respective plurality of single dimension input data vectors, wherein each respective single dimension input data vector of the plurality of single dimension input data vectors corresponds to a respective channel of the plurality of channels; and

combining the respective plurality of single dimension input data vectors to obtain an input data stream vector, wherein the input data stream vector is a single dimension vector comprising data values corresponding to each of the plurality of channels;

while processing the image data;

generating a weighted matrix for convolution processing of the image data, the weighted matrix comprising a combination of a first single dimension filter vector of a first convolution filter and a second single dimension filter vector of a second convolution filter;

storing the weighted matrix in a memory of a neural network accelerator; and

applying the weighted matrix to a memristor crossbar array of a processing unit of the neural network accelerator;

performing the convolutional processing of the image data to perform at least one of blurring, sharpening, embossing, or edge detection of the input data stream vector by performing a matrix vector multiplication operation comprising applying the weighted matrix applied to the memristor crossbar array of the processing unit to the input data stream vector;

obtaining an output element array based on the convolutional processing of the image data; and

generating a layered output stream comprising a first layer output corresponding to the first convolution filter and a second layer output corresponding to the second convolution filter, by storing a first portion of the output element array associated with the first convolution filter in the first layer output and storing a second portion of the output element array associated with the second convolution filter in the second layer output.

11 . The method as claimed in claim 10 , wherein the input data stream vector is obtained based on a filter window corresponding to a size of the first convolution filter and the second convolution filter.

12 . The method as claimed in claim 11 , wherein flattening the portion of the respective matrix for each channel of the plurality of channels comprises:

selecting a first set of elements of a first respective matrix of a first channel of the plurality of channels based on the filter window, wherein the filter window is multidimensional;

reformatting the first set of elements into a first single dimension input vector of the plurality of single dimension input data vectors;

selecting a second set of elements of a second respective matrix of a second channel of the plurality of channels based on the filter window; and

reformatting the second set of elements into a second single dimension input vector of the plurality of single dimension input data vectors.

13 . The method as claimed in claim 12 , comprising:

selecting a subsequent set of elements of the first respective matrix based on the filter window and a stride factor;

reformatting the subsequent set of elements into a subsequent single dimension input vector;

obtaining a subsequent input data stream vector based on the subsequent single dimension input vector; and

processing the subsequent input data stream vector using the weighted matrix; and

obtaining a subsequent output element array based on the processing.

14 . The method as claimed in claim 10 , wherein the predefined number of channels of the input data stream is equal to the number of channels of the first convolution filter and the second convolution filter.

15 . The method as claimed in claim 10 , wherein the weighted matrix is obtained by:

flattening the first convolution filter and the second convolution filter to the first single dimension filter vector and the second single dimension filter vector, respectively; and

combining the first single dimension filter vector and the second single dimension filter vector to obtain the weighted matrix, wherein the weighted matrix comprises a two dimensional matrix.

16 . The method as claimed in claim 10 , wherein processing the input data stream vector comprises determining a dot product of the input data stream vector and the weighted matrix.

17 . A non-transitory computer-readable medium comprising instructions for performing a convolution operation using a neural network accelerator, the instructions being executable by a processing resource to:

obtain a plurality of convolution filters;

flatten each convolution filter of the plurality of the convolution filters to obtain a corresponding plurality of single dimensional filter vectors;

merge the plurality of single dimensional filter vectors to obtain a weighted matrix for convolutional processing of image data, the convolutional processing comprising at least one of: sharpening, embossing, or edge detection of the image data;

store the weighted matrix onto a processing unit of the neural network accelerator;

apply the weighted matrix to a memristor crossbar array of the processing unit of the neural network accelerator;

obtain an input data stream comprising a plurality of channels, wherein each channel of the plurality of channels comprises a respective matrix of the input data stream;

flatten at least a portion of the respective matrix of each channel of the plurality of channels to obtain a respective plurality of single dimension input data vectors, wherein each respective single dimension input data vector of the plurality of single dimension input data vectors corresponds to a respective channel of the plurality of channels;

merge the respective plurality of single dimension input data vectors to obtain an input data stream vector, wherein the input data stream vector is a single dimension vector comprising data values corresponding to each of the plurality of channels;

cause the processing unit to perform the convolutional processing, by performance of a matrix vector multiplication operation on the input data stream vector based on the weighted matrix applied to the memristor crossbar array to generate a set of output data streams, wherein each of the set of the output data streams corresponds to each of the plurality of the convolution filters; and

generate a layered output stream comprising a first layer output corresponding to a first convolution filter of the plurality of convolution filters and a second layer output corresponding to a second convolution filter of the plurality of convolution filters, by storing a first portion of an output element array associated with the first convolution filter in the first layer output and storing a second portion of the output element array associated with the second convolution filter in the second layer output.

18 . The computer-readable medium as claimed in claim 17 , wherein flattening the portion of the respective matrix for each channel of the plurality of channels comprises:

selecting a first set of elements of a first respective matrix of a first channel of the plurality of channels based on a filter window, wherein the filter window is multidimensional;

reformatting the first set of elements into a first single dimension input vector of the plurality of single dimension input data vectors;

selecting a second set of elements of a second respective matrix of a second channel of the plurality of channels based on the filter window; and

reformatting the second set of elements into a second single dimension input vector of the plurality of single dimension input data vectors.

19 . The computer-readable medium as claimed in claim 17 , wherein the input data stream corresponds to a digital image having three channels.

20 . The computer-readable medium as claimed in claim 19 , wherein each of the plurality of convolution filters comprises three channels.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Sep 26, 2021
From: KOLHE, JITENDRA ONKAR; CHATTERJEE, SOUMITRA
To: HEWLETT PACKARD ENTERPRISE DEVELOPMENT LP
Reel/Frame 057600/0341 →
Priority Claims (1)
IN 2020041055652 · Dec 21, 2020 · national
Continuity (1)
Related Publication 20220198250A1 · Jun 23, 2022
References Cited (16)
US 20010021271A1 · Ishibashi · 2001 [cited by examiner]
US 20180341495A1 · Culurciello et al. · 2018 [cited by applicant]
US 20190220709A1 · Freeman · 2019 [cited by examiner]
US 20200034148A1 · Sumbul et al. · 2020 [cited by applicant]
US 20200110984A1 · Chatterjee et al. · 2020 [cited by applicant]
US 20220012013A1 · Sebastian · 2022 [cited by examiner]
Jin, Jonghoon, Aysegul Dundar, and Eugenio Culurciello. “Flattened convolutional neural networks for feedforward acceleration.” arXiv preprint arXiv:1412.5474. (Year: 2015). [cited by examiner]
Ankit et al., “PUMA: A Programmable Ultra-efficient Memristor-based Accelerator for Machine Learning Inference”, Jan. 30, 2019, 17 pages. [cited by applicant]
Bruel et al., “Generalize or Die: Operating Systems Support for Memristor-based Accelerators”, (Research Paper), 2017 IEEE International Conference on Rebooting Computing (ICRC), Nov. 8-9, 2017, 8 Pages. [cited by applicant]
Hu et al., “Dot-product engine for neuromorphic computing: Programming 1T1M crossbar to accelerate matrix-vector multiplication”, Hewlett Packard Labs, 53nd ACM/EDAC/IEEE Design Automation Conference (DAC), Jun. 5-9, 20… [cited by applicant]
Shen et al., “Maximizing CNN Accelerator Efficiency Through Resource Partitioning”, (Research Paper), 2017 ACM/IEEE 44th Annual International Symposium on Computer Architecture (ISCA), Apr. 12, 2018, 13 Pgs. [cited by applicant]
Strachan et al., “The Dot-Product Engine (DPE): exploring high efficiency analog multiplication with memristorarrays”, HPE, Dec. 11, 2015, 29 pages. [cited by applicant]
Zhang et al., “Mitigate Parasitic Resistance in Resistive Crossbar-based Convolutional Neural Networks”, (Research Paper), ACM Journal on Emerging Technologies in Computing Systems, Dec. 17, 2019, 20 Pages. [cited by applicant]
Naoyuki Ichimura, “Accelerating Convolution Operations by GPU (CUDA), Part 1: Fundamentals with Example Code Using Only Global Memory”, available online at <https://qiita.com/naoyuki_ichimura/items/8c80e67a10d99c2fb53c>… [cited by applicant]
NVIDIA, “Convolution”, available online at <https://developer.nvidia.com/discover/convolution>, 2025, 3 pages. [cited by applicant]
NVIDIA, “Convolutional Neural Network (CNN)”, available online at <https://developer.nvidia.com/discover/convolutional-neural-network>, 2025, 4 pages. [cited by applicant]