IP Library › Granted Patent US 12,450,165
Granted Patent B2
US 12,450,165 · App. 18/732,865 · Granted Oct 21, 2025

Vector based matrix multiplication

Inventors: Asheesh Bhardwaj (Allen, TX); Mujibur Rahman (Plano, TX); Timothy David Anderson (University Park, TX)
Assignee: TEXAS INSTRUMENTS INCORPORATED
G06F12/1045G06F7/24G06F7/487G06F7/4876G06F7/49915G06F7/53G06F7/57G06F9/3001G06F9/30014G06F9/30018G06F9/30021G06F9/30032G06F9/30036G06F9/30065G06F9/30072G06F9/30098G06F9/30112G06F9/30145G06F9/30149G06F9/3016G06F9/32G06F9/345G06F9/3802G06F9/3818G06F9/383G06F9/3836G06F9/3851G06F9/3856G06F9/3867G06F9/3887G06F9/48G06F11/00G06F11/1048G06F12/0862G06F12/0875G06F12/0897G06F12/1009G06F17/16H03H17/0664G06F9/325G06F9/381G06F9/3822G06F11/10G06F15/7807G06F15/781G06F2212/452G06F2212/60G06F2212/602G06F2212/68
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,450,165
App. No.
18/732,865
Granted
Oct 21, 2025
Kind
B2
Abstract

A method is provided that includes performing, by a processor in response to a vector matrix multiply instruction, multiplying an m×n matrix (A matrix) and a n×p matrix (B matrix) to generate elements of an m×p matrix (R matrix), and storing the elements of the R matrix in a storage location specified by the vector matrix multiply instruction.

Claims (67)

1. A system comprising:

a first permute component including a first plurality of permute networks;

a first multiplexer including an output and inputs coupled to the first plurality of permute networks;

a second permute component including a second plurality of permute networks;

a second multiplexer including an output and inputs coupled to the second plurality of permute networks;

a third multiplexer including:

a first input coupled to the output of the first multiplexer;

a second input coupled to the output of the second multiplexer; and

an output; and

a multiplication component coupled to the output of the third multiplexer.

2. The system of claim 1 , wherein the first plurality of permute networks includes a first permute network configured to permute data for performing a finite impulse response operation.

3. The system of claim 1 , wherein the first plurality of permute networks includes a second permute network configured to permute data for performing matrix multiplication.

4. The system of claim 1 , further comprising a logical OR circuit coupled to the third multiplexer and the multiplication component.

5. The system of claim 4 , further comprising:

a third permute component; and

a fourth multiplexer coupled to the third permute component, wherein the logical OR circuit is coupled to the fourth multiplexer.

6. The system of claim 5 , wherein the logical OR circuit is configured to concatenate an output of the third multiplexer and an output of the fourth multiplexer.

7. The system of claim 1 , further comprising:

a first buffer coupled to the first permute component; and

a second buffer coupled to the second permute component.

8. The system of claim 7 , wherein input packets of a first set of data elements and a second set of data elements are alternated in a ping pong fashion between the first buffer and the second buffer.

9. The system of claim 7 ,

wherein the first and second buffers are configured to receive a set of data elements, and

wherein the third multiplexer is configured to select between the first permute component and the second permute component based on which of the first buffer and the second buffer stores a next data element of the set of data elements.

10. The system of claim 9 , further comprising an instruction decoder circuit configured to:

generate a select signal based on which buffer of the first buffer and the second buffer stores the next data element; and

send the select signal to the third multiplexer.

11. The system of claim 1 , wherein the multiplication component is configured to perform vector multiplication.

12. A method comprising:

selecting, by a first multiplexer, among a first plurality of permute networks;

selecting, by a second multiplexer, among a second plurality of permute networks;

selecting, by a third multiplexer, among a third plurality of permute networks;

selecting, by a fourth multiplexer, among a fourth plurality of permute networks;

selecting, by a fifth multiplexer, between an output of the first multiplexer and an output of the second multiplexer;

selecting, by a sixth multiplexer, between an output of the third multiplexer and an output of the fourth multiplexer;

performing a logical OR operation on an output of the fifth multiplexer and an output of the sixth multiplexer; and

performing a multiplication operation on an output of the logical OR operation.

13. The method of claim 12 , further comprising:

by a first permute network in the first plurality of permute networks, permuting data for performing a finite impulse response operation; and

by a second permute network in the first plurality of permute networks, permuting data for performing matrix multiplication.

14. The method of claim 12 , wherein performing the logical OR operation comprises concatenating the output of the fifth multiplexer and the output of the sixth multiplexer.

15. A system comprising:

a first permute component including a first plurality of permute networks;

a first multiplexer coupled to the first plurality of permute networks;

a second permute component including a second plurality of permute networks;

a second multiplexer coupled to the second plurality of permute networks;

a third permute component including a third plurality of permute networks;

a third multiplexer coupled to the third plurality of permute networks;

a fourth permute component including a fourth plurality of permute networks;

a fourth multiplexer coupled to the fourth plurality of permute networks;

a fifth multiplexer coupled to the first multiplexer and the second multiplexer;

a sixth multiplexer coupled to the third multiplexer and the fourth multiplexer;

a logical OR circuit coupled to the fifth multiplexer and the sixth multiplexer; and

a multiplication component coupled to the logical OR circuit.

16. The system of claim 15 , wherein the first plurality of permute networks includes:

a first permute network configured to permute data for performing a finite impulse response operation; and

a second permute network configured to permute data for performing matrix multiplication.

17. The system of claim 15 , wherein the logical OR circuit is configured to concatenate an output of the fifth multiplexer and an output of the sixth multiplexer.

18. The system of claim 15 , further comprising:

a first buffer coupled to the first permute component; and

a second buffer coupled to the second permute component,

wherein the third multiplexer is configured to select between the first permute component and the second permute component based on which of the first buffer and the second buffer stores a next data element of a set of data elements, and

wherein input packets of a first set of data elements and a second set of data elements are alternated in a ping pong fashion between the first buffer and the second buffer.

19. The system of claim 18 , further comprising an instruction decoder circuit configured to:

generate a select signal based on which buffer of the first buffer and the second buffer stores the next data element; and

send the select signal to the third multiplexer.

20. The system of claim 15 , wherein the multiplication component is configured to perform vector multiplication.

Continuity (4)
Continuation 17749671 · May 20, 2022
Continuation 16878607 · May 20, 2020
Provisional Application 62852870 · May 24, 2019
Related Publication 20240354259A1 · Oct 24, 2024
References Cited (29)
US 5107504A · Nakamura · 1992 [cited by examiner]
US 5181183A · Miyazaki · 1993 [cited by examiner]
US 5303172A · Magar et al. · 1994 [cited by applicant]
US 5867289A · Gerstel · 1999 [cited by examiner]
US 6606686B1 · Agarwala et al. · 2003 [cited by applicant]
US 7873812B1 · Mimar · 2011 [cited by applicant]
US 8918445B2 · Anderson et al. · 2014 [cited by applicant]
US 9904645B2 · Thompson et al. · 2018 [cited by applicant]
US 9959247B1 · Woo et al. · 2018 [cited by applicant]
US 10747531B1 · Langer et al. · 2020 [cited by applicant]
US 20020031220A1 · Lee et al. · 2002 [cited by applicant]
US 20030014457A1 · Desai · 2003 [cited by examiner]
US 20050004962A1 · Ju · 2005 [cited by applicant]
US 20070089019A1 · Tang et al. · 2007 [cited by applicant]
US 20080285652A1 · Oxman et al. · 2008 [cited by applicant]
US 20100039567A1 · Gruijters · 2010 [cited by examiner]
US 20120314775A1 · Laksono et al. · 2012 [cited by applicant]
US 20140101358A1 · Kaltenbach et al. · 2014 [cited by applicant]
US 20140365548A1 · Mortensen · 2014 [cited by applicant]
US 20160041813A1 · Narayanamoorthy et al. · 2016 [cited by applicant]
US 20160321537A1 · Akopyan · 2016 [cited by examiner]
US 20190012295A1 · Yinger et al. · 2019 [cited by applicant]
US 20190018673A1 · Langhammer et al. · 2019 [cited by applicant]
US 20190171448A1 · Chen et al. · 2019 [cited by applicant]
US 20200082253A1 · Park · 2020 [cited by examiner]
US 20200210188A1 · Ould-Ahmed-Vall et al. · 2020 [cited by applicant]
EP 0408226A2 · 1991 [cited by applicant]
“66AK2Hx Keystone Multicore DSP+ARM System-on-Chips”, SPRT651A, Texas Instruments, Inc., 2013, pp. 1-3. [cited by applicant]
Pradeep Kumar Kumawat and Gajendra Sujediya, “Design and Comparison of 8×8 Wallace Tree Multiplier using CMOS and GDI Technology”, IOSR Journal of VLSI and Signal Processing (IOSR-JVSP), vol. 7, Issue 4, Version 1, Jun.… [cited by applicant]