IP Library › Granted Patent US 12,613,699
Granted Patent B2
US 12,613,699 · App. 18/956,439 · Granted Apr 28, 2026

DMA controller and LSU to transpose data arrays stored in main memory for storage in processor registers

Inventor: Jeppe Oland (San Francisco,, CA)
Assignee: Sony Interactive Entertainment Inc.
G06F9/30043G06F9/30029G06F9/30032
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,613,699
App. No.
18/956,439
Granted
Apr 28, 2026
Kind
B2
Abstract

One or more hardware elements operate on an array of data from a memory to generate a shuffled array of data. A subsequent of hardware operations on the shuffled array of data produces a transposed array of data in which rows and columns of the array of data are transposed. A load-store unit may then load the array of transposed data into a plurality of processor registers.

Claims (29)

1 . A computer system comprising:

a processor having a plurality of registers;

a main memory;

a local memory;

a direct memory access (DMA) controller configured to perform one or more operations on an array of data from the memory to generate a shuffled array of data from the array of data from the memory and store the shuffled array in the local memory; and

a load-store unit configured to perform one or more operations on the shuffled array of data stored in the local memory to produce a transposed array of data in which rows and columns of the array of data are transposed, and wherein the load-store unit is configured to load the transposed array of data into the plurality of registers.

2 . The system of claim 1 , wherein the DMA controller is configured to perform the one or more operations on the array of data from the memory by performing an incrementing XOR operation in each row of data in the array of data to shuffle lanes of data in each said row of data to generate the shuffled array of data.

3 . The system of claim 2 , wherein the load-store unit is configured to perform the one or more operations on the shuffled array of data by loading each lane in each row of the shuffled array of data from corresponding incremented memory indices of a local memory according to an incrementing XOR pattern to produce the transposed array of data.

4 . The system of claim 1 , wherein the DMA controller is configured to perform the one or more operations on the array of data from the memory by performing one or more swap operations between one or more pairs of data entries in two or more rows of data in the array of data to shuffle lanes of data in each said row of data to generate the shuffled array of data.

5 . The system of claim 4 , wherein the load-store unit is configured to perform the one or more operations on the shuffled array of data by performing one or more swap operations between one or more pairs of data entries in two or more rows of data in the shuffled array of data to produce the transposed array of data.

6 . A method for transferring transposed data from a memory to a plurality of registers, comprising: performing one or more operations on an array of data from a memory to generate a shuffled array of data with a direct memory access controller; and

performing one or more operations on the shuffled array of data with a load-store unit to produce a transposed array of data in which rows and columns of the array of data are transposed; and

loading the array of transposed data into a plurality of registers of a processor with the load-store unit.

7 . The method of claim 6 , wherein performing the one or more operations on the array of data from the memory includes performing an incrementing XOR operation in each row of data in the array of data to shuffle lanes of data in each said row of data to generate the shuffled array of data.

8 . The method of claim 7 , wherein performing the one or more operations on the shuffled array of data includes loading each lane in each row of the shuffled array of data from corresponding incremented memory indices of a local memory along with an incrementing XOR pattern to produce the transposed array of data.

9 . The method of claim 6 , wherein performing the one or more operations on the array of data from the memory includes performing one or more swap operations between one or more pairs of data entries in two or more rows of data in the array of data to shuffle lanes of data in each said row of data to generate the shuffled array of data.

10 . The method of claim 9 , wherein performing the one or more operations on the shuffled array of data includes performing one or more swap operations between one or more pairs of data entries in two or more rows of data in the shuffled array of data to produce the transposed array of data.

11 . A computer system comprising:

a main memory;

a direct memory access (DMA) controller;

a load-store unit (LSU), and;

a processor having a plurality of registers;

wherein the DMA controller, LSU, and processor include hardware configured to perform one or more operations on an array of data from the memory to generate a shuffled array of data from the array of data from the memory and to perform one or more operations on the shuffled array of data to produce a transposed array of data in which rows and columns of the array of data are transposed, and wherein the load-store unit is configured to load the transposed array of data into the processor's plurality of registers.

12 . The system of claim 11 , wherein the DMA controller is configured to perform the one or more operations on the array of data from the memory by performing an incrementing XOR operation in each row of data in the array of data to shuffle lanes of data in each said row of data to generate the shuffled array of data.

13 . The system of claim 12 , wherein the load-store unit is configured to perform the one or more operations on the shuffled array of data by loading each lane in each row of the shuffled array of data from corresponding incremented memory indices according to an incrementing XOR pattern to produce the transposed array of data.

14 . The system of claim 11 , wherein the DMA controller or the LSU is configured to perform the one or more operations on the array of data from the memory by performing one or more swap operations between one or more pairs of data entries in two or more rows of data in the array of data to shuffle lanes of data in each said row of data to generate the shuffled array of data.

15 . The system of claim 14 , wherein the DMA controller or the LSU is configured to perform the one or more operations on the shuffled array of data by performing one or more swap operations between one or more pairs of data entries in two or more rows of data in the shuffled array of data to produce the transposed array of data.

16 . The system of claim 11 , wherein the DMA controller is configured to perform the one or more operations on the array of data from the memory by performing one or more swap operations between one or more pairs of data entries in two or more rows of data in the array of data to shuffle lanes of data in each said row of data to generate the shuffled array of data and wherein the DMA controller configured to perform the one or more operations on the shuffled array of data by performing one or more swap operations between one or more pairs of data entries in two or more rows of data in the shuffled array of data to produce the transposed array of data.

17 . The system of claim 11 , wherein the LSU is configured to perform the one or more operations on the array of data from the memory by performing one or more swap operations between one or more pairs of data entries in two or more rows of data in the array of data to shuffle lanes of data in each said row of data to generate the shuffled array of data and wherein the LSU configured to perform the one or more operations on the shuffled array of data by performing one or more swap operations between one or more pairs of data entries in two or more rows of data in the shuffled array of data to produce the transposed array of data.

Assignments (2)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Feb 5, 2025
From: OLAND, JEPPE
To: SONY INTERACTIVE ENTERTAINMENT INC.
Reel/Frame 070117/0532 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Nov 22, 2024
From: OLAND, JEPPE
To: SONY INTERACTIVE ENTERTAINMENT INC.
Reel/Frame 069371/0709 →
Continuity (2)
Provisional Application 63612968 · Dec 20, 2023
Related Publication 20250208870A1 · Jun 26, 2025
References Cited (59)
US 5502747A · McGrath · 1996 [cited by applicant]
US 5619198A · Blackham et al. · 1997 [cited by applicant]
US 5815421A · Dulong et al. · 1998 [cited by applicant]
US 5862407A · Sriti · 1999 [cited by examiner]
US 5872965A · Petrick · 1999 [cited by applicant]
US 6006245A · Thayer · 1999 [cited by applicant]
US 6021206A · McGrath · 2000 [cited by applicant]
US 6259795B1 · McGrath · 2001 [cited by applicant]
US 6373416B1 · McGrath · 2002 [cited by applicant]
US 6389390B1 · Reilly · 2002 [cited by applicant]
US 6421697B1 · McGrath et al. · 2002 [cited by applicant]
US 6505223B1 · Haitsma et al. · 2003 [cited by applicant]
US 6625629B1 · Garcia · 2003 [cited by applicant]
US 6804771B1 · Jung · 2004 [cited by examiner]
US 9431987B2 · Betbeder et al. · 2016 [cited by applicant]
US 9614713B1 · Dias et al. · 2017 [cited by applicant]
US 11669464B1 · Bilski · 2023 [cited by examiner]
US 20040215683A1 · Beaumont · 2004 [cited by applicant]
US 20060038710A1 · Staszewski et al. · 2006 [cited by applicant]
US 20060051088A1 · Lee et al. · 2006 [cited by applicant]
US 20060114214A1 · Griffin · 2006 [cited by examiner]
US 20100293214A1 · Longley · 2010 [cited by applicant]
US 20110107060A1 · McAllister et al. · 2011 [cited by applicant]
US 20110246787A1 · Farrugia et al. · 2011 [cited by applicant]
US 20110307459A1 · Jacob et al. · 2011 [cited by applicant]
US 20120113133A1 · Shpigelblat · 2012 [cited by applicant]
US 20130262880A1 · Pong et al. · 2013 [cited by applicant]
US 20140355786A1 · Betbeder et al. · 2014 [cited by applicant]
US 20160306566A1 · Lu et al. · 2016 [cited by applicant]
US 20200363978A1 · Kajigaya · 2020 [cited by applicant]
US 20210097375A1 · Huynh et al. · 2021 [cited by applicant]
US 20220130404A1 · Fuchs et al. · 2022 [cited by applicant]
US 20220207108A1 · Dupont et al. · 2022 [cited by applicant]
US 20220244919A1 · Bondarenko · 2022 [cited by examiner]
US 20230185570A1 · Kerr · 2023 [cited by examiner]
US 20230259469A1 · Zheng · 2023 [cited by examiner]
US 20250210050A1 · Oland · 2025 [cited by applicant]
CN 1291766A · 2001 [cited by applicant]
CN 101050971A · 2007 [cited by applicant]
CN 100353664C · 2007 [cited by applicant]
CN 101283520A · 2008 [cited by applicant]
CN 101617235A · 2009 [cited by applicant]
CN 102387270A · 2012 [cited by applicant]
CN 102419396A · 2012 [cited by applicant]
CN 103001605A · 2013 [cited by applicant]
EP 1081684A2 · 2001 [cited by applicant]
WO WO1999049574A1 · 1999 [cited by applicant]
WO WO2003012405A2 · 2003 [cited by applicant]
[No Author Listed], “Digital signal processing with v850 and v850e devices,” Renesas Electronics Corporation, May 2005, 56 pages. [cited by applicant]
Armelloni et al., “Implementation of real-time partitioned convolution on a DSP board,” 2003 IEEE Workshop on Applications of Signal Processing to Audio and Acoustics, Oct. 19-22, 2003, pp. 71-74 (4 pages). [cited by applicant]
Battenberg et al., “Implementing real-time partitioned convolution algorithms on conventional operating systems,” Proc. of the 14th Int. Conference on Digital Audio Effects (DAFx-11), Sep. 19-23, 2011, pp. 313-320 (DAFX… [cited by applicant]
engineeringproductivitytools.com [online], “Practical Considerations.” Engineering Productivity Tools, available on or before Jul. 1, 2007, via Internet Archive: Wayback Machine URL<https://web.archive.org/web/200707011… [cited by applicant]
Kabal et al., “Rounding and Scaling in Fixed-Point FFT Implementations,” Rapport Technique De l'INRS—Telecommunications, Jun. 12, 1985, 48 pages. [cited by applicant]
Torger et al., a, A, “Real-time partitioned convolution for Ambiophonics surround sound,” 2001 IEEE Workshop on the Applications of Signal Processing to Audio and Acoustics, 2001, pp. 195-198 (W2001-1-4, 4 pages). [cited by applicant]
wikipedia.org [online], “Complex Number,” Oct. 17, 2023, retrieved on Oct. 26, 2023, retrieved from URL<https://en.wikipedia.org/wiki/Complex_number>, 13 pages. [cited by applicant]
Chen, “Digital signal processing on MMX technology,” Programmable Digital Signal Processors, 2001, retrieved on Dec. 3, 2025, retrieved from URL<https:/ /citeseerx.ist. psu.edu/document?repid=rep 1 &type=pdf &doi=074ee0… [cited by applicant]
Courtaud et al., “Improving prediction accuracy of memory interferences for multicore platforms,” RTSS 2019—40th IEEE Real-Time Systems Symposium, Dec. 2019, retrieved on Dec. 3, 2025, retrieved from URL<https://inria.h… [cited by applicant]
Dun et al., “Towards efficient canonical polyadic decomposition on sunway many-core processor,” Information Sciences, 549 (2021):221-248. [cited by applicant]
International Search Report and Written Opinion in International Application No. PCT/US2024/057395, mailed on Feb. 13, 2025, 14 pages. [cited by applicant]