IP Library › Granted Patent US 12,541,468
Granted Patent B2
US 12,541,468 · App. 18/584,151 · Granted Feb 3, 2026

Methods and apparatus to reduce read-modify-write cycles for non-aligned writes

Inventors: Naveen Bhoria (Plano, TX); Timothy David Anderson (University Park, TX); Pete Michael Hippleheuser (Murphy, TX)
Assignee: Texas Instruments Incorporated
G06F12/128G06F9/3001G06F9/30043G06F9/30047G06F9/546G06F11/1064G06F12/0215G06F12/0238G06F12/0292G06F12/0802G06F12/0804G06F12/0806G06F12/0811G06F12/0815G06F12/082G06F12/0853G06F12/0855G06F12/0864G06F12/0884G06F12/0888G06F12/0891G06F12/0895G06F12/0897G06F12/1027G06F12/12G06F12/121G06F12/126G06F12/127G06F13/1605G06F13/1642G06F13/1673G06F13/1689G06F15/8069G11C5/066G11C7/10G11C7/1015G11C7/106G11C7/1075G11C7/1078G11C7/1087G11C7/222G11C29/42G11C29/44G06F2212/1016G06F2212/1021G06F2212/1024G06F2212/1041G06F2212/1044G06F2212/301G06F2212/454G06F2212/603G06F2212/6032G06F2212/6042G06F2212/608G06F2212/62
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,541,468
App. No.
18/584,151
Granted
Feb 3, 2026
Kind
B2
Abstract

Methods, apparatus, systems and articles of manufacture are disclosed to reduce read-modify-write cycles for non-aligned writes. An example apparatus includes a memory that includes a plurality of memory banks, an interface configured to be coupled to a central processing unit, the interface to obtain a write operation from the central processing unit, wherein the write operation is to write a subset of the plurality of memory banks, and bank processing logic coupled to the interface and to the memory, the bank processing logic to determine the subset of the plurality of memory banks to write based on the write operation, and determine whether to cause a read operation to be performed in response to the write operation based on whether a number of addresses in the subset of the plurality of memory banks to write satisfies a threshold.

Claims (68)

1 . A system, comprising:

a cache memory comprising a first memory bank and a second memory bank;

an interface configured to receive a write operation indicating to write first data to the first memory bank and second data to the second memory bank; and

bank processing logic configured to:

determine whether a combination of the first data and the second data corresponds to a sum of a data width of the first memory bank and a data width of the second memory bank; and

based on determining that the combination of the first data and the second data is less than the sum of the data width of the first memory bank and the data width of the second memory bank,

cause write of the first data to the first memory bank without read of first existing data from the first memory bank;

cause read of second existing data from the second memory bank;

cause merge of the second data and the second existing data read from the second memory bank to generate second merged data; and

cause write of the second merged data to the second memory bank.

2 . The system of claim 1 , wherein:

the bank processing logic is configured to:

based on determining that the combination of the first data and the second data corresponds to the sum of the data width of the first memory bank and the data width of the second memory bank,

cause write of the first data to the first memory bank without read of the first existing data from the first memory bank; and

cause write of the second data to the second memory bank without read of the second existing data from the second memory bank.

3 . The system of claim 1 , wherein:

the interface is configured to receive another write operation indicating to write third data to the first memory bank; and

the bank processing logic is further configured to:

determine whether the third data corresponds to the data width of the first memory bank; and

based on determining that the third data corresponds to the data width of the first memory bank, cause write of the third data to the first memory bank without read of the first existing data from the first memory bank.

4 . The system of claim 3 , wherein the bank processing logic is further configured to:

based on determining that the third data is less than the data width of the first memory bank,

cause read of the first existing data from the first memory bank;

cause merge of the third data and the first existing data read from the first memory bank to generate first merged data; and

cause write of the first merged data to the first memory bank.

5 . The system of claim 1 , wherein the cache memory includes sixteen memory banks configured to be accessible independently from each other.

6 . The system of claim 1 , further comprising:

a first store queue coupled to the first memory bank; and

a second store queue coupled to the second memory bank.

7 . The system of claim 6 , wherein:

read of the first existing data from the first data to the first memory bank is performed based on the first store queue; and

read of the second existing data from the second memory bank is performed based on the second store queue.

8 . The system of claim 1 , wherein the cache memory is a victim cache.

9 . The system of claim 1 , wherein the cache memory is a level one (L1) cache memory.

10 . A system, comprising:

a cache memory comprising a first memory bank and a second memory bank;

a processor configured to generate a write operation indicating to write first data to the first memory bank and second data to the second memory bank; and

bank processing logic configured to:

determine whether a combination of the first data and the second data is same as a sum of a data width of the first memory bank and a data width of the second memory bank; and

based on determining that the combination of the first data and the second data is less than the sum of the data width of the first memory bank and the data width of the second memory bank,

cause write of the first data to the first memory bank without read of first existing data from the first memory bank;

cause read of second existing data from the second memory bank;

cause merge of the second data and the second existing data read from the second memory bank to generate second merged data; and

cause write of the second merged data to the second memory bank.

11 . The system of claim 10 , wherein:

the bank processing logic is configured to:

based on determining that the combination of the first data and the second data is same as the sum of the data width of the first memory bank and the data width of the second memory bank,

cause write of the first data to the first memory bank without read of the first existing data from the first memory bank; and

cause write of the second data to the second memory bank without read of the second existing data from the second memory bank.

12 . The system of claim 10 , wherein:

the processor is configured to generate another write operation indicating to write third data to the first memory bank; and

the bank processing logic is further configured to:

determine whether the third data is same as the data width of the first memory bank; and

based on determining that the third data is same as the data width of the first memory bank, cause write of the third data to the first memory bank without read of the first existing data from the first memory bank.

13 . The system of claim 12 , wherein the bank processing logic is further configured to:

based on determining that the third data is less than the data width of the first memory bank,

cause read of the first existing data from the first memory bank;

cause merge of the third data and the first existing data read from the first memory bank to generate first merged data; and

cause write of the first merged data to the first memory bank.

14 . The system of claim 10 , wherein the cache memory includes sixteen memory banks configured to be accessible independently from each other.

15 . The system of claim 10 , further comprising:

a first store queue coupled to the first memory bank; and

a second store queue coupled to the second memory bank.

16 . The system of claim 15 , wherein:

read of the first existing data from the first memory bank is performed based on the first store queue; and

read of the second existing data from the second memory bank is performed based on the second store queue.

17 . The system of claim 10 , wherein the cache memory is a victim cache.

18 . The system of claim 10 , wherein the cache memory is a level one (L1) cache memory.

Continuity (3)
Continuation 16882234 · May 22, 2020
Provisional Application 62852494 · May 24, 2019
Related Publication 20240232100A1 · Jul 11, 2024
References Cited (10)
US 7779307B1 · Favor · 2010 [cited by examiner]
US 10719058B1 · Johnson et al. · 2020 [cited by applicant]
US 20020087821A1 · Saulsbury et al. · 2002 [cited by applicant]
US 20050273564A1 · Lakshmanamurthy et al. · 2005 [cited by applicant]
US 20060036817A1 · Oza · 2006 [cited by examiner]
US 20090094380A1 · Qiu et al. · 2009 [cited by applicant]
US 20140181377A1 · Kimmel et al. · 2014 [cited by applicant]
US 20170052721A1 · Yamaji · 2017 [cited by applicant]
US 20190235759A1 · Sen et al. · 2019 [cited by applicant]
US 20200097406A1 · Ou · 2020 [cited by applicant]