IP Library › Granted Patent US 12,530,295
Granted Patent B2
US 12,530,295 · App. 18/587,289 · Granted Jan 20, 2026

Decoupling atomicity from operation size

Inventors: Francesco Spadini (Sunset Valley, TX); Gideon Levinsky (Cedar Park, TX); Mridul Agarwal (Sunnyvale, CA)
Assignee: Apple Inc.
G06F12/0804G06F9/30043G06F9/3826G06F9/3834G06F2212/601
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,530,295
App. No.
18/587,289
Granted
Jan 20, 2026
Kind
B2
Abstract

In an embodiment, a processor implements a different atomicity size (for memory consistency order) than the operation size. More particularly, the processor may implement a smaller atomicity size than the operation size. For example, for multiple register loads, the atomicity size may be the register size. In another example, the vector element size may be the atomicity size for vector load instructions. In yet another example, multiple contiguous vector elements, but fewer than all the vector elements in a vector register, may be the atomicity size for vector load instructions.

Claims (39)

1 . A processor, comprising:

a decoder circuit configured to decode a load instruction into a load memory operation; and

a load/store unit configured to:

execute the load memory operation having an operation size that is larger than an atomicity size implemented for compliance with a memory ordering model;

verify, for compliance with the memory ordering model, that each of a plurality of atomic elements associated with the load memory operation is fully sourced from either a store memory operation in a store queue or from a different source, wherein the plurality of atomic elements are of the atomicity size; and

permit forwarding from the store queue responsive to a verification that each of the plurality of atomic elements is fully sourced from either a store memory operation in the store queue or from a different source.

2 . The processor of claim 1 , wherein the load/store unit is configured to:

prevent completion of the load memory operation responsive to a detection that:

at least one of the plurality of atomic elements is not fully sourced from a single source; and

at least one of the plurality of atomic elements is sourced from the store queue.

3 . The processor of claim 2 , wherein the load/store unit is configured to prevent the completion of the load memory operation until one or more store memory operations targeted by the load memory operation have drained from the store queue.

4 . The processor of claim 1 , wherein the load memory operation targets a plurality of registers, wherein the atomicity size is a size of a register, and wherein a given atomic element comprises bytes written into a given register of the plurality of registers.

5 . The processor of claim 1 , wherein the load memory operation is a vector load memory operation, and the atomicity size is a size of a vector element in a vector that is read by the load memory operation.

6 . The processor of claim 1 , wherein at least one of the plurality of atomic elements is fully sourced from a store memory operation in the store queue and at least one of the plurality of atomic elements is fully sourced from the different source.

7 . The processor of claim 1 , further comprising a data cache that corresponds to the different source.

8 . The processor of claim 1 , wherein the load/store unit is configured to ensure ordering of the plurality of atomic elements responsive to the load memory operation receiving bytes from a plurality of cache lines.

9 . The processor of claim 1 , further comprising:

a reorder buffer, wherein the load memory operation is one of a plurality of load memory operations corresponding to the load instruction, wherein the load/store unit is configured to signal a flush of the load memory operation if an ordering violation is detected, and wherein the reorder buffer is configured to enforce in order execution of the plurality of load memory operations responsive to detecting that the load memory operation is within a threshold number of entries of a head of the reorder buffer when the flush is signaled.

10 . A load/store unit, comprising:

a store queue configured to queue one or more store memory operations that write data to one or more memory locations; and

an execution circuit coupled to the store queue and configured to:

execute a load memory operation having an operation size that is larger than an atomicity size implemented for compliance with a memory ordering model;

verify, for compliance with the memory ordering model, that each of a plurality of atomic elements associated with the load memory operation is fully sourced from either a store memory operation in the store queue or from a different source, wherein the plurality of atomic elements are of the atomicity size; and

permit forwarding of data from a particular store memory operation in the store queue for the load memory operation responsive to a verification that each of the plurality of atomic elements is fully sourced from either a store memory operation in the store queue or from a different source.

11 . The load/store unit of claim 10 , wherein the execution circuit is configured to, in response to a detection that at least one of the plurality of atomic elements is not fully sourced from either a store memory operation in the store queue or from a different source, stall the load memory operation until the store queue has emptied.

12 . The load/store unit of claim 10 , wherein the load memory operation targets a plurality of registers, and the atomicity size is a size of a register.

13 . The load/store unit of claim 10 , wherein the load memory operation is a vector load memory operation, and the atomicity size is based on a vector element size of a vector element in a vector that is read by the load memory operation.

14 . The load/store unit of claim 13 , wherein the atomicity size is the vector element size, and a given atomic element of the plurality of atomic elements corresponds to a vector element in the vector.

15 . The load/store unit of claim 13 , wherein the atomicity size is a multiple of the vector element size, and a given atomic element of the plurality of atomic elements corresponds to a plurality of adjacent vector elements in the vector.

16 . The load/store unit of claim 10 , wherein the operation size is an integer multiple of the atomicity size.

17 . A method, comprising:

executing, by a processor, a load memory operation, wherein the load memory operation has an operation size that is larger than an atomicity size implemented for compliance with a memory ordering model, and wherein the load memory operation is one of a plurality of load memory operations corresponding to a load instruction;

detecting, by the processor, an ordering violation based on a plurality of atomic elements associated with the load memory operation;

in response to the detecting, the processor signaling a flush of the load memory operation; and

enforcing, by the processor, in order execution of the plurality of load memory operations responsive to detecting that the load memory operation is within a threshold number of entries of a head of a reorder buffer of the processor when the flush is signaled.

18 . The method of claim 17 , wherein the operation size and an address accessed by the load memory operation cause a crossing of a cache line boundary between a plurality of cache lines, and wherein the ordering violation is detected in response to the plurality of atomic elements receiving bytes from the plurality of cache lines out of order.

19 . The method of claim 17 , wherein the enforcing includes:

tagging, by the processor, the plurality of load memory operations with an in-order-load indication to force the plurality of load memory operations to execute in order.

20 . The method of claim 17 , wherein the threshold number is programmable in the reorder buffer.

Continuity (2)
Continuation 16907740 · Jun 22, 2020
Related Publication 20240248844A1 · Jul 25, 2024
References Cited (65)
US 5265233A · Frailong et al. · 1993 [cited by applicant]
US 6170001B1 · Hinds et al. · 2001 [cited by applicant]
US 6192465B1 · Roberts · 2001 [cited by applicant]
US 6304963B1 · Elwood · 2001 [cited by applicant]
US 8086801B2 · Hrusecky et al. · 2011 [cited by applicant]
US 8656103B2 · Yang et al. · 2014 [cited by applicant]
US 9052889B2 · Jacobi et al. · 2015 [cited by applicant]
US 9092345B2 · Nystad · 2015 [cited by examiner]
US 9411542B2 · Higham et al. · 2016 [cited by applicant]
US 9424034B2 · Hinton et al. · 2016 [cited by applicant]
US 9786338B2 · Hinton et al. · 2017 [cited by applicant]
US 9971626B2 · Busaba et al. · 2018 [cited by applicant]
US 10102888B2 · Hinton et al. · 2018 [cited by applicant]
US 10141033B2 · Hinton et al. · 2018 [cited by applicant]
US 10146538B2 · Sade et al. · 2018 [cited by applicant]
US 10153011B2 · Hinton et al. · 2018 [cited by applicant]
US 10153012B2 · Hinton et al. · 2018 [cited by applicant]
US 10163468B2 · Hinton et al. · 2018 [cited by applicant]
US 10170165B2 · Hinton et al. · 2019 [cited by applicant]
US 10228951B1 · Kothari et al. · 2019 [cited by applicant]
US 11113065B2 · Kalamatianos · 2021 [cited by examiner]
US 11119767B1 · Mestan et al. · 2021 [cited by applicant]
US 11119920B2 · Ros et al. · 2021 [cited by applicant]
US 11314509B2 · Abhishek Raja · 2022 [cited by applicant]
US 11334485B2 · Kaxiras · 2022 [cited by examiner]
US 11841802B2 · Favor · 2023 [cited by examiner]
US 20050149703A1 · Hammong et al. · 2005 [cited by applicant]
US 20050223201A1 · Tremblay et al. · 2005 [cited by applicant]
US 20090019231A1 · Cypher et al. · 2009 [cited by applicant]
US 20100250850A1 · Yang et al. · 2010 [cited by applicant]
US 20100262781A1 · Hruseky et al. · 2010 [cited by applicant]
US 20120173848A1 · Sun et al. · 2012 [cited by applicant]
US 20120290791A1 · Yang et al. · 2012 [cited by applicant]
US 20140181421A1 · O'Connor et al. · 2014 [cited by applicant]
US 20140215191A1 · Kanapathipillai et al. · 2014 [cited by applicant]
US 20150006848A1 · Hinton et al. · 2015 [cited by applicant]
US 20150046655A1 · Nystad · 2015 [cited by examiner]
US 20150067230A1 · Avudaiyappan et al. · 2015 [cited by applicant]
US 20150199272A1 · Goel et al. · 2015 [cited by applicant]
US 20150205605A1 · Abdallah · 2015 [cited by examiner]
US 20150248294A1 · Singh et al. · 2015 [cited by applicant]
US 20160283237A1 · Pardo et al. · 2016 [cited by applicant]
US 20160358636A1 · Hinton et al. · 2016 [cited by applicant]
US 20170286113A1 · Shanbhogue et al. · 2017 [cited by applicant]
US 20180033468A1 · Hinton et al. · 2018 [cited by applicant]
US 20180122429A1 · Hinton et al. · 2018 [cited by applicant]
US 20180122430A1 · Hinton et al. · 2018 [cited by applicant]
US 20180122431A1 · Hinton et al. · 2018 [cited by applicant]
US 20180122432A1 · Hinton et al. · 2018 [cited by applicant]
US 20180122433A1 · Hinton et al. · 2018 [cited by applicant]
US 20180307492A1 · Di · 2018 [cited by applicant]
US 20200192801A1 · Kaxiras · 2020 [cited by examiner]
US 20200319889A1 · Kalamatianos · 2020 [cited by examiner]
US 20210294607A1 · Abhishek · 2021 [cited by applicant]
US 20220091846A1 · Mestan et al. · 2022 [cited by applicant]
US 20220358047A1 · Favor · 2022 [cited by examiner]
CN 102754069A · 2012 [cited by applicant]
CN 104205820A · 2014 [cited by applicant]
GB 2602814A · 2022 [cited by applicant]
KR 1020170131379A · 2017 [cited by applicant]
‘Understanding Atomics and Memory Ordering’ from DEV Community, posted Apr. 8, 2021. (Year: 2021). [cited by examiner]
‘Memory Consistency Models: A Tutorial’ by James Bornholt, Feb. 17, 2016. (Year: 2016). [cited by examiner]
Office Action in Korean Application No. 10-2022-7044214 mailed Sep. 30, 2024, 7 pages. [cited by applicant]
Office Action in United Kingdom Appl. No. 2218443.6 mailed Jul. 31, 2024, 6 pages. [cited by applicant]
ISRWO, PCT/US2021/034512, mailed Sep. 23, 2021, 15 pages. [cited by applicant]