IP Library Granted Patent US 12,474,926
Granted Patent B2
US 12,474,926 · App. 18/190,620 · Granted Nov 18, 2025

Fused data generation and associated communication

Inventors: Shaizeen Dilawarhusen Aga (Santa Clara, CA); Suchita Pati (Madison, WI); Nuwan S. Jayasena (Cupertino, CA)
Assignee: Advanced Micro Devices, Inc.
G06F9/30036G06F9/3834
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,474,926
App. No.
18/190,620
Granted
Nov 18, 2025
Kind
B2
Abstract

Fused data generation and associated communication techniques are described. In an implementation, a system includes processing system having a plurality of processors. A data generation and communication tracking module is configured to track programmatically defined data generation and associated communication as performed by the plurality of processors. A targeted communication module is configured to trigger targeted communication of data between the plurality of processors based on the tracked programmatically defined data generation and associated communication.

Claims (31)

1 . A system comprising:

a processing system including a plurality of processors for training a machine learning model; and

a memory controller configured to:

track data generation by the plurality of processors and associated communication of the generated data between the plurality of processors during the training of the machine learning model; and

trigger targeted communication of the generated data between the plurality of processors based on the tracked data generation and associated communication.

2 . The system of claim 1 , wherein the data generation and associated communication of the generated data includes generation of the data by at least one processor of the plurality of processors and a targeted update to transmit the data by the at least one processor to another processor of the plurality of processors.

3 . The system of claim 2 , wherein the targeted update is triggered upon completion of the generation of the data by the at least one processor.

4 . The system of claim 2 , wherein the targeted update is triggered based on a remote communication event received at the at the at least one processor as implemented by a data mover engine by another processor of the plurality of processors.

5 . The system of claim 4 , wherein the remote communication event is part of a bulk operation involving communication of the data.

6 . The system of claim 1 , wherein the data generation and associated communication of the generated data are defined using a single fused data generation and associated communication operation.

7 . The system of claim 6 , wherein the fused data generation and associated communication operation identifies another processor of the plurality of processors to receive the data.

8 . The system of claim 6 , wherein the fused data generation and associated communication operation identifies an address range that is a source of the data or an address range that is a destination to transmit the data.

9 . The system of claim 1 , wherein at least one processor of the plurality of processors is configured to support concurrent updates to the data in physical memory.

10 . The system of claim 9 , wherein a processor-in-memory component of a memory module that includes the physical memory is configured to implement the concurrent updates.

11 . The system of claim 1 , wherein the data generation and associated communication of the generated data is configured to control a data generation order by respective processors of the plurality of processors.

12 . The system of claim 1 , wherein the training includes multiple stages and wherein targeted communication is triggered by causing communication of data generated by a first stage of the multiple stages of the training to overlap data generation of a next stage of the multiple stages of the training.

13 . The system of claim 12 , wherein the training comprises large network matrix multiplication (GEMM) operations, and wherein a GEMM write of generated data in the first stage automatically triggers communication of the generated data.

14 . The system of claim 1 , wherein the at memory controller is configured to track data generation and associated communication of the generated data by monitoring a table structure.

15 . The system of claim 14 , wherein the table structure includes table entries for each of the plurality of processors.

16 . The system of claim 14 , wherein the memory controller is configured to trigger the targeted communication by causing communication of generated data to specific processors of the plurality of processors.

17 . A device comprising:

a processing system including a plurality of processors for training a machine learning model; and

a memory controller configured to:

trigger communication of generated data between at least one processor and another processor of plurality of processors as part of data generation and associated communication of the generated data during the training of the machine learning model; and

resolve concurrent updates to the generated data in physical memory.

18 . The device of claim 17 , wherein a processor-in-memory component of a memory module that includes the physical memory is configured to resolve the concurrent updates to the data in the physical memory.

19 . A method comprising:

tracking data generation and associated communication as performed between a plurality of processors of a processing system during training of a machine learning model;

triggering targeted communication of data between the plurality of processors as part of the data generation and associated communication; and

resolving concurrent updates to physical memory involving the data generated by the plurality of processors.

20 . The method of claim 19 , wherein the data generation and associated communication is configured to identify an address range that is a source of the data or an address range that is a destination to transmit the data.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Mar 28, 2023
From: AGA, SHAIZEEN DILAWARHUSEN; PATI, SUCHITA; JAYASENA, NUWAN S.
To: ADVANCED MICRO DEVICES, INC.
Reel/Frame 063135/0630 →
Continuity (2)
Provisional Application 63387434 · Dec 14, 2022
Related Publication 20240201990A1 · Jun 20, 2024
References Cited (21)
US 6417797B1 · Cousins et al. · 2002 [cited by applicant]
US 8713197B1 · Li · 2014 [cited by applicant]
US 9160607B1 · Froese · 2015 [cited by examiner]
US 9632832B2 · Solihin · 2017 [cited by examiner]
US 10771589B1 · Brito · 2020 [cited by examiner]
US 11442945B1 · Tsai · 2022 [cited by examiner]
US 20080046666A1 · Termaine · 2008 [cited by examiner]
US 20110107059A1 · Oh et al. · 2011 [cited by applicant]
US 20160162424A1 · Dobbs et al. · 2016 [cited by applicant]
US 20180107510A1 · Carlough et al. · 2018 [cited by applicant]
US 20180329958A1 · Choudhury · 2018 [cited by examiner]
US 20190102574A1 · Roberts · 2019 [cited by examiner]
US 20210027136A1 · Hwang · 2021 [cited by examiner]
US 20220029986A1 · Neumann · 2022 [cited by examiner]
US 20230195375A1 · Puthoor · 2023 [cited by examiner]
US 20230342588A1 · Kramer · 2023 [cited by examiner]
CN 120344964A · 2025 [cited by applicant]
Gasparakis, Harris , “US Application as Filed”, U.S. Appl. No. 18/091,441, filed Dec. 30, 2022, 48 pages. [cited by applicant]
Pati, Suchita , et al., “US Application as Filed”, U.S. Appl. No. 18/091,443, filed Dec. 30, 2022, 48 pages. [cited by applicant]
Rajbhandari, Samyam , et al., “ZeRO: Memory Optimizations Toward Training Trillion Parameter Models”, Cornell University arXiv, arXiv.org [retrieved May 4, 2023]. Retrieved from the Internet <https://arxiv.org/pdf/1910.… [cited by applicant]
PCT/US2023/082982 , “International Search Report and Written Opinion”, International Application No. PCT/US2023/082982, Apr. 19, 2024, 8 pages. [cited by applicant]