IP Library Granted Patent US 12,373,716
Granted Patent B2
US 12,373,716 · App. 17/072,757 · Granted Jul 29, 2025

Automated synchronization of clone directed acyclic graphs

Inventor: Bradley David Safnuk (Victoria, CA)
Assignee: Ford Global Technologies, LLC
G06N7/01G06N3/045G06N3/08G06N3/084G06N3/098
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,373,716
App. No.
17/072,757
Granted
Jul 29, 2025
Kind
B2
Abstract

A method is disclosed for synchronization of clone directed acyclic graphs. The method can include identifying a directed acyclic graph (“DAG”) including a plurality of vertices linked in pairwise relationships via a plurality of edges. At least one clone DAG can be created, which at least one clone DAG can be identical to at least a portion of the DAG. For each of the vertices of the DAG, a corresponding clone vertex from the clone vertices of the at least one clone DAG can be identified. Aggregate gradient data can be calculated based on gradient data from each of the clone vertices and its corresponding vertex in the DAG, and at least one weight of the DAG and of the at least one clone DAG can be updated based on the aggregate gradient data.

Claims (38)

1. A method of distributed computing, the method comprising:

identifying a directed acyclic graph (“DAG”), the DAG comprising a plurality of vertices linked in pairwise relationships via a plurality of edges;

creating at least one clone DAG identical to at least a portion of the DAG, the at least one clone DAG comprising a plurality of clone vertices;

for each of the vertices of the DAG, identifying a corresponding clone vertex from the clone vertices of the at least one clone DAG;

training the DAG and the least one clone DAG, each of the vertices of the DAG and each of the corresponding clone vertices of the at least one clone DAG have a gradient as a result of the training;

traversing the DAG and the at least one clone DAG in a first direction, wherein the clone vertices of the at least one clone DAG and its corresponding vertices of the DAG exchange first gradient data during traversal of the DAG and the at least one clone DAG;

calculating first gradient data for related nodes of the DAG and the at least one clone DAG based on the first gradient data exchanged between each of the clone vertices and its corresponding vertex in the DAG during traversal of the DAG and the at least one clone DAG in the first direction;

traversing the DAG and the at least one clone DAG in a second direction that is a reverse direction with respect to the first direction;

calculating second gradient data for related nodes of the DAG and the at least one clone DAG based on second gradient data exchanged between each of the clone vertices and its corresponding vertex in the DAG during traversal of the DAG and the at least one clone DAG in the second direction;

calculating aggregate gradient data based on the first gradient data and the second gradient data; and

updating at least one weight of the DAG and of the at least one clone DAG based on the aggregate gradient data.

2. The method of claim 1 , further comprising identifying the vertices of the DAG, and wherein aggregate gradient data based on gradient data from each of the clone vertices and its corresponding vertex in the DAG are calculated during training.

3. The method of claim 2 , wherein identifying the corresponding clone vertex from the clone vertices of the at least one clone DAG comprises: applying incrementing naming across the clone vertices of the at least one clone DAG; notifying vertices of the DAG of their corresponding clone vertices of the at least one clone DAG; and notifying clone vertices of the at least one clone DAG of their corresponding vertices of the DAG.

4. The method of claim 1 , wherein training each of the DAG and the at least one clone DAG comprises: ingesting first data into the DAG and ingesting second data into the at least one clone DAG.

5. The method of claim 4 , wherein the first data and the second data are non-identical.

6. The method of claim 1 , wherein calculating aggregate gradient data based on the first gradient data and the second gradient data comprises calculating mean gradient data.

7. The method of claim 1 , wherein the updating at least one weight of the DAG and of the at least one clone DAG based on the aggregate gradient data is performed according to synchronous gradient updates.

8. The method of claim 1 , wherein the updating at least one weight of the DAG and of the at least one clone DAG based on the aggregate gradient data is performed according to asynchronous gradient updates.

9. A system comprising:

one or more processors; and

one or more memories storing computer-executable instructions that, when executed by the one or more processors, configure the one or more processors to:

identify a directed acyclic graph (“DAG”), the DAG comprising a plurality of vertices linked in pairwise relationships via a plurality of edges;

create at least one clone DAG identical to at least a portion of the DAG, the at least one clone DAG comprising a plurality of clone vertices;

for each of the vertices of the DAG, identify a corresponding clone vertex from the clone vertices of the at least one clone DAG;

train the DAG and the least one clone DAG, each of the vertices of the DAG and each of the corresponding clone vertices of the at least one clone DAG have a gradient as a result of the training;

traverse the DAG and the at least one clone DAG in a first direction, wherein the clone vertices of the at least one clone DAG and its corresponding vertices of the DAG exchange first gradient data during traversal of the DAG and the at least one clone DAG;

calculate first gradient data for related nodes of the DAG and the at least one clone DAG based on the first gradient data exchanged between each of the clone vertices and its corresponding vertex in the DAG during traversal of the DAG and the at least one clone DAG in the first direction;

traverse the DAG and the at least one clone DAG in a second direction that is a reverse direction with respect to the first direction;

calculate second gradient data for related nodes of the DAG and the at least one clone DAG based on second gradient data exchanged between each of the clone vertices and its corresponding vertex in the DAG during traversal of the DAG and the at least one clone DAG in the second direction;

calculate aggregate gradient data based on the first gradient data and the second gradient data; and

update at least one weight of the DAG and of the at least one clone DAG based on the aggregate gradient data.

10. The system of claim 9 , wherein the computer-executable instructions that, when executed by the one or more processors, further configure the one or more processors to identify the vertices of the DAG, and wherein aggregate gradient data based on gradient data from each of the clone vertices and its corresponding vertex in the DAG are calculated during training.

11. The system of claim 10 , wherein identifying the corresponding clone vertex from the clone vertices of the at least one clone DAG comprises: applying incrementing naming across the clone vertices of the at least one clone DAG; notifying vertices of the DAG of their corresponding clone vertices of the at least one clone DAG; and notifying clone vertices of the at least one clone DAG of their corresponding vertices of the DAG.

12. The system of claim 9 , wherein training each of the DAG and the at least one clone DAG comprises: ingesting first data into the DAG and ingesting second data into the at least one clone DAG.

13. The system of claim 12 , wherein the first data and the second data are non-identical.

14. The system of claim 9 , wherein calculating aggregate gradient data based on the first gradient data and the second gradient data comprises calculating mean gradient data.

15. The system of claim 9 , wherein the updating at least one weight of the DAG and of the at least one clone DAG based on the aggregate gradient data is configured to be performed according to synchronous gradient updates.

16. The system of claim 9 , wherein the updating at least one weight of the DAG and of the at least one clone DAG based on the aggregate gradient data is configured to be performed according to asynchronous gradient updates.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Oct 16, 2020
From: SAFNUK, BRADLEY DAVID
To: FORD GLOBAL TECHNOLOGIES, LLC
Reel/Frame 054080/0807 →
Continuity (1)
Related Publication 20220121974A1 · Apr 21, 2022
References Cited (28)
US 8229968B2 · Wang et al. · 2012 [cited by applicant]
US 9811775B2 · Krizhevsky et al. · 2017 [cited by applicant]
US 10325340B2 · Wu et al. · 2019 [cited by applicant]
US 10540588B2 · Burger et al. · 2020 [cited by applicant]
US 10846245B2 · Islam et al. · 2020 [cited by applicant]
US 20170038919A1 · Moss · 2017 [cited by examiner]
US 20190042946A1 · Sur · 2019 [cited by examiner]
US 20190114534A1 · Teng et al. · 2019 [cited by applicant]
CN 110941494A · 2020 [cited by applicant]
WO 2016186801A1 · 2016 [cited by applicant]
WO WO2019042571A1 · 2019 [cited by examiner]
Lazo, Visual intuition on ring-Allreduce for distributed Deep Learning, Towards Data Science, 2019 (Year: 2019). [cited by examiner]
Qi Intro Distributed Deep Learning _ Xiandong, https://xiandong79.github.io/ 2017 (Year: 2017). [cited by examiner]
Zhao, Sparse Allreduce: Efficient Scalable Communication for Power-Law Data, arXiv, 2013 (Year: 2013). [cited by examiner]
Naumov, Maxim, Alysson Vrielink, and Michael Garland. “Parallel depth-first search for directed acyclic graphs.” Proceedings of the Seventh Workshop on Irregular Applications: Architectures and Algorithms. 2017. https:/… [cited by examiner]
Shi, S., et al., “A DAG Model of Synchronous Stochastic Gradient Descent in Distributed Deep Learning”, 2018 IEEE 24th International Conference on Parallel and Distributed Systems (ICPADS), Last revised on Oct. 31, 2018… [cited by applicant]
Eustace, P., et al., “SpiNNaker: A 1-W 18-Core System-on-Chip for Massively-Parallel Neural Network Simulation,” in IEEE Journal of Solid-State Circuits, vol. 48, No. 8, pp. 1943-1953, Aug. 2013, DOI: 10.1109/JSSC.2013.… [cited by applicant]
Takahiro, S., et al., “Structure discovery of deep neural network based on evolutionary algorithms,” 2015 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), Brisbane, QLD, 2015, pp. 4979-… [cited by applicant]
Hedge, V., et al., “Parallel and Distributed Deep Learning”, (2016), 8 pages. [cited by applicant]
Ben-Nun, T., et al., “Demystifying Parallel and Distributed Deep Learning: An In-depth Concurrency Analysis”, ACM Computing Surveys, vol. 52, Issue 4, Article 65, (Last revised Sep. 15, 2018), 43 pages. arXiv:1802.09941… [cited by applicant]
Zhang, J., et al., “A Parallel Strategy for Convolutional Neural Network Based on Heterogeneous Cluster for Mobile Information System”, Hindawi, Mobile InformationSystems, vol. 2017, Article ID 3824765, published Mar. 2… [cited by applicant]
Costa et al., “Evaluating the performance of convolutional neural networks with direct acyclic graph architectures in automatic segmentation of breast lesion in US images.” BMC Medical Imaging 19 (2019): 1-13. (Year: 20… [cited by applicant]
Mayer et al., “The TensorFlow partitioning and scheduling problem: it's the critical path!.”, In Proceedings of the 1st Workshop on Distributed Infrastructures for Deep Learning (pp. 1-6). (Year: Dec. 2017). [cited by applicant]
U.S. Appl. No. 17/072,709, filed Oct. 16, 2020, Non-Final Office Action mailed Jul. 28, 2023, all pages. [cited by applicant]
U.S. Appl. No. 17/072,709 , Final Office Action, Mailed On Mar. 28, 2024, 46 pages. [cited by applicant]
U.S. Appl. No. 17/072,709 , Non-Final Office Action, Mailed On Nov. 6, 2024, 23 pages. [cited by applicant]
Abadi et al., “TensorFlow: Large-Scale Machine Learning on Heterogeneous Distributed Systems”, Available Online at: https://arxiv.org/pdf/1603.04467, Mar. 16, 2016, pp. 1-19. [cited by applicant]
Veeramanikandan et al., “Data Flow and Distributed Deep Neural Network Based Low Latency IoT-edge Computation Model for Big Data Environment”, Engineering Applications of Artificial Intelligence, vol. 94, No. 2, Sep. 20… [cited by applicant]