IP Library › Granted Patent US 12,493,801
Granted Patent B2
US 12,493,801 · App. 18/696,993 · Granted Dec 9, 2025

Deep neural network checkpoint optimization system and method based on non-volatile memory

Inventors: Shu Yin (Shanghai, CN); Tianyuan Wu (Shanghai, CN); Yuanhao Li (Shanghai, CN)
Assignee: ShanghaiTech University
G06N3/10G06T1/60
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,493,801
App. No.
18/696,993
Granted
Dec 9, 2025
Kind
B2
Abstract

Deep neural network checkpoint optimization system and method based on non-volatile memory are provided, where the client module and the server module register corresponding network structures in non-volatile memory and create data indexes and data communication protocols based on remote direct memory access (RDMA) before the start of training deep neural networks, and during the neural network training process, provide zero-copy, asynchronous, end-to-end neural network data persistence, which allows users to perform fine-grained checkpointing to ensure fault tolerance and data persistence without affecting the training speed.

Claims (21)

1 . A deep neural network checkpoint optimization system based on non-volatile memory, comprising: a client module in a computing node equipped with a GPU memory, and a server module in a storage node equipped with a non-volatile memory;

wherein before a start of each neural network model training, the client module sends a network structure obtained by initializing a corresponding neural network model stored in the GPU memory to the server module and constructs an index structure of the neural network model on the non-volatile memory, to establish end-to-end communication between the GPU memory and the non-volatile memory; and

wherein when the server module receives a checkpoint request from the client module during training of the neural network model, corresponding model data is directly read from the GPU memory to the non-volatile memory based on the index structure of the neural network model.

2 . The deep neural network checkpoint optimization system based on the non-volatile memory according to claim 1 , wherein a method for the client module to initialize the corresponding neural network model stored in the GPU memory to obtain the network structure comprises:

collecting GPU memory pointers corresponding to each level of the neural network model through a neural network framework;

using an NVIDIA Peer Memory kernel module to register a GPU address space of each level of the neural network model as a RDMA memory region based on the GPU memory pointers corresponding to the level of the neural network model, and providing each RDMA memory region with a unique identifier; and

matching metadata of each level of the neural network model with the unique identifier of the RDMA memory region corresponding to the level of the neural network model, and aggregating all identifier-metadata pairs into a model structure package.

3 . The deep neural network checkpoint optimization system based on the non-volatile memory according to claim 2 , wherein a method for the client module to construct the index structure of the neural network model on the non-volatile memory comprises:

upon receiving the model structure package, selecting a thread from a thread pool, and constructing the index structure corresponding to the neural network model on the non-volatile memory by this thread based on the model structure package, to respectively map each level of the neural network model to a checkpoint structure.

4 . The deep neural network checkpoint optimization system based on the non-volatile memory according to claim 3 , wherein the index structure is a three-level index structure comprising a model table at a first level, model metadata at a second level, and model data information at a third level.

5 . The deep neural network checkpoint optimization system based on the non-volatile memory according to claim 4 ,

wherein before the server module receives the checkpoint request from the client module, the client module obtains a corresponding GPU memory pointer, and generates and sends the checkpoint request to the server module after the client module receives a user checkpoint request during the neural network model training process;

wherein a method for the server module to directly read the corresponding model data from the GPU memory to the non-volatile memory based on the index structure of the neural network model comprises: according to the checkpoint request from the client module, controlling a corresponding thread and directly reading the corresponding model data from the GPU memory to the non-volatile memory through a RDMA read operation based on the index structure.

6 . The deep neural network checkpoint optimization system based on the non-volatile memory according to claim 1 , wherein when the server module receives a data recovery request from the client module, the server module actively writes the corresponding model data from the non-volatile memory into the GPU memory, based on the index structure of the neural network model.

7 . The deep neural network checkpoint optimization system based on the non-volatile memory according to claim 1 , wherein the client module communicates with the server module via a TCP protocol.

8 . The deep neural network checkpoint optimization system based on the non-volatile memory according to claim 2 , wherein the neural network framework is a PyTorch software library.

9 . A deep neural network checkpoint optimization method based on non-volatile memory, which is applied to a deep neural network checkpoint optimization system based on non-volatile memory comprising a client module in a computing node equipped with a GPU memory and a server module in a storage node equipped with a non-volatile memory, the method comprising:

before a start of training each neural network model, sending a network structure obtained by initializing one corresponding neural network model stored in the GPU memory to the server module by the client module, constructing an index structure of the neural network model on the non-volatile memory, and establishing end-to-end communication between the GPU memory and the non-volatile memory; and

when the server module receives a checkpoint request from the client module during training of the neural network model, directly reading corresponding model data from the GPU memory to the non-volatile memory based on the index structure of the neural network model.

10 . The deep neural network checkpoint optimization method based on non-volatile memory according to claim 9 , further comprising:

when the server module receives a data recovery request from the client module, based on the index structure of the neural network model, actively writing the corresponding model data from the non-volatile memory into the GPU memory.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Apr 3, 2024
From: YIN, SHU; WU, TIANYUAN; LI, YUANHAO
To: SHANGHAITECH UNIVERSITY
Reel/Frame 066985/0576 →
Priority Claims (1)
CN 202310153129.5 · Feb 22, 2023 · national
Continuity (1)
Related Publication 20250371374A1 · Dec 4, 2025
References Cited (9)
US 10776164B2 · Zhao · 2020 [cited by examiner]
US 20210383240A1 · Gaidon et al. · 2021 [cited by applicant]
CN 111078607A · 2020 [cited by applicant]
CN 114282665A · 2022 [cited by applicant]
CN 115310605A · 2022 [cited by applicant]
CN 115345285A · 2022 [cited by applicant]
Rojas et al (“A Study of Checkpointing in Large Scale Training of Deep Neural Networks” 2021) (Year: 2021). [cited by examiner]
Wang et al (“Enabling Efficient Large-Scale Deep Learning Training with Cache Coherent Disaggregated Memory Systems” 2022) (Year: 2022). [cited by examiner]
Mohan et al (“CheckFreq: Frequent, Fine-Grained DNN Checkpointing” 2021) (Year: 2021). [cited by examiner]