IP Library › Granted Patent US 12,645,942
Granted Patent B2
US 12,645,942 · App. 18/385,263 · Granted Jun 2, 2026

Replacement of neural network gradient values in transmission data

Inventor: Dennis Charles Abts (Rochester, MN)
Assignee: NVIDIA Corporation
G06N3/084G06N3/045G06N3/10H04L43/0847
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,645,942
App. No.
18/385,263
Granted
Jun 2, 2026
Kind
B2
Abstract

Apparatuses, systems, and techniques to replace one or more corrupt neural network gradient values with one or more corresponding surrogate neural network gradient values. In at least one embodiment, faulty gradient values are identified in a transmitted data packet and are replaced according to a replacement policy with a suitable surrogate value.

Claims (26)

1 . One or more processors, comprising:

circuitry to replace one or more neural network gradient values with one or more corresponding surrogate neural network gradient values, wherein the circuitry is to train a neural network using the one or more surrogate neural network gradient values, and wherein the one or more corresponding surrogate neural network gradient values are all zero values or are based, at least in part, on at least one of an average of recently received gradient values or a gradient value most recently received.

2 . The one or more processors of claim 1 , wherein the one or more neural network gradient values are transmitted in a packet, and wherein the circuitry is to perform a cyclic redundancy check on a first checksum value of a packet header portion and a second checksum value of a packet data portion.

3 . The one or more processors of claim 1 , wherein the circuitry is to store the one or more corresponding surrogate neural network gradient values in memory.

4 . The one or more processors of claim 1 , wherein the average of the recently received gradient values is computed based on a mean and standard deviation of the recently received gradient values.

5 . The one or more processors of claim 1 , wherein the circuitry is to perform forward error correction (FEC) on the one or more neural network gradient values prior to replacement of the one or more neural network gradient values with the one or more corresponding surrogate neural network gradient values.

6 . The one or more processors of claim 1 , wherein the circuitry is to generate a neural network loss function using stochastic gradient descent based, at least in part, on the one or more surrogate neural network gradient values.

7 . The one or more processors of claim 1 , wherein the one or more neural network gradient values have been corrupted during transmission within a communication network.

8 . A system comprising:

one or more processors to replace one or more neural network gradient values with one or more corresponding surrogate neural network gradient values, wherein the one or more processors are to train a neural network using the one or more surrogate neural network gradient values, and wherein the one or more corresponding surrogate neural network gradient values are all zero values or are based, at least in part, on at least one of an average of recently received gradient values or a gradient value most recently received.

9 . The system of claim 8 , wherein the one or more neural network gradient values are transmitted in a packet, and wherein the one or more processors are to perform a cyclic redundancy check on a first checksum value of a packet header portion and a second checksum value of a packet data portion.

10 . The system of claim 8 , wherein the one or more processors are to store the one or more corresponding surrogate neural network gradient values in memory.

11 . The system of claim 8 , wherein the average of the recently received gradient values is computed based on a mean and standard deviation of the recently received gradient values.

12 . The system of claim 8 , wherein the one or more processors are to perform forward error correction (FEC) on the one or more neural network gradient values prior to replacement of the one or more neural network gradient values with the one or more corresponding surrogate neural network gradient values.

13 . The system of claim 8 , wherein the one or more processors are to generate a neural network loss function using stochastic gradient descent based, at least in part, on the one or more surrogate neural network gradient values.

14 . The system of claim 8 , wherein the one or more neural network gradient values have been corrupted during transmission within a communication network.

15 . A method comprising:

replacing one or more neural network gradient values with one or more corresponding surrogate neural network gradient values; and

training a neural network using the one or more surrogate neural network gradient values, wherein the one or more corresponding surrogate neural network gradient values are all zero values or are based, at least in part, on at least one of an average of recently received gradient values or a gradient value most recently received.

16 . The method of claim 15 , further comprising:

transmitting the one or more neural network gradient values in a packet; and

performing a cyclic redundancy check on a first checksum value of a packet header portion and a second checksum value of a packet data portion.

17 . The method of claim 15 , further comprising storing the one or more corresponding surrogate neural network gradient values in memory.

18 . The method of claim 15 , further comprising computing the average of the recently received gradient values based on a mean and standard deviation of the recently received gradient values.

19 . The method of claim 15 , further comprising performing forward error correction (FEC) on the one or more neural network gradient values prior to replacement of the one or more neural network gradient values with the one or more corresponding surrogate neural network gradient values.

20 . The method of claim 15 , further comprising generating a neural network loss function using stochastic gradient descent based, at least in part, on the one or more surrogate neural network gradient values.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Nov 13, 2023
From: ABTS, DENNIS CHARLES
To: NVIDIA CORPORATION
Reel/Frame 065539/0719 →
Continuity (1)
Related Publication 20250139439A1 · May 1, 2025
References Cited (15)
US 20240202505A1 · Keuninckx · 2024 [cited by examiner]
US 20240412042A1 · Savinov · 2024 [cited by examiner]
US 20250005341A1 · Kim · 2025 [cited by examiner]
US 20250086449A1 · Deng · 2025 [cited by examiner]
Brownlee, Jason, A Gentle Introduction to Exploding Gradients in Neural Networks, Aug. 14, 2019, https://machinelearningmastery.com/exploding-gradients-in-neural-networks/. [cited by examiner]
IEEE “IEEE Standard for Floating-Point Arithmetic”, Microprocessor Standards Committee of the IEEE Computer Society, IEEE Std 754-2008, dated Jun. 12, 2008, 70 pages. [cited by applicant]
Jouppi et al., “TPU v4: An Optically Reconfigurable Supercomputer for Machine Learning with Hardware Support for Embeddings,” Apr. 20, 2023, 14 Pages. [cited by applicant]
Society of Automotive Engineers On-Road Automated Vehicle Standards Committee “Taxonomy and Definitions for Terms Related to Driving Automation Systems for On-Road Motor Vehicles”, Standard No. J3016-201806, dated Jun. … [cited by applicant]
Society of Automotive Engineers On-Road Automated Vehicle Standards Committee, “Taxonomy and Definitions for Terms Related to Driving Automation Systems for On-Road Motor Vehicles,” Standard No. J3016-201609, issued Jan… [cited by applicant]
Srivastava et al., “Dropout: A Simple Way to Prevent Neural Networks from Overfitting,” Journal of Machine Learning Research, 15(56) 2014, 30 pages. [cited by applicant]
Wangni et al., “Gradient Sparsification for Communication-Efficient Distributed Optimization,” Oct. 26, 2017, 22 Pages. [cited by applicant]
International Search Report and Written Opinion for Application No. PCT/US2024/053725, mailed Apr. 29, 2025, 13 pages. [cited by applicant]
Chen et al., “Boosting Distributed Machine Learning Training Through Loss-tolerant Transmission Protocol,” IEEE/ACM 31st International Symposium on Quality of Service, 2023, 10 pages. [cited by applicant]
Ma et al., “Approximate Wireless Communication for Federated Learning,” Data Science, E-Learning and Information Systems, 2023, 6 pages. [cited by applicant]
Wang et al., “NeuroMessenger: Towards Error Tolerant Distributed Machine Learning Over Edge Networks,” IEEE Conference on Computer Communications, 2022, 10 pages. [cited by applicant]