IP Library › Granted Patent US 12,743,876
Granted Patent B2
US 12,743,876 · App. 18/203,562 · Granted Sep 22, 2026

Dynamic class weighting for training one or more neural networks

Inventors: Yue Zhu (Shanghai, CN); Yongjun Huang (Santa Clara, CA); Yongliang Wu (Santa Clara, CA)
Assignee: NVIDIA Corporation
G06V10/82
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,743,876
App. No.
18/203,562
Granted
Sep 22, 2026
Kind
B2
Abstract

Apparatuses, systems, and techniques are presented to train neural networks and use those neural networks for inferencing tasks. In at least one embodiment, one or more neural networks are caused to be trained using weight parameters based, at least in part, on an amount of training data used to train the one or more neural networks.

Claims (59)

1 . A processor, comprising:

one or more circuits to cause one or more neural networks to be trained to classify data according to a plurality of classes, wherein to cause the one or more neural networks to be trained the one or more circuits are further to:

calculate respective losses according to a loss function for different classes of the plurality of classes based on inferences of the different classes by the one or more neural networks, wherein calculation of the respective losses comprises adjusting a respective loss for each of at least two of the different classes using a different class-specific weight value parameter for each of the at least two different classes during training;

cause network parameters of the one or more neural networks to be updated during the training of the one or more neural networks based on the respective losses as adjusted using the different class-specific weight values; and

update one or more of the class-specific weight values for subsequent training based, at least in part, on an amount of loss according to the loss function for a corresponding class of the at least two different classes during the training.

2 . The processor of claim 1 , wherein the class-specific weight values correspond to different classes of objects to be detected by the one or more neural networks.

3 . The processor of claim 2 , wherein the one or more circuits are further to adjust the class-specific weight values during training based, at least in part, upon the relative amount of training data in each of the different classes of objects.

4 . The processor of claim 3 , wherein the one or more circuits are further to determine offsets from a set of default anchors to one or more objects detected in the individual training images.

5 . The processor of claim 4 , wherein the one or more circuits are further to determine a loss per class per anchor for the different classes of objects, and to adjust the class-specific weight values based at least in part upon the determined losses per class per anchor.

6 . The processor of claim 2 , wherein the one or more circuits are further to adjust the class-specific weight values using a momentum vector for each of a number of training epochs.

7 . A system, comprising:

one or more processors to cause one or more neural networks to be trained to classify data according to a plurality of classes, wherein to cause the one or more neural networks to be trained the one or more processors are further to:

calculate respective losses according to a loss function for different classes of the plurality of classes based on inferences of the different classes by the one or more neural networks, wherein calculation of the respective losses comprises adjusting a respective loss for each of at least two of the different classes using a different class-specific weight value for each of the at least two different classes during training;

cause network parameters of the one or more neural networks to be updated during the training of the one or more neural networks based on the respective losses as adjusted using the different class-specific weight values; and

update one or more of the class-specific weight values for subsequent training based, at least in part, on an amount of loss according to the loss function for a corresponding class of the at least two different classes during the training.

8 . The system of claim 7 , wherein the class-specific weight values correspond to different classes of objects to be detected by the one or more neural networks.

9 . The system of claim 8 , wherein the one or more processors are further to adjust the class-specific weight values during training based, at least in part, upon the relative amount of training data in each of the different classes of objects.

10 . The system of claim 9 , wherein the one or more processors are further to determine offsets from a set of default anchors to one or more objects detected in the individual training images.

11 . The system of claim 10 , wherein the one or more processors are further to determine a loss per class per anchor for the different classes of objects, and to adjust the class-specific weight values based at least in part upon the determined losses per class per anchor.

12 . The system of claim 8 , wherein the one or more processors are further to adjust the class-specific weight values using a momentum vector for each of a number of training epochs.

13 . A method, comprising:

causing one or more neural networks to be trained to classify data according to a plurality of classes, wherein said causing comprises:

calculating respective losses according to a loss function for different classes of the plurality of classes based on inferences of the different classes by the one or more neural networks, wherein calculation of the respective losses comprises adjusting a respective loss for each of at least two of the different classes using a different class-specific weight value for each of the at least two different classes during training;

causing network parameters of the one or more neural networks to be updated during the training of the one or more neural networks based on the respective losses as adjusted using the different class-specific weight values; and

updating one or more of the class-specific weight values for subsequent training based, at least in part, on an amount of loss according to the loss function for a corresponding class of the at least two different classes during the training.

14 . The method of claim 13 , wherein the class-specific weight values correspond to different classes of objects to be detected by the one or more neural networks.

15 . The method of claim 14 , further comprising:

adjusting the class-specific weight values during training based, at least in part, upon the relative amount of training data in each of the different classes of objects.

16 . The method of claim 15 , further comprising:

determining offsets from a set of default anchors to one or more objects detected in the individual training images.

17 . The method of claim 16 , further comprising:

determining a loss per class per anchor for the different classes of objects, and to adjust the class-specific weight values based at least in part upon the determined losses per class per anchor.

18 . The method of claim 14 , further comprising:

adjusting the class-specific weight values using a momentum vector for each of a number of training epochs.

19 . A non-transitory machine-readable medium having stored thereon a set of instructions, which if performed by one or more processors, cause the one or more processors to at least:

cause one or more neural networks to be trained to classify data according to a plurality of classes, wherein to cause the one or more neural networks to be trained the instructions if performed further cause the one or more processors to:

calculate respective losses according to a loss function for different classes of the plurality of classes based on inferences of the different classes by the one or more neural networks, wherein calculation of the respective losses comprises adjusting a respective loss for each of at least two of the different classes using a different class-specific weight value for each of the at least two different classes during training;

cause network parameters of the one or more neural networks to be updated during the training of the one or more neural networks based on the respective losses as adjusted using the different class-specific weight values; and

update one or more of the class-specific weight values for subsequent training based, at least in part, on an amount of loss according to the loss function for a corresponding class of the at least two different classes during the training.

20 . The non-transitory machine-readable medium of claim 19 , wherein the class-specific weight values correspond to different classes of objects to be detected by the one or more neural networks.

21 . The non-transitory machine-readable medium of claim 20 , wherein the instructions if performed further cause the one or more processors to:

adjust the class-specific weight values during training based, at least in part, upon the relative amount of training data in each of the different classes of objects.

22 . The non-transitory machine-readable medium of claim 21 , wherein the instructions if performed further cause the one or more processors to:

determine offsets from a set of default anchors to one or more objects detected in the individual training images.

23 . The non-transitory machine-readable medium of claim 22 , wherein the instructions if performed further cause the one or more processors to:

determine a loss per class per anchor for the different classes of objects, and to adjust the class-specific weight values based at least in part upon the determined losses per class per anchor.

24 . The non-transitory machine-readable medium of claim 20 , wherein the instructions if performed further cause the one or more processors to:

adjust the class-specific weight values using a momentum vector for each of a number of training epochs.

25 . A network training system, comprising:

one or more processors to cause one or more neural networks to be trained to classify data according to a plurality of classes, wherein to cause the one or more neural networks to be trained the one or more processors are further to:

calculate respective losses according to a loss function for different classes of the plurality of classes based on inferences of the different classes by the one or more neural networks, wherein calculation of the respective losses comprises adjusting a respective loss for each of at least two of the different classes using a different class-specific weight value for each of the at least two different classes during training;

cause network parameters of the one or more neural networks to be updated during the training of the one or more neural networks based on the respective losses as adjusted using the different class-specific weight values; and

update one or more of the class-specific weight values for subsequent training based, at least in part, on an amount of loss according to the loss function for a corresponding class of the at least two different classes during the training; and

memory for storing the network parameters for the one or more neural networks.

26 . The network training system of claim 25 , wherein the class-specific weight values correspond to different classes of objects to be detected by the one or more neural networks.

27 . The network training system of claim 26 , wherein the one or more processors are further to adjust the class-specific weight values during training based, at least in part, upon the relative amount of training data in each of the different classes of objects.

28 . The network training system of claim 27 , wherein the one or more processors are further to determine offsets from a set of default anchors to one or more objects detected in the individual training images.

29 . The network training system of claim 28 , wherein the one or more processors are further to determine a loss per class per anchor for the different classes of objects, and to adjust the class-specific weight values based at least in part upon the determined losses per class per anchor.

30 . The network training system of claim 26 , wherein the one or more processors are further to adjust the class-specific weight values using a momentum vector for each of a number of training epochs.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jun 27, 2023
From: ZHU, YUE; HUANG, YONGJUN; WU, YONGLIANG
To: NVIDIA CORPORATION
Reel/Frame 064085/0510 →
Continuity (2)
Continuation PCTCN2022095881 · May 30, 2022
Related Publication 20230386191A1 · Nov 30, 2023
References Cited (23)
US 10217346B1 · Zhang · 2019 [cited by examiner]
US 12169875B2 · Chen · 2024 [cited by examiner]
US 20190287264A1 · Chalamala · 2019 [cited by examiner]
US 20190340492A1 · Burger · 2019 [cited by examiner]
US 20200074234A1 · Tong · 2020 [cited by examiner]
US 20200184268A1 · Lewis · 2020 [cited by examiner]
US 20210264300A1 · Staudinger · 2021 [cited by examiner]
US 20210287089A1 · Mayer · 2021 [cited by examiner]
US 20210406603A1 · Narwariya · 2021 [cited by examiner]
US 20210407090A1 · Li · 2021 [cited by examiner]
US 20220125400A1 · Wu · 2022 [cited by examiner]
US 20220187076A1 · Gonzalez Diaz · 2022 [cited by examiner]
US 20230085401A1 · Deng · 2023 [cited by examiner]
US 20230306050A1 · Gupta · 2023 [cited by examiner]
CN 109902722A · 2019 [cited by applicant]
CN 110705685A · 2020 [cited by applicant]
CN 111144566A · 2020 [cited by applicant]
EP 3828777A1 · 2021 [cited by applicant]
EP 3832341A1 · 2021 [cited by applicant]
IEEE “IEEE Standard for Floating-Point Arithmetic”, Microprocessor Standards Committee of the IEEE Computer Society, IEEE Std 754-2008, dated Jun. 12, 2008, 70 pages. [cited by applicant]
International Search Report and Written Opinion for Application No. PCT/CN2022/095881, mailed Dec. 28, 2022, filed May 30, 2022, 9 pages. [cited by applicant]
Society of Automotive Engineers On-Road Automated Vehicle Standards Committee “Taxonomy and Definitions for Terms Related to Driving Automation Systems for On-Road Motor Vehicles”, Standard No. J3016-201806, dated Jun. … [cited by applicant]
Society of Automotive Engineers On-Road Automated Vehicle Standards Committee, “Taxonomy and Definitions for Terms Related to Driving Automation Systems for On-Road Motor Vehicles,” Standard No. J3016-201609, issued Jan… [cited by applicant]